Reliable & Auditable LLM Agents
Evaluation and control for tool-using and multi-agent systems, with an emphasis on action semantics, order sensitivity, progress attribution, replayability, and failure diagnosis.
Research
I investigate the evaluation principles and system mechanisms that make intelligent agents auditable, adaptive, and grounded across long-horizon interactions. My work spans action execution, persistent memory, multi-agent simulation, and multimodal embodied reasoning.
Evaluation and control for tool-using and multi-agent systems, with an emphasis on action semantics, order sensitivity, progress attribution, replayability, and failure diagnosis.
Mechanisms for agents to acquire, update, retrieve, and forget long-term memories while preserving user preferences, temporal consistency, and controllable behavior.
Executable environments for studying coordination, interaction, and emergent behavior through traceable world-state transitions and reproducible counterfactual experiments.
Multimodal models that connect vision, language, and geometry for tiny-object perception, pose understanding, 3D scene reasoning, and embodied decision-making.
Publications
Paper links lead to the corresponding arXiv records or publisher pages.
An evaluation audit of skeleton-based exercise correctness classification, separating estimands, subset structure, and temporal controls.
A typed settlement contract and audit framework for order sensitivity, useful progress, and replay consistency in multi-agent environments.
Input-aware dynamic downsampling for more efficient tiny-object detection in unmanned aerial imagery.
A mobile pose-estimation system that links exercise feedback with responsive visual rewards.