Reliable Tool-Using Agents
Question. How can runtime defenses block unsafe tool actions while preserving useful task completion—and how should recovery be evaluated after an intervention?
Contribution. Designed evaluation of safety, task completion, structured block feedback, and post-intervention recovery; integrated online action interception and performed leakage auditing; co-developed the benchmark, evaluation pipeline, and closed-loop trace-replay studies.
Ongoing collaborative research. No publication status, private research artifact, or sole-ownership claim is implied.