Sentry: Learning to Recover from LLM Agent Failures at Test Time (github.com)

🤖 AI Summary
Sentry has been introduced as an innovative external runtime failure management layer designed specifically for large language model (LLM) agents. This system enhances the reliability of LLMs by detecting execution failures, guiding recovery processes, and learning from successful recoveries—all without disrupting the existing agent environment. Unlike traditional runtime monitors, Sentry evolves by acquiring reusable lessons from verified recoveries, boasting an impressive 81.7% recovery rate across 939 detected failures and achieving a 37% improvement over existing intervention methods in various tasks. The significance of Sentry lies in its ability to streamline LLM operations, making them more robust and efficient while maintaining lower token usage compared to other runtime solutions. With an average reduction of 31.7 seconds per task in SWE-bench Lite and 45.6 seconds in AppWorld, Sentry offers a compelling alternative to existing techniques. Key features include hard and soft repair methods, recovery verification, and online playbook learning that ensures agents learn from their mistakes. This approach not only improves the agents' performance on seen tasks but also equips them with lessons applicable to unseen challenges, highlighting the critical interplay between task-level and failure-level learning in advancing AI capabilities.
Loading comments...
loading comments...