LLM Agents Do Not Reliably Follow Company Policies
A new benchmark that tests whether AI agents can follow long company policies during realistic work
A new benchmark that tests whether AI agents can follow long company policies during realistic work
A large-scale study on long-horizon document tasks.
Self-generated agent context files don't help.
Curated skills boost agent performance by 16 points; self-generated ones don't help at all.