LLM Agents Do Not Reliably Follow Company Policies
A new benchmark that tests whether AI agents can follow long company policies during realistic work
A new benchmark that tests whether AI agents can follow long company policies during realistic work
An LLM-judge approach that brings interpretability and actionability to your scores.