LLM Agents Do Not Reliably Follow Company Policies
A new benchmark that tests whether AI agents can follow long company policies during realistic work
A new benchmark that tests whether AI agents can follow long company policies during realistic work