AgentDock

1.7k
Research

When should an agent act on its own?

We build SteerBench, a benchmark for the decision an agent makes before it does something real: act, or hold for a person.

SteerBench

Two benchmarks for agent judgment

Work scores a single decision. Missions scores a whole job.

SteerBench-Work

The moment before an agent acts: send the message, update the record, or charge the card. It scores one call, proceed or hold for a person, and counts both mistakes.

GitHubPaper

SteerBench-Missions

Whether an AI can hold a job across many sessions. It grades whether the work got done and every action stayed inside the authority the agent held at that step.

Working on agent evaluation? [email protected]