We build SteerBench, a benchmark for the decision an agent makes before it does something real: act, or hold for a person.
Work scores a single decision. Missions scores a whole job.
The moment before an agent acts: send the message, update the record, or charge the card. It scores one call, proceed or hold for a person, and counts both mistakes.
Whether an AI can hold a job across many sessions. It grades whether the work got done and every action stayed inside the authority the agent held at that step.
Working on agent evaluation? [email protected]