Recruiter Bench
Multi-stakeholder recruiting workflows, including recovery when people behave like people.
10+ hard-band scenarios 2 difficulty bandsEnvironments · verifiers · replay
Train and evaluate agents on long-horizon workflows authored by verified practitioners, on resettable software, with a human baseline you can measure against.
Every verifier reproduces the authoring expert's own recorded run before it ships
The most valuable work never made it to the internet.
We turn the edge cases, trade-offs, and judgment behind real work into RL environments where agents learn how the work actually gets done.

Supply
Enterprise-grade, self-hosted and resettable, with full action telemetry.
resettable 02Multi-hour workflows elicited from working practitioners and confirmed by them.
expert-authored 03Expert-weighted rubrics compiled into code that grades the state, not the prose.
deterministic 04Frozen cuts baselined against the authoring expert's own recorded cold run.
baselinedQuality control
Most agent benchmarks reward an answer that reads correctly. We grade the state the work leaves behind, and no grader ships until it reproduces the expert who authored the task.
Benchmarks
Build, fix, harden and migrate integrations against a live applicant tracking system.
50 public tasks 15 fictional vendorsMulti-stakeholder recruiting workflows, including recovery when people behave like people.
10+ hard-band scenarios 2 difficulty bandsRequest supply