rlsupply · research supply for reinforcement learning integration bench / rev 01 live

Environments · verifiers · replay

The supply layer for
production grade agents

Train and evaluate agents on long-horizon workflows authored by verified practitioners, on resettable software, with a human baseline you can measure against.

env / 0042 verified
domain
Payroll and HRIS
software
Live HRIS instance
horizon
Multi-hour workflow
reset
Per episode, snapshot
seed
Deterministic
baseline
Expert cold run
rev 03 / pinned replay available

Every verifier reproduces the authoring expert's own recorded run before it ships

The most valuable work never made it to the internet.

We turn the edge cases, trade-offs, and judgment behind real work into RL environments where agents learn how the work actually gets done.

A figure running up a flight of steps, caught mid-stride in motion blur

Quality control

Higher reward integrity.

Most agent benchmarks reward an answer that reads correctly. We grade the state the work leaves behind, and no grader ships until it reproduces the expert who authored the task.

graded on
System state
reconciled
Ledgers and totals
checked
Audit trails
enforced
Held invariants

Benchmarks

Public leaderboards. Private holdouts.

View all research
bench / 001 rev 01

Integration Bench

Build, fix, harden and migrate integrations against a live applicant tracking system.

50 public tasks 15 fictional vendors
bench / 002 in development

Recruiter Bench

Multi-stakeholder recruiting workflows, including recovery when people behave like people.

10+ hard-band scenarios 2 difficulty bands

Request supply

Tell us where your model breaks.
We build the environment.