RL environments built from real expert work.

Train and evaluate agents on long-horizon workflows authored by verified practitioners, using resettable software, deterministic verifiers, and human performance baselines.

A winding road traversing a mountain valley, traced by a single route line

The most valuable work never made it to the internet. It lives in how a payroll specialist closes a period, how a claims adjuster reads a file, how a dispatcher talks a rate down on a live call.

Environments compound. Datasets deplete. We turn expert workflows into environments you can rerun, grade, and improve against, not one-time datasets.

Reward integrity is the product.

Many agent benchmarks reward outputs that look correct without verifying the resulting system state. Ours grade with code: state assertions, reconciled ledgers, and audit-trail checks. Model-graded criteria are capped at a fifth of any corpus, labeled, and expert-audited.

Every verifier must reproduce the authoring expert's own recorded run before it ships.

  • db_assert

    State assertions against the environment's own database. The record either exists in the right state or it does not.

  • numeric_reconcile

    Totals, ledgers, and registers reconciled within expert-set tolerance bands. The ledger ties or the episode fails.

  • file_diff

    Produced artifacts diffed against acceptable outcome sets, not a single golden file.

  • audit_log_assert

    The path matters. Required intermediate actions are checked in the environment's audit trail.

  • invariant_assert

    Mandatory invariants that must hold at every step. Breaking one is a disqualifying condition with a terminal penalty.

From working expert to running environment

Every step hardens real expert work into something a lab can train against.

  1. 01

    Source the practitioner

    Our private expert network reaches verified specialists in occupations labs cannot staff. Identity and skill verification gate every cohort.

  2. 02

    Elicit the work

    We capture how the work actually gets done. The decisions, the exceptions, the tricks that never made it to the internet.

  3. 03

    Author in the sandbox

    The expert walks their own tasks inside a recorded environment on real software. That cold run becomes the human baseline and the ground truth.

  4. 04

    Compile the reward

    Rubric weights compile into shaped reward functions. Checkpoints become intermediate signals, disqualifiers become terminal penalties.

  5. 05

    Verify deterministically

    Every verifier must pass the authoring expert's own recorded run before it ships. If the check cannot reproduce the expert, it does not ship.

  6. 06

    Refresh on the model clock

    Agent failures reopen the task list at higher difficulty. Every frontier release gets fresh variants, harder bands, and a re-run leaderboard.

Operational expertise that is difficult to recruit inside a lab

Payroll administrators, claims adjusters, dispatchers, and other working specialists are rarely represented in technical hiring networks. We built a sourcing and verification process to reach them directly.

Payroll & HRIS first cohort live
Recruiting operations benchmarked
Medical billing in development
Insurance claims in development
Freight dispatch & TMS in development
Bookkeeping & close in development
Hotel property management roadmap
Dental practice management roadmap
Warehouse management roadmap
Support ticketing roadmap

Tell us where your model breaks. We build the environment.

Sample task packets, environment access, and benchmark holdout licensing for AI researchers, data teams, and labs. Exclusive cuts available.