Browse the docs

Design guidance

Measuring whether a simulation worked

Decide during design what "worked" means, so every simulation ships with a measurement story.

On this page

Most designers are asked "did the training work?" after launch, without having decided what "worked" means. Settle it during design. It takes two questions, and every simulation ships with a measurement story built in.

Three levels, cheapest first

  1. In-simulation evidence (automatic once evaluation is on). Pass rate, average score, and per-criterion performance. devlin.ai records these for each completed attempt that is evaluated, and shows them on the simulation's Results tab. Evaluation is off by default: turn it on and add at least one criterion, or nothing is graded. Each evaluation uses credits. Feedback-only evaluations give written feedback with no pass/fail or score, so they are left out of the pass rate and average score. Your own test runs are not counted. Trends matter more than snapshots: a learner who fails on attempt one and passes by attempt three is practice working. This level answers "can they do it in the scenario?"
  2. Behavior on the job. Did the practiced behavior show up in real work? Sources the org usually already has: QA/call scores, manager observation checklists, mystery shops, audit findings. Pick ONE existing signal rather than inventing new instrumentation.
  3. The business number. The metric that motivated the training: escalation rate, close rate, ramp time, complaint volume, safety incidents, CSAT. Slow-moving and confounded. Frame it as directional evidence, not proof.

Getting to a metric when you don't have one

  • Normalize it first: most requests arrive as a topic ("we need negotiation training"), not a metric. That's a normal starting point, not a gap.
  • Work backward from pain: "What goes wrong today when someone does this badly?" The cost of that failure is the candidate metric.
  • List 2-3 candidates inferred from the goal and let stakeholders react. Recognizing the right metric is easier than producing one.
  • Whatever they choose (or don't), agree on the level-1 proxy: which evaluation criteria stand in for the outcome, and what pass rate would satisfy them. That way there is ALWAYS something to report.

Design implications

  • Success measures belong in the build brief: evaluation criteria should be the best available in-scenario evidence for the chosen signal.
  • Set a review checkpoint: after ~25-50 attempts, read the Criteria Breakdown and the average score over time on the Results tab, open failing attempts to read their transcripts, then tune difficulty or criteria.
  • For stakeholder reporting, pair the pass rate and the average score trend with 2-3 transcript excerpts showing the target behavior. You can copy a transcript from an attempt or export transcripts as CSV. Numbers plus evidence beats either alone.