# Measuring whether a simulation worked

> Decide during design what "worked" means, so every simulation ships with a measurement story.

- Canonical: https://docs.devlin.ai/design-guidance/measuring-simulation-impact

Most designers are asked "did the training work?" after launch, without having decided what "worked" means. Settle it during design. It takes two questions, and every simulation ships with a measurement story built in.

## Three levels, cheapest first
1. **In-simulation evidence (automatic once evaluation is on).** Pass rate, average score, and per-criterion performance. devlin.ai records these for each completed attempt that is evaluated, and shows them on the simulation's Results tab. Evaluation is off by default: turn it on and add at least one criterion, or nothing is graded. Each evaluation uses credits. Feedback-only evaluations give written feedback with no pass/fail or score, so they are left out of the pass rate and average score. Your own test runs are not counted. Trends matter more than snapshots: a learner who fails on attempt one and passes by attempt three is practice working. This level answers "can they do it in the scenario?"
2. **Behavior on the job.** Did the practiced behavior show up in real work? Sources the org usually already has: QA/call scores, manager observation checklists, mystery shops, audit findings. Pick ONE existing signal rather than inventing new instrumentation.
3. **The business number.** The metric that motivated the training: escalation rate, close rate, ramp time, complaint volume, safety incidents, CSAT. Slow-moving and confounded. Frame it as directional evidence, not proof.

## Getting to a metric when you don't have one
- Normalize it first: most requests arrive as a topic ("we need negotiation training"), not a metric. That's a normal starting point, not a gap.
- Work backward from pain: "What goes wrong today when someone does this badly?" The cost of that failure is the candidate metric.
- List 2-3 candidates inferred from the goal and let stakeholders react. Recognizing the right metric is easier than producing one.
- Whatever they choose (or don't), agree on the level-1 proxy: which evaluation criteria stand in for the outcome, and what pass rate would satisfy them. That way there is ALWAYS something to report.

## Design implications
- Success measures belong in the build brief: evaluation criteria should be the best available in-scenario evidence for the chosen signal.
- Set a review checkpoint: after ~25-50 attempts, read the Criteria Breakdown and the average score over time on the Results tab, open failing attempts to read their transcripts, then tune difficulty or criteria.
- For stakeholder reporting, pair the pass rate and the average score trend with 2-3 transcript excerpts showing the target behavior. You can copy a transcript from an attempt or export transcripts as CSV. Numbers plus evidence beats either alone.
