Browse the docs

Design guidance

Designing evaluation criteria and pass thresholds

Well-designed evaluation criteria make a simulation's score meaningful; vague criteria make it a coin flip.

On this page

Evaluation criteria are how a simulation scores the learner. Well-designed criteria make the score mean something; vague criteria make it a coin flip that erodes learner trust.

Writing criteria that score reliably

  • Describe observable transcript evidence: "Asked at least two open-ended questions before proposing a solution," not "showed curiosity."
  • One behavior per criterion. Compound criteria ("greeted warmly and set an agenda") half-pass constantly and produce mushy scores.
  • Phrase positively where possible (what good looks like); use deductions for specific disqualifying moves (interrupting, quoting an off-policy discount) rather than folding negatives into every criterion.
  • Include the threshold of evidence when it matters: "at least once," "before offering a discount," "in their own words."

Setting the pass threshold

  • The default 70% suits most skill-practice simulations: it rewards competence while tolerating one weak area.
  • Raise it (80-90%+) for compliance, safety, or certification scenarios where a single critical miss is unacceptable, and pair it with a deduction for that critical miss, worth enough points to pull the score below the threshold, so the score can't average it away.
  • Lower it (50-60%) for early-funnel or confidence-building practice where the goal is attempting the behavior, not mastery.
  • A threshold is a claim about readiness. If you can't say what passing means on the job, the criteria need work before the threshold does.

Calibrate against real attempts

  • After launch, watch the pass rate and per-criterion performance (the simulation's Results tab shows the pass rate and a Criteria Breakdown). A pass rate near 100% usually means criteria are too easy or the character collapses too readily; near 0% means criteria are unrealistic, the scenario is too hard, or a criterion is misfiring.
  • Open a few failing attempts from the Results tab and read their transcripts before adjusting anything. The fix is as often the character's behavior or a criterion's wording as the threshold.
  • Very low attempt counts (a handful of runs) are too little signal to judge calibration; wait for more attempts rather than over-reading two data points.