Evaluation criteria are how a simulation scores the learner. Well-designed criteria make the score mean something; vague criteria make it a coin flip that erodes learner trust.
Writing criteria that score reliably
- Describe observable transcript evidence: "Asked at least two open-ended questions before proposing a solution," not "showed curiosity."
- One behavior per criterion. Compound criteria ("greeted warmly and set an agenda") half-pass constantly and produce mushy scores.
- Phrase positively where possible (what good looks like); use deductions for specific disqualifying moves (interrupting, quoting an off-policy discount) rather than folding negatives into every criterion.
- Include the threshold of evidence when it matters: "at least once," "before offering a discount," "in their own words."
Setting the pass threshold
- The default 70% suits most skill-practice simulations: it rewards competence while tolerating one weak area.
- Raise it (80-90%+) for compliance, safety, or certification scenarios where a single critical miss is unacceptable, and pair it with a deduction for that critical miss, worth enough points to pull the score below the threshold, so the score can't average it away.
- Lower it (50-60%) for early-funnel or confidence-building practice where the goal is attempting the behavior, not mastery.
- A threshold is a claim about readiness. If you can't say what passing means on the job, the criteria need work before the threshold does.
Calibrate against real attempts
- After launch, watch the pass rate and per-criterion performance (the simulation's Results tab shows the pass rate and a Criteria Breakdown). A pass rate near 100% usually means criteria are too easy or the character collapses too readily; near 0% means criteria are unrealistic, the scenario is too hard, or a criterion is misfiring.
- Open a few failing attempts from the Results tab and read their transcripts before adjusting anything. The fix is as often the character's behavior or a criterion's wording as the threshold.
- Very low attempt counts (a handful of runs) are too little signal to judge calibration; wait for more attempts rather than over-reading two data points.