An Evaluator scores content a learner submits, consult notes typed during a call, a completed assessment, a care plan, any open-text work product, against a designer's rubric, and returns a pass/fail verdict, a score, a per-criterion breakdown, and written feedback. It works on its own (submit text, get a scored result) or linked to a simulation, where the conversation transcript and the submitted work can be judged together and the two scores combine into one weighted session result. If your workspace disallows scored evaluations in its Data and privacy settings, results carry written feedback only, with no verdict or score.
Where Evaluators live
- Studio has an Evaluators tab listing every evaluator with its publish status and field count. Creating one takes a name; everything else is configured in its editor.
- The editor uses the same lifecycle stages as the simulation and coach editors: Design (with Fields, Simulation Link, Evaluation, and Form sections), Test (a live preview of the hosted form), Publish (embed key, hosted page link, Storyline guide, an AI-agent prompt, API example), and Results.
- The Results stage shows the most recent real submissions (test runs are excluded), each expandable to the full criteria breakdown with evidence quotes and the exact input that was evaluated, plus a CSV export of all results.
Fields and stages
- Fields are the declared in-roads, up to 12 per evaluator. Each has a key that integrations submit content under, a per-field length cap (default 2000 characters, up to 8000), a required flag, and optional designer guidance that tells the scoring judge what good looks like (guidance is never shown to learners). The label learners see on the hosted form is set in the Form section; a field without a label shows its key as a readable name. One session can hold up to 32000 characters across all fields.
- Fields can be assigned to a stage (for example pre-work), with up to 5 stages besides the built-in Final stage, and submitted at any point in a session. Everything accumulates and is scored together in one final evaluation, so submission order never matters. If a field is submitted more than once, the latest value is used. Required fields must be filled before the final evaluation can run.
- A field can also be mapped to a course variable name (a Storyline variable, or a variable in an AI-built course), which lets a Storyline course collect the learner's typed work in ordinary text-entry fields and submit it with a single JavaScript trigger.
The evaluation rubric
- In the editor's Evaluation section, criteria earn points and deductions subtract them. Each criterion can optionally be scoped to specific sources, the linked simulation's conversation, or particular submitted fields, so the judge knows where its evidence should come from. A criterion with no scope is judged against everything.
- Every rating the judge gives must cite verbatim quotes from the submission or transcript. The platform checks each quote against the source text and asks the judge to correct any that do not match. If quotes still do not match after that one retry, the result is accepted and those quotes are flagged as unverified.
- The pass threshold is a percentage of available points. Whether written feedback is returned to the learner, and its maximum length, are configurable; turning feedback off also hides the per-criterion breakdown from the learner. When a simulation is linked, the combined-score weighting and combined pass threshold are set here too.
Ways to integrate
- Hosted page: turn on the standalone form in the Design stage's Form section and a published evaluator gets a ready-made form page at its own link: fields, a submit button, and the scored result. Share the link directly or embed it in a course or LMS with an iframe. No code required. The standalone form is off by default; with it off, the link shows a message that the evaluator has no standalone form.
- Storyline: load the bridge script on the first slide, map fields to Storyline variables, and call
submitEvaluator("<embed key>")from a Submit button trigger. Results come back into the course asEval_Pass,Eval_Score, andEval_Feedbackvariables, so slide logic can branch on them; when the evaluator is linked to a simulation in the same course session and both results are in,Overall_PassandOverall_Scorecarry the combined verdict. When the hosted page is embedded in a course that loads the bridge script, it writes the same variables and also reports the evaluator's score into the course's LMS result using the designer's weights, if your plan includes LMS reporting. - AI agents: the Publish stage has a "Use with Claude / AI agents" section with a one-click-copy prompt that teaches Claude Design, Claude Code, or any AI agent the evaluator's fields, stages, endpoint, and response shape, so an AI-built course or app can integrate it without hand-written JSON. The prompt reflects the last saved version, so copy it again after changing fields.
- API: an application can submit fields directly to the public submit endpoint with the evaluator's embed key. Use unguessable session IDs (UUIDs); staged submissions accumulate and the final call returns the scored result. The evaluator must be published for the endpoint to accept submissions.
- Test before publishing: the editor's Test stage shows the hosted page in preview mode, drafts included and whether or not the standalone form is on; submissions made there are marked as tests and never appear in results or exports.
Linking to a simulation
Linking an evaluator to a simulation lets one session cover both a conversation and written work: the final evaluation can see the call transcript ("the plan is consistent with what was discussed") unless you turn that option off, and the simulation's score and the evaluator's score combine into a single weighted session result with its own pass threshold. Weights are relative and don't need to sum to 100. The simulation is still evaluated on its own; the combined result is the summative verdict, and it exists only once both the simulation and the evaluator have been scored.
Billing
Each final evaluation run uses credits based on the AI usage of that run, which grows with the amount of submitted text and transcript. Stage submissions that only store content don't run the AI and cost nothing; repeat submissions of an already-scored session return the stored result without a new run. If the workspace is out of credits, the final submission is refused and nothing is scored.