# Evaluation and scoring: criteria, pass threshold and feedback-only mode

> How a simulation is evaluated when it ends, from criteria, rating scales and deductions to the pass threshold and feedback-only mode.

- Plans: All plans
- Canonical: https://docs.devlin.ai/simulations/evaluation-and-scoring

Evaluation has an AI judge read the finished conversation against criteria you write, then produce written feedback and, in scored mode, a score and a pass or fail. Evaluation is available on every plan. Attaching reference documents for the judge needs a plan that includes knowledge (Core, Team, Enterprise).

All of the settings on this page are in the simulation editor, in the **Design** stage under **Evaluation**.

## Turning evaluation on
Evaluation is off for a new simulation. Check **Enable evaluation** to turn it on. Each evaluation uses credits in addition to the conversation itself.

An evaluation only runs when evaluation is on and the simulation has at least one criterion. With evaluation on and no criteria, the simulation ends without a result.

## Criteria
Under **Evaluation Criteria**, each criterion is one thing the learner should demonstrate. Use **+ Add criterion** to add one. A criterion has:

- A description of up to 1000 characters. Learners and anyone reading results can see this text, so write it for them. See [evaluation criteria design](https://docs.devlin.ai/design-guidance/evaluation-criteria-design) for how to write criteria the judge can apply consistently.
- A point value. The editor adds up the point values and shows the total possible below the list.
- A rating scale (optional, see below).
- Scoring guidance (optional, see below).

## Rating scales
When the judge scores a criterion it picks one level from that criterion's rating scale. Each level has a label, which appears on results, and a percent of the criterion's points that the learner earns for it.

A simulation that sets no scale uses the standard one: "No credit", "Partial credit" (half the points) and "Full credit".

Open **Rating scale** above the criteria list to set a default scale for the whole simulation, or open **Rating scale** inside a criterion to give that criterion its own. A criterion without its own scale follows the simulation default, and follows the standard scale when there is no default. Editing any level of an inherited scale creates a custom scale at that spot. The reset link under the levels returns to the inherited scale.

Rules for a scale:

- It has between 2 and 6 levels. **+ Add level** inserts one just below the top level.
- Every level needs a label, up to 60 characters.
- Percents are whole numbers, and each level must be worth more than the one before it.
- The last level must be worth the criterion's full points. The first level may be worth more than nothing, which guarantees every learner some credit.

Levels are always ordered by what they are worth, so there is no reordering. The editor shows the first rule a scale breaks. A scale that still breaks a rule when it is saved is ignored when the simulation is evaluated: the criterion is judged on the simulation default, or on the standard scale.

After you change the simulation default, criteria with their own scale keep it. **Apply to all criteria** replaces those custom scales with the default, after a confirmation. Scoring guidance written on a level is kept wherever the default scale has a level with the same name or value.

## Scoring guidance for the judge
**Scoring guidance** is text that only the AI judge reads. Use it to tell the judge how to apply a criterion without changing the wording learners see. It is available on each criterion, on each deduction and on each level of a rating scale, with up to 2000 characters in each. Guidance is never shown to learners and is not copied into stored results.

## Deductions
Under **Deductions**, each deduction is a behavior that costs the learner points, with a description (the same length limit as a criterion) and a point value. Use **+ Add deduction** to add one. The judge decides for each deduction whether the learner did it. A deduction is all or nothing: when it applies, its full point value is subtracted.

## Pass threshold
Under **Pass Threshold**, **Passing score** is the percent of the total possible points a learner needs to pass. The default is 70 percent. The editor shows how many points that is for your current criteria.

## How the score is computed
The judge never does arithmetic. It picks a level for each criterion and says whether each deduction applies, and devlin.ai computes the rest:

1. Each criterion earns its level's percent of the criterion's points, rounded to the nearest whole point, with halves rounded up.
2. The earned points are added up and the applied deductions are subtracted. The score cannot go below zero.
3. The score is divided by the total possible points to get a percentage. Deductions do not change the total possible.
4. The learner passes when the percentage is at or above the pass threshold. The comparison uses the exact percentage. The percentage shown on results and sent to a course is rounded to the nearest whole number.

## Feedback-only mode
Under **Feedback**, check **Feedback only (no score)** to keep the written feedback and drop the numbers. The judge still reads the conversation against your criteria, but no points, percentage or pass or fail are computed, shown, stored or sent to your course. No per-criterion results are kept either, so results show the written feedback only.

While feedback-only mode is on, the point values, rating scales and pass threshold are dimmed in the editor. They stay saved, so turning the setting off brings scoring back without rebuilding anything.

What learners see in each mode, and how many attempts they get, is covered in [criteria breakdown and attempt limits](https://docs.devlin.ai/results-and-analytics/learner-results-and-attempts).

## Disallowing scores for a whole workspace
A workspace can force every simulation into feedback-only mode, for example while a legal or privacy review of AI-scored performance is pending. Workspace owners and admins set this in Settings, under **Data & privacy**, in the **AI scoring** card: **Disallow scored evaluations (feedback only everywhere)**. It is available on every plan. Other members see a note that these settings are managed by the owner and admins.

While it is on:

- Every evaluation in the workspace runs in feedback-only mode. This is enforced each time an evaluation runs, whatever the simulation's own setting says.
- In the simulation editor, **Feedback only (no score)** shows checked and cannot be changed, with a note explaining why.
- Each simulation's own setting is left untouched. When the workspace setting is turned off again, simulations follow their own scoring settings.

## Reference materials for the judge
Under **Reference materials** you can attach policies or guides so the judge can check whether what the learner said is accurate. Choose documents or whole knowledge bases from your library, or upload new ones. Changes to the attached documents save automatically.

These documents are only for scoring. The character never reads them, and documents attached for the character in the **Knowledge** section are not read by the judge unless you attach them here too.

Attaching documents is not enough on its own. Turn on **Let the judge reference source materials while scoring**, which is off by default. When it is on, evaluations may use more credits. The judge uses the materials to check accuracy and never quotes them as evidence of what the learner said. If the materials cannot be retrieved for an evaluation, the evaluation still runs without them.

**Draft criteria from documents** suggests criteria, with scoring guidance and point values, based on the attached documents. It uses credits. Nothing is added until you choose **Add** on a suggestion or **Add all**, and added criteria are not saved until you save the simulation. Documents that are still processing are left out, and the editor says so.

## When evaluation runs
Evaluation runs once, when the simulation ends. See [how a simulation ends](https://docs.devlin.ai/simulations/ending-a-simulation) for what ends it. The judge is also told how long the session lasted, so a criterion can refer to duration.

If the evaluation request fails, the simulation retries it once automatically. If it still fails, the course variable that signals a finished evaluation (`Sim_EvalComplete` by default) is set to `failed` instead of `true`, so course logic can branch instead of waiting. This applies to text and voice simulations.

In a text simulation the learner sees "Results loading" while the evaluation runs, and "Results could not be loaded." with a **Retry** link if it fails.

An evaluation does not run when the workspace is out of credits. The simulation treats that as a failed evaluation.

A conversation is evaluated only once: asking again returns the stored result. devlin.ai also checks in the background, for a limited time, for evaluations that were interrupted. A voice session that completed without a result (for example because the learner closed the tab) is evaluated then, and so is a text simulation whose evaluation started and then failed. The result appears in the simulation's results when it finishes.

## Evidence in results
For every level above the lowest one, the judge has to support its rating with a quote from the conversation. devlin.ai checks each quote against the transcript, and asks the judge to correct itself once if a quote does not match.

In the simulation's **Results** stage, opening a scored attempt shows each criterion with its level, points, the judge's reason and the supporting quotes, each tagged with the turn it came from. A quote that still could not be matched exactly is marked "unverified". Feedback-only attempts have no per-criterion results, so they show no quotes.

## Related
- [Evaluation criteria design](https://docs.devlin.ai/design-guidance/evaluation-criteria-design)
- [Criteria breakdown and attempt limits](https://docs.devlin.ai/results-and-analytics/learner-results-and-attempts)
- [How a simulation ends](https://docs.devlin.ai/simulations/ending-a-simulation)
- [Building a simulation](https://docs.devlin.ai/simulations/build-a-simulation)
