# Results for one simulation or coach: attempts, transcripts and human review

> Read the results of one simulation or coach, open an attempt's transcript and evaluation, and review, override or delete a result.

- Plans: All plans
- Canonical: https://docs.devlin.ai/results-and-analytics/results-for-a-simulation-or-coach

Every simulation and every coach has a **Results** stage in its editor. For a simulation it shows summary numbers, a list of learner attempts, and for each attempt the transcript, the evaluation and its review history. For a coach it shows session counts and the transcript of each session. Every member of the workspace can read these results on every plan. Only owners and admins can review an evaluation, override a result or delete an attempt.

## Before you start
- You need access to the workspace that holds the simulation or coach. Results are read from the active workspace only.
- Reviewing, overriding and deleting need the owner or admin role. In a personal workspace you are the owner. See [Members, roles, invitations and the activity log](https://docs.devlin.ai/workspace-and-team/members-and-roles).
- Members of an organization workspace see learner identifiers and identity field values only when the workspace allows it. Otherwise the learner column shows a short session code and the identity field columns show a dash. See [Collecting learner identity](https://docs.devlin.ai/workspace-and-team/learner-identity).
- Scores, pass rates and criteria only exist when the simulation has evaluation turned on. See [Evaluation and scoring: criteria, pass threshold and feedback-only mode](https://docs.devlin.ai/simulations/evaluation-and-scoring).

## What counts as an attempt
The **Attempts** table lists one row per learner session on the simulation. A session is listed when all of these are true:

- It started inside the selected date range.
- The learner sent at least one message or spoke at least once, or the session has an evaluation. Opening the simulation without taking a turn does not create a row.
- It is not a test run from the editor's preview. Test runs never appear in results.
- It is not the earlier part of a voice session that dropped and was resumed. A resumed voice session is one attempt, and its transcript includes the turns from before the drop.

Each attempt has a status:

- **Completed**: the attempt ended. A voice attempt is completed when the voice session ends for any reason, including the learner hanging up or closing the page. A text attempt is completed when the simulation reaches its end and its evaluation is requested, so text attempts on a simulation with evaluation turned off are not marked completed.
- **In progress**: the attempt has not ended and its last message is less than 60 minutes old.
- **Abandoned**: the attempt has not ended and has been quiet for longer than that.

## Steps
1. Open the simulation in the editor and select the **Results** stage. The page is headed **Simulation Metrics**.
2. Choose the date range: **24h**, **7d**, **30d**, **90d** or **All**. The page opens on **90d**. A range ends at the moment you select it and reaches back that long. Select **Refresh** to move the range up to now and reload the numbers.
3. To look at one kind of attempt, use the mode filter (all, text or voice). It changes the summary cards, the chart and the table together.
4. Scroll to **Attempts**. The table opens on **Completed**. Switch to **uncompleted** (shown capitalized) to see attempts that are in progress or abandoned, or to **All**.
5. If the workspace collects identity fields with a fixed list of options, a dropdown for each one appears at the top of the page. Choosing a value narrows the **Attempts** table and the exports when you are allowed to see learner identity. For members who are not, the dropdown has no effect. It does not change the summary cards or the chart.
6. Select a row to open **Conversation Detail** for that attempt.

## Result
You see the numbers for the range and mode you chose, a list of attempts with 25 rows per page (**Prev** and **Next** move between pages), and a side panel with everything recorded for the attempt you opened.

## The summary numbers
The **Completion & Score** cards are built from evaluations, so they appear once at least one attempt in the range has been evaluated. Until then the section shows **No evaluations yet**, even when the **Attempts** table already has rows.

| Card | What it counts |
|---|---|
| **Completions** | Evaluations created in the range. |
| **Pass Rate** | Passed evaluations as a share of scored evaluations. |
| **Avg Score** | The average score, as a percentage, of scored evaluations. For evaluations older than a workspace retention window the stored point score is used, which equals the percentage only when the criteria total one hundred points. |
| **Unique Learners** | Distinct learner IDs among the sessions started in the range. A session without a learner ID counts as its own learner. |
| **Total Starts** | Sessions started in the range. |
| **Completion Rate** | **Completions** as a share of **Total Starts**. |

Read them with these rules in mind:

- Sessions are placed in the range by their start time. Evaluations are placed by the time they were created; with the mode filter on text or voice, an evaluation is counted only when its session also started in the range.
- **Total Starts** counts every started session, including a voice session in which the learner never spoke. The **Attempts** table leaves those out, so the two numbers can differ.
- Test runs and the earlier parts of resumed voice sessions are left out of every card.
- Feedback-only evaluations have no score and no pass or fail. They count toward **Completions** and are left out of **Pass Rate** and **Avg Score**. A note under **Pass / Fail** says how many there are, and both cards show a dash when nothing in the range was scored.
- The cards use the AI's original result. A human override changes the result shown for that attempt, not the cards. See "Review or override an evaluation" below.

Below the cards:

- **Pass / Fail** shows the passed and failed counts as one bar.
- **Score Distribution** groups scored evaluations into score bands.
- **Criteria Breakdown** shows, for each criterion in the simulation's current evaluation settings, the share of evaluations rated **Full**, **Partial** or **None**. It only uses evaluations that cover every current criterion, and says so when that is fewer than all of them. Feedback-only evaluations are not part of it.
- **Deductions Frequency** shows how often each deduction was applied.

The **Activity** chart plots the range over time. Switch it between **Sessions** (attempts by start time), **Completions** (the ones that ended), **Credits** and **Avg score**. **Credits** is only offered while the mode filter is on all.

**Credit Usage** shows **Avg credits / completion** (with the figure excluding evaluation credits beneath it), **Avg credits / session** and **Unique learners**. These cards only count sessions that used credits. With the mode filter on all, each card also splits the figure into text and voice. When the range starts before per-session credit tracking began, a note gives the date the credit figures start from. For how credits are charged, see [Credits, usage and low-credit alerts](https://docs.devlin.ai/plans-and-billing/credits-and-usage).

## The Attempts table
Attempts are listed newest first.

| Column | What it shows |
|---|---|
| **Started** | When the attempt started. |
| **Learner / Session** | The learner identifier, if one was collected and you are allowed to see it. Selecting it opens that learner's history. Otherwise a short code for the learner or session. |
| Identity fields | One column for each identity field the workspace collects. |
| **Mode** | Text or voice. |
| **Language** | The language the attempt ran in. |
| **Status** | **Completed**, **In progress** or **Abandoned**. |
| **Credits** | Credits charged to the attempt. A dash means the attempt is older than per-session credit tracking. |
| **Result** | **Pass** or **Fail** with the score as a percentage, **Not scored** for a feedback-only evaluation, or a dash when the attempt has no evaluation. For an evaluation older than a workspace retention window the stored point score is shown, which equals the percentage only when the criteria total one hundred points. |

A **Reviewed** badge next to a result means a person has reviewed it. When a result was overridden, the **Result** column shows the adjusted pass or fail and score.

The **Export** menu downloads the attempts, transcripts or evaluations for the current range, mode and identity filters. The attempt status filter applies to attempts and transcripts only, and the evaluations option is unavailable until the range has an evaluation. See [Exporting results, transcripts and evaluations](https://docs.devlin.ai/results-and-analytics/exports).

## Read one attempt
**Conversation Detail** has these sections:

- **Evaluation**: the result (**Pass** or **Fail**) with the score, points and percentage, then the written feedback. A feedback-only evaluation shows **Feedback only, not scored** and the feedback, with no criteria. The section is absent when the attempt was not evaluated.
- The criteria: each criterion with its rating, the points earned out of the points possible, and the evaluation's reasoning. Where the evaluation quoted the conversation, each quote carries a turn reference (the letter T and the turn number), and a quote that could not be matched to the transcript is marked **unverified**.
- **Deductions**: every deduction in the evaluation, with the points taken off or **Not applied**.
- **Human review**: the review history. See the next section.
- **Summary**: a short summary of the conversation, when one has been generated, for example for a linked coach.
- **Details**: **Mode**, **Messages**, **Started** and, when present, **Learner ID**.
- **Transcript**: the conversation, labelled with the character's name and **Learner**. **Copy Transcript** copies it as text, and **Export this attempt (CSV)** downloads this one attempt.
- **Credit Breakdown**: credits for **Chat**, for **Voice time** on a voice attempt, and for **Evaluation**, with the **Session total**.

How the transcript is built:

- Text: the messages exactly as they were sent. Each character reply also shows the credits it used, when credits were recorded for the attempt.
- Voice: the transcript of what was said, taken from the voice provider's final record of the session when it is available. This is the same transcript the evaluation reads. When it is not available, the transcript is built from the turns recorded during the session. Voice transcripts do not show credits per turn.
- The panel shows up to 500 stored messages for one attempt.
- If the workspace has a retention window, the stored messages of an attempt older than the window are deleted, so a text transcript is empty. A voice transcript can still be shown while the voice provider's record of the session is available, because it is read again from that record. See [Data privacy, retention and learner data requests](https://docs.devlin.ai/workspace-and-team/data-privacy-and-retention) for what is kept.

## Review or override an evaluation
Owners and admins can record a human review on any scored evaluation. Members see the review history but not the buttons. Feedback-only evaluations cannot be reviewed, because there is no score or pass to confirm or change, and an attempt without an evaluation has nothing to review.

The AI's result is never edited. A review is added to the attempt's history, and the history keeps every entry with the reviewer's name or email and the time.

To confirm a result without changing it:

1. Open the attempt and find **Human review**.
2. Select **Mark reviewed**.

To change a result:

1. Select **Override result**.
2. For each criterion you disagree with, choose a different rating. The AI's rating is marked "(AI)" in the list.
3. Leave **Overall result** on **Recompute from ratings**, or set it to **Pass** or **Fail**.
4. Enter the reason. It is required and can be up to 2,000 characters.
5. Select **Save override**.

An override has to change at least one rating or the overall result.

How the adjusted result is worked out:

- A changed criterion is re-scored with the points and rating levels that applied when the attempt was evaluated. Criteria you did not change keep their points, and deductions stay as the AI applied them. Editing the simulation's criteria later does not move a past result.
- With **Recompute from ratings**, the adjusted score is compared with the simulation's pass threshold at the time of the override to decide pass or fail.
- Setting **Overall result** to **Pass** or **Fail** decides the outcome directly. If you change no ratings, the score stays the AI's score.
- You can override the same attempt again. The latest override is the current adjusted result.

Where an override shows and counts:

- In **Human review**, the **Adjusted result** appears next to the AI's result, with the reason.
- In the **Attempts** table, the **Result** column shows the adjusted result with the **Reviewed** badge.
- In exports, the adjusted result is added in its own columns next to the AI's result. See [Exporting results, transcripts and evaluations](https://docs.devlin.ai/results-and-analytics/exports).
- The summary cards, **Pass / Fail**, **Score Distribution** and **Criteria Breakdown** on this page keep using the AI's original result.
- A score that was already sent to your course or LMS is not changed.

## Delete an attempt
Owners and admins can delete a simulation attempt, for example a test run that was made outside the editor's preview and is mixed into real learner data.

1. In the **Attempts** table, point at the row and select the trash icon at its end (**Delete attempt**).
2. Confirm in **Delete this attempt?**.

The transcript and the evaluation are permanently deleted, and the attempt stops counting toward the attempts, completions, scores and **Credit Usage** cards on the page. The **Credits** measure in the **Activity** chart still includes the credits it used. Credits the attempt used stay in your usage history. Members do not see the delete control. Coach sessions cannot be deleted from the **Results** stage.

## Human review on an evaluator's results
The same review tools serve evaluators, which score work a learner submits. Open the evaluator in the editor, select the **Results** stage and expand a row under **Recent results**:

- The row shows **Pass** or **Fail** with the score as a percentage, or **Not scored** when the workspace does not allow scored evaluations. As in the **Attempts** table, the latest override decides the result shown, and a **Reviewed** badge marks a result a person has reviewed.
- The expanded result shows the written feedback, each criterion with its rating, points, reasoning and quoted evidence, and the deductions that were applied.
- **Human review** works as described above: owners and admins can select **Mark reviewed** or **Override result**, members see the history only, and a result that was not scored cannot be reviewed.
- **What was evaluated** shows the exact input the result was scored on: the turns of the linked simulation under **Conversation**, each with the turn reference the evidence quotes point to, and the learner's answers under **Submitted answers**, each under its field label. A result stored before inputs were recorded has no stored input and says so.

For everything else about evaluators, see [Evaluators, scoring learner-submitted work](https://docs.devlin.ai/evaluators/evaluators-scoring-submitted-work).

## Results for a coach
Open the coach in the editor and select the **Results** stage. The page is headed **Coach Metrics**. A coach session is an open conversation, so there is no evaluation, score or pass rate.

- Choose the range: **7d**, **30d**, **90d** or **All**. The page opens on **90d**.
- The cards show **Sessions**, **Unique learners** and **Avg credits / session**. **Sessions** counts sessions in which the learner sent at least one message or spoke at least once, and only from the date per-session credit tracking began. A note gives that date when your range starts earlier.
- **Sessions** below the cards lists the coach's sessions in the range, newest first, 25 per page. Test sessions from a coach's voice preview and the earlier parts of resumed voice sessions are left out. The list includes sessions in which the learner said nothing, and the **Sessions** card includes voice-preview test sessions that the list leaves out, so the two numbers can differ in either direction.
- Each row shows the learner, **Voice** or **Text**, the number of messages and the start time. Coach sessions are stored without a learner identifier, so the row reads **Anonymous learner**.
- Select a row to open **Coach session**: the mode, the start time and the transcript, labelled **Learner** and **Coach**. Voice transcripts are built the same way as for a simulation, and the same limit of 500 stored messages applies.

Before any session is returned, devlin.ai checks that the coach belongs to your active workspace. For how a coach session continues across slides and restarts, see [Coach sessions across slides, reloads, and course restarts](https://docs.devlin.ai/coaches/coach-sessions-and-restarts).

## AI tools over MCP
An AI tool connected through the [MCP server](https://docs.devlin.ai/integrations/mcp-server-and-ai-tool-connectors) can read the same results:

- It can get the analytics of one simulation or coach over a time window: sessions, unique learners, the last activity and the split between text and voice, plus the pass rate, average score and per-criterion performance of a simulation. These figures use the AI's original results.
- It can list the most recent sessions of a simulation or a coach, newest first: 20 by default and up to 50. Unlike the **Attempts** table, this list has no date range and includes sessions a learner opened without taking a turn. Simulation attempts carry the pass or fail, the score, the feedback and any adjusted result from an override. Feedback-only attempts are marked as not scored.
- It can read one attempt in full: the result, each criterion's rating, points, reasoning and quoted evidence, the deductions that were applied, and the transcript. A very long transcript is cut at the end and flagged as truncated.
- For an evaluator it can list recent results and read one in full, including the submitted answers as they were scored. Very long answers are shortened and flagged as truncated. The analytics tool covers simulations and coaches only.
- It cannot mark an evaluation reviewed, override a result or delete an attempt. Those are done by a person in the app.

## Related
- [The Results dashboard](https://docs.devlin.ai/results-and-analytics/results-dashboard)
- [Learner history and AI summaries](https://docs.devlin.ai/results-and-analytics/learner-history)
- [Exporting results, transcripts and evaluations](https://docs.devlin.ai/results-and-analytics/exports)
- [Criteria breakdown and attempt limits](https://docs.devlin.ai/results-and-analytics/learner-results-and-attempts)
- [Evaluation and scoring: criteria, pass threshold and feedback-only mode](https://docs.devlin.ai/simulations/evaluation-and-scoring)
- [Evaluators, scoring learner-submitted work](https://docs.devlin.ai/evaluators/evaluators-scoring-submitted-work)
- [Collecting learner identity](https://docs.devlin.ai/workspace-and-team/learner-identity)
- [MCP server and AI tool connectors](https://docs.devlin.ai/integrations/mcp-server-and-ai-tool-connectors)
