Browse the docs

Results and analytics

Results for one simulation or coach: attempts, transcripts and human review

Read the results of one simulation or coach, open an attempt's transcript and evaluation, and review, override or delete a result.

Plans: All plansView as Markdown
On this page

Every simulation and every coach has a Results stage in its editor. For a simulation it shows summary numbers, a list of learner attempts, and for each attempt the transcript, the evaluation and its review history. For a coach it shows session counts and the transcript of each session. Every member of the workspace can read these results on every plan. Only owners and admins can review an evaluation, override a result or delete an attempt.

Before you start

What counts as an attempt

The Attempts table lists one row per learner session on the simulation. A session is listed when all of these are true:

  • It started inside the selected date range.
  • The learner sent at least one message or spoke at least once, or the session has an evaluation. Opening the simulation without taking a turn does not create a row.
  • It is not a test run from the editor's preview. Test runs never appear in results.
  • It is not the earlier part of a voice session that dropped and was resumed. A resumed voice session is one attempt, and its transcript includes the turns from before the drop.

Each attempt has a status:

  • Completed: the attempt ended. A voice attempt is completed when the voice session ends for any reason, including the learner hanging up or closing the page. A text attempt is completed when the simulation reaches its end and its evaluation is requested, so text attempts on a simulation with evaluation turned off are not marked completed.
  • In progress: the attempt has not ended and its last message is less than 60 minutes old.
  • Abandoned: the attempt has not ended and has been quiet for longer than that.

Steps

  1. Open the simulation in the editor and select the Results stage. The page is headed Simulation Metrics.
  2. Choose the date range: 24h, 7d, 30d, 90d or All. The page opens on 90d. A range ends at the moment you select it and reaches back that long. Select Refresh to move the range up to now and reload the numbers.
  3. To look at one kind of attempt, use the mode filter (all, text or voice). It changes the summary cards, the chart and the table together.
  4. Scroll to Attempts. The table opens on Completed. Switch to uncompleted (shown capitalized) to see attempts that are in progress or abandoned, or to All.
  5. If the workspace collects identity fields with a fixed list of options, a dropdown for each one appears at the top of the page. Choosing a value narrows the Attempts table and the exports when you are allowed to see learner identity. For members who are not, the dropdown has no effect. It does not change the summary cards or the chart.
  6. Select a row to open Conversation Detail for that attempt.

Result

You see the numbers for the range and mode you chose, a list of attempts with 25 rows per page (Prev and Next move between pages), and a side panel with everything recorded for the attempt you opened.

The summary numbers

The Completion & Score cards are built from evaluations, so they appear once at least one attempt in the range has been evaluated. Until then the section shows No evaluations yet, even when the Attempts table already has rows.

Card What it counts
Completions Evaluations created in the range.
Pass Rate Passed evaluations as a share of scored evaluations.
Avg Score The average score, as a percentage, of scored evaluations. For evaluations older than a workspace retention window the stored point score is used, which equals the percentage only when the criteria total one hundred points.
Unique Learners Distinct learner IDs among the sessions started in the range. A session without a learner ID counts as its own learner.
Total Starts Sessions started in the range.
Completion Rate Completions as a share of Total Starts.

Read them with these rules in mind:

  • Sessions are placed in the range by their start time. Evaluations are placed by the time they were created; with the mode filter on text or voice, an evaluation is counted only when its session also started in the range.
  • Total Starts counts every started session, including a voice session in which the learner never spoke. The Attempts table leaves those out, so the two numbers can differ.
  • Test runs and the earlier parts of resumed voice sessions are left out of every card.
  • Feedback-only evaluations have no score and no pass or fail. They count toward Completions and are left out of Pass Rate and Avg Score. A note under Pass / Fail says how many there are, and both cards show a dash when nothing in the range was scored.
  • The cards use the AI's original result. A human override changes the result shown for that attempt, not the cards. See "Review or override an evaluation" below.

Below the cards:

  • Pass / Fail shows the passed and failed counts as one bar.
  • Score Distribution groups scored evaluations into score bands.
  • Criteria Breakdown shows, for each criterion in the simulation's current evaluation settings, the share of evaluations rated Full, Partial or None. It only uses evaluations that cover every current criterion, and says so when that is fewer than all of them. Feedback-only evaluations are not part of it.
  • Deductions Frequency shows how often each deduction was applied.

The Activity chart plots the range over time. Switch it between Sessions (attempts by start time), Completions (the ones that ended), Credits and Avg score. Credits is only offered while the mode filter is on all.

Credit Usage shows Avg credits / completion (with the figure excluding evaluation credits beneath it), Avg credits / session and Unique learners. These cards only count sessions that used credits. With the mode filter on all, each card also splits the figure into text and voice. When the range starts before per-session credit tracking began, a note gives the date the credit figures start from. For how credits are charged, see Credits, usage and low-credit alerts.

The Attempts table

Attempts are listed newest first.

Column What it shows
Started When the attempt started.
Learner / Session The learner identifier, if one was collected and you are allowed to see it. Selecting it opens that learner's history. Otherwise a short code for the learner or session.
Identity fields One column for each identity field the workspace collects.
Mode Text or voice.
Language The language the attempt ran in.
Status Completed, In progress or Abandoned.
Credits Credits charged to the attempt. A dash means the attempt is older than per-session credit tracking.
Result Pass or Fail with the score as a percentage, Not scored for a feedback-only evaluation, or a dash when the attempt has no evaluation. For an evaluation older than a workspace retention window the stored point score is shown, which equals the percentage only when the criteria total one hundred points.

A Reviewed badge next to a result means a person has reviewed it. When a result was overridden, the Result column shows the adjusted pass or fail and score.

The Export menu downloads the attempts, transcripts or evaluations for the current range, mode and identity filters. The attempt status filter applies to attempts and transcripts only, and the evaluations option is unavailable until the range has an evaluation. See Exporting results, transcripts and evaluations.

Read one attempt

Conversation Detail has these sections:

  • Evaluation: the result (Pass or Fail) with the score, points and percentage, then the written feedback. A feedback-only evaluation shows Feedback only, not scored and the feedback, with no criteria. The section is absent when the attempt was not evaluated.
  • The criteria: each criterion with its rating, the points earned out of the points possible, and the evaluation's reasoning. Where the evaluation quoted the conversation, each quote carries a turn reference (the letter T and the turn number), and a quote that could not be matched to the transcript is marked unverified.
  • Deductions: every deduction in the evaluation, with the points taken off or Not applied.
  • Human review: the review history. See the next section.
  • Summary: a short summary of the conversation, when one has been generated, for example for a linked coach.
  • Details: Mode, Messages, Started and, when present, Learner ID.
  • Transcript: the conversation, labelled with the character's name and Learner. Copy Transcript copies it as text, and Export this attempt (CSV) downloads this one attempt.
  • Credit Breakdown: credits for Chat, for Voice time on a voice attempt, and for Evaluation, with the Session total.

How the transcript is built:

  • Text: the messages exactly as they were sent. Each character reply also shows the credits it used, when credits were recorded for the attempt.
  • Voice: the transcript of what was said, taken from the voice provider's final record of the session when it is available. This is the same transcript the evaluation reads. When it is not available, the transcript is built from the turns recorded during the session. Voice transcripts do not show credits per turn.
  • The panel shows up to 500 stored messages for one attempt.
  • If the workspace has a retention window, the stored messages of an attempt older than the window are deleted, so a text transcript is empty. A voice transcript can still be shown while the voice provider's record of the session is available, because it is read again from that record. See Data privacy, retention and learner data requests for what is kept.

Review or override an evaluation

Owners and admins can record a human review on any scored evaluation. Members see the review history but not the buttons. Feedback-only evaluations cannot be reviewed, because there is no score or pass to confirm or change, and an attempt without an evaluation has nothing to review.

The AI's result is never edited. A review is added to the attempt's history, and the history keeps every entry with the reviewer's name or email and the time.

To confirm a result without changing it:

  1. Open the attempt and find Human review.
  2. Select Mark reviewed.

To change a result:

  1. Select Override result.
  2. For each criterion you disagree with, choose a different rating. The AI's rating is marked "(AI)" in the list.
  3. Leave Overall result on Recompute from ratings, or set it to Pass or Fail.
  4. Enter the reason. It is required and can be up to 2,000 characters.
  5. Select Save override.

An override has to change at least one rating or the overall result.

How the adjusted result is worked out:

  • A changed criterion is re-scored with the points and rating levels that applied when the attempt was evaluated. Criteria you did not change keep their points, and deductions stay as the AI applied them. Editing the simulation's criteria later does not move a past result.
  • With Recompute from ratings, the adjusted score is compared with the simulation's pass threshold at the time of the override to decide pass or fail.
  • Setting Overall result to Pass or Fail decides the outcome directly. If you change no ratings, the score stays the AI's score.
  • You can override the same attempt again. The latest override is the current adjusted result.

Where an override shows and counts:

  • In Human review, the Adjusted result appears next to the AI's result, with the reason.
  • In the Attempts table, the Result column shows the adjusted result with the Reviewed badge.
  • In exports, the adjusted result is added in its own columns next to the AI's result. See Exporting results, transcripts and evaluations.
  • The summary cards, Pass / Fail, Score Distribution and Criteria Breakdown on this page keep using the AI's original result.
  • A score that was already sent to your course or LMS is not changed.

Delete an attempt

Owners and admins can delete a simulation attempt, for example a test run that was made outside the editor's preview and is mixed into real learner data.

  1. In the Attempts table, point at the row and select the trash icon at its end (Delete attempt).
  2. Confirm in Delete this attempt?.

The transcript and the evaluation are permanently deleted, and the attempt stops counting toward the attempts, completions, scores and Credit Usage cards on the page. The Credits measure in the Activity chart still includes the credits it used. Credits the attempt used stay in your usage history. Members do not see the delete control. Coach sessions cannot be deleted from the Results stage.

Human review on an evaluator's results

The same review tools serve evaluators, which score work a learner submits. Open the evaluator in the editor, select the Results stage and expand a row under Recent results:

  • The row shows Pass or Fail with the score as a percentage, or Not scored when the workspace does not allow scored evaluations. As in the Attempts table, the latest override decides the result shown, and a Reviewed badge marks a result a person has reviewed.
  • The expanded result shows the written feedback, each criterion with its rating, points, reasoning and quoted evidence, and the deductions that were applied.
  • Human review works as described above: owners and admins can select Mark reviewed or Override result, members see the history only, and a result that was not scored cannot be reviewed.
  • What was evaluated shows the exact input the result was scored on: the turns of the linked simulation under Conversation, each with the turn reference the evidence quotes point to, and the learner's answers under Submitted answers, each under its field label. A result stored before inputs were recorded has no stored input and says so.

For everything else about evaluators, see Evaluators, scoring learner-submitted work.

Results for a coach

Open the coach in the editor and select the Results stage. The page is headed Coach Metrics. A coach session is an open conversation, so there is no evaluation, score or pass rate.

  • Choose the range: 7d, 30d, 90d or All. The page opens on 90d.
  • The cards show Sessions, Unique learners and Avg credits / session. Sessions counts sessions in which the learner sent at least one message or spoke at least once, and only from the date per-session credit tracking began. A note gives that date when your range starts earlier.
  • Sessions below the cards lists the coach's sessions in the range, newest first, 25 per page. Test sessions from a coach's voice preview and the earlier parts of resumed voice sessions are left out. The list includes sessions in which the learner said nothing, and the Sessions card includes voice-preview test sessions that the list leaves out, so the two numbers can differ in either direction.
  • Each row shows the learner, Voice or Text, the number of messages and the start time. Coach sessions are stored without a learner identifier, so the row reads Anonymous learner.
  • Select a row to open Coach session: the mode, the start time and the transcript, labelled Learner and Coach. Voice transcripts are built the same way as for a simulation, and the same limit of 500 stored messages applies.

Before any session is returned, devlin.ai checks that the coach belongs to your active workspace. For how a coach session continues across slides and restarts, see Coach sessions across slides, reloads, and course restarts.

AI tools over MCP

An AI tool connected through the MCP server can read the same results:

  • It can get the analytics of one simulation or coach over a time window: sessions, unique learners, the last activity and the split between text and voice, plus the pass rate, average score and per-criterion performance of a simulation. These figures use the AI's original results.
  • It can list the most recent sessions of a simulation or a coach, newest first: 20 by default and up to 50. Unlike the Attempts table, this list has no date range and includes sessions a learner opened without taking a turn. Simulation attempts carry the pass or fail, the score, the feedback and any adjusted result from an override. Feedback-only attempts are marked as not scored.
  • It can read one attempt in full: the result, each criterion's rating, points, reasoning and quoted evidence, the deductions that were applied, and the transcript. A very long transcript is cut at the end and flagged as truncated.
  • For an evaluator it can list recent results and read one in full, including the submitted answers as they were scored. Very long answers are shortened and flagged as truncated. The analytics tool covers simulations and coaches only.
  • It cannot mark an evaluation reviewed, override a result or delete an attempt. Those are done by a person in the app.