Browse the docs

Results and analytics

Exporting results, transcripts and evaluations

Download attempts, transcripts, evaluations, evaluator results and workspace totals as CSV, and see what each file contains.

Plans: All plansView as Markdown
On this page

devlin.ai exports results as CSV files that open in any spreadsheet. There are three places to export from: the Results stage of one simulation, which gives you that simulation's attempts, transcripts and evaluations; the Results stage of one evaluator, which gives you its results; and the workspace Results dashboard, which gives you totals for every simulation and coach. Most exports follow the filters set on the screen, so set the filters first. Each section below says which filters its file honors.

Before you start

  • Exports are available on every plan, to every member of the workspace who can open the simulation, the evaluator or the dashboard. There is no owner or admin requirement.
  • In an organization workspace, learner identity columns are filled only for people allowed to see learner identity. Owners and admins always are. Members are only when Learner identity visibility in Settings → Data & privacy is set to include all members. See Collecting learner identity.
  • Simulation test runs from the editor, and coach voice previews, are never included in any export on this page. A text conversation with a coach in the editor's preview is not marked as a test, so it counts in the Coaches (CSV) file and appears in a coach transcript export over MCP.
  • Old conversation text can be removed by the workspace's retention setting. See Data privacy and retention.
  • The export of an evaluator's results has its own section below. See Evaluators for how evaluators work.
  • A single coach has no CSV export in the app. Coach totals are in the workspace export, and coach transcripts can be exported by an AI tool over MCP (both are described below).

Steps

To export the results of one simulation:

  1. Open the simulation in the editor and select the Results stage. The export controls are in the header of Simulation Metrics.
  2. Set the mode filter to all, text or voice (the buttons show these words capitalized).
  3. Set the date range: 24h, 7d, 30d, 90d or All. The page opens on 90d.
  4. If your workspace has learner identity fields with a list of options, a dropdown appears for each one. Choose a value to keep only the attempts with that value.
  5. Above the Attempts table, set the status filter to completed, uncompleted or all (also shown capitalized). The page opens on completed, so uncompleted attempts are left out of the attempts and transcripts files until you change it.
  6. Select Export, then Attempts (CSV), Transcripts (CSV) or Evaluations (CSV).

To export the transcript of a single attempt, select the attempt in the Attempts table. In Conversation Detail, next to the Transcript heading, select Export this attempt (CSV). This file ignores the filters and contains only that attempt.

Result

The browser downloads a CSV file named after the simulation: its name in lowercase with hyphens, followed by -attempts.csv, -transcripts.csv or -evaluations.csv. A single-attempt transcript also carries a short piece of the attempt's ID in the name. A file with no matching rows contains only the header row.

All dates and times in every file are in UTC, in ISO format.

How the filters and date range apply

  • The date range ends at the moment the page was loaded, or the moment you last selected Refresh or chose a date range, not at the moment you select Export. Select Refresh first if learners are active right now. All has no start date.
  • Attempts and transcripts are selected by the time the attempt started. Evaluations are selected by the time the evaluation was created, which is usually a little after the attempt ended.
  • The mode filter and the learner identity dropdowns apply to all three files. In the evaluations file, when either is set, an evaluation is included only if its attempt also started inside the date range. The identity dropdowns do nothing for a member who is not allowed to see learner identity.
  • The status filter applies to the attempts and transcripts files only. The evaluations file ignores it.
  • Evaluations (CSV) is greyed out when the selected date range and mode contain no evaluations.

What counts as an attempt

The attempts and transcripts files use the same list as the Attempts table. An attempt is included when all of these are true:

  • It started inside the date range.
  • It is not a test run.
  • The learner sent at least one message or spoke at least once, or the attempt has an evaluation.
  • It is not the dropped half of a voice session that the learner resumed. The resumed session carries the whole attempt.

An attempt is completed when the session has ended and uncompleted otherwise.

Attempts file

One row per attempt, newest first. Columns:

Column What it contains
conversation_id The attempt's ID. Use it to match rows across the three files.
session_id The learner's browser session.
learner_id The random ID given to the learner's browser. In a workspace that has pseudonymized learner IDs, a pseudonym appears here instead. That setting has no control in the app.
learner_identifier The identifier the learner entered or the course passed in. Empty if none was collected, or if you are not allowed to see learner identity.
One column per learner identity field Headed by the field's label, directly after learner_identifier. Empty when the attempt recorded no value. These columns are left out entirely if you are not allowed to see learner identity.
started_at, ended_at When the attempt started and ended. ended_at is empty for an uncompleted attempt.
mode text or voice.
language The language code of the attempt. Empty means the simulation's primary language.
status completed or uncompleted.
total_credits, chat_credits, eval_credits Credits the attempt used in total, the part used by the conversation (the total minus evaluation), and the part used by evaluation.
result pass or fail from the AI evaluation. Empty when the attempt has no evaluation or was evaluated for feedback only.
score_percent The AI evaluation's score as a whole percentage.
feedback The feedback text of the evaluation.
reviewed yes when a person has reviewed the evaluation, otherwise no.
adjusted_result, adjusted_score_percent, review_reason The latest human override, if there is one. The AI's own result and score_percent are never rewritten.
Two columns per evaluation criterion Headed by the criterion's description followed by (rating) and (points). The rating is the level the AI chose. Points are written as earned "of" possible, in words rather than with a slash, so a spreadsheet does not turn them into a date.
Two columns per deduction Headed by the deduction's description followed by (applied) and (points). Applied is yes or no.

Details worth knowing:

  • If an attempt was evaluated more than once, the row shows the latest evaluation.
  • The criterion and deduction columns come from the simulation's evaluation setup as it is now. An attempt that was scored against criteria you have since removed or replaced has empty cells in those columns.

Transcripts file

One row per turn. Columns: conversation_id, started_at (of the attempt), turn_index, speaker, content, mode, credits.

  • speaker is Learner for the learner and the character's name for the simulation. If the simulation has no character name, it is Assistant.
  • turn_index counts the turns of one attempt in order. The first turn is numbered zero.
  • Several character messages in a row are joined into one row. This happens mostly in voice.
  • credits is the AI cost of the character's reply. Learner rows always show zero, and so does an opening message that was written in advance.
  • The file has no learner columns. To see who a transcript belongs to, match conversation_id with the attempts file.
  • An attempt whose messages have been removed by the retention setting still appears in the attempts file but has no rows here.

Evaluations file

One row per evaluation, newest first. Columns: evaluation_id, conversation_id, session_id, result, score_percent, feedback, completed_at, reviewed, adjusted_result, adjusted_score_percent, review_reason.

  • result is pass, fail, or not_scored for an evaluation that gave feedback without a score.
  • completed_at is when the evaluation was created.
  • session_id is shortened in this file. Match rows to the attempts file by conversation_id, not by session_id.
  • The review columns work as in the attempts file: the AI's result stays as it was, and the latest human override sits beside it.
  • An attempt that was evaluated more than once has one row per evaluation.
  • The file has no per-criterion columns and no learner columns. Use the attempts file for those.

Evaluator results file

To export the results of one evaluator, open the evaluator in the editor and select the Results stage. Next to Recent results, select Export CSV. The button appears only when the list shows at least one result.

  • The file contains every result of the evaluator, not only the recent ones in the list. There is no date range.
  • If your workspace has learner identity fields with a list of options, the dropdowns next to the button filter both the list and the file. They do nothing for a member who is not allowed to see learner identity.
  • Results from test submissions are never included.
  • The file is named after the evaluator, in lowercase with hyphens, followed by -results.csv.

One row per result, newest first. Columns:

Column What it contains
session_id The submission session the result belongs to.
learner_id The learner's ID, or a pseudonym in a workspace that has pseudonymized learner IDs.
learner_identifier The identifier collected for the session. Empty if none was collected, or if you are not allowed to see learner identity.
One column per learner identity field Headed by the field's label. Left out entirely if you are not allowed to see learner identity.
scope Which result of the session the row is. final marks the final result of a session.
result pass, fail, or not_scored for a result that gave feedback without a score.
score, max_score, percentage Points earned, points possible, and the score as a whole percentage.
combined_score, combined_result For an evaluator linked to a simulation: the combined score and verdict for the simulation and the submission together. Filled only on the final row of a session, and only once the combined verdict is final.
reviewed yes when a person has reviewed the result. Empty otherwise (this file does not write no).
adjusted_result, adjusted_score_percent, review_reason The latest human override, if there is one. The AI's own result is never rewritten.
criteria All criteria in one cell, separated by a vertical bar: each criterion's description, the rating, and points earned and possible in brackets.
deductions The deductions that were applied, in one cell, each with the points taken off. Deductions that were not applied are not listed.
feedback The feedback text.
completed_at When the result was created.

Criteria are in one cell rather than one column each, because an evaluator's results can span changes to its criteria.

Workspace exports

Open Results, choose a date range (24h, 7d, 30d, 90d or All; the dashboard opens on 30d), then select Export. As in a simulation, the range ends when the dashboard was loaded, last refreshed, or the range was last chosen. Each file is named after the workspace.

Simulations (CSV): one row per simulation in the workspace, including simulations with no activity in the range, ordered by number of sessions.

Column What it contains
simulation_id, name, status Which simulation, and whether it is published.
sessions Attempts started in the range in which the learner sent at least one message.
completions Those sessions that have ended.
unique_learners Different learner IDs among those sessions. A session without a learner ID counts by its session.
evaluations Evaluations created in the range.
scored_evaluations Evaluations that produced a pass or fail. Feedback-only evaluations are not counted here.
pass_rate_percent Passes divided by scored evaluations, as a whole percentage. Empty, not zero, when nothing was scored.
avg_score_percent The average score. Empty when nothing was scored.
credits_used Credits charged to the simulation in the range.
last_activity_at When the most recent session in the range started.

Coaches (CSV): one row per coach, with coach_id, name, status, sessions, unique_learners, credits_used and last_activity_at, counted the same way. Coaches have no evaluations, so there are no pass or score columns.

Daily activity (CSV): one row per day that had any activity, with bucket_start, sessions, completions, coach_sessions, evaluations, scored_evaluations, pass_count, avg_score_percent and credits_used.

  • Days start at midnight UTC.
  • Days with no activity have no row.
  • With All selected, each row is a week rather than a day, and weeks start on Monday.
  • sessions and completions count simulations only. Coach sessions are in coach_sessions.
  • credits_used is every credit the workspace used in that period, not only credits used by simulations and coaches.

The simulations and coaches files hold at most 10,000 rows. If the workspace has more, the file ends with a line that starts with # truncated at and gives the number of rows written and the total.

Human review overrides do not change the workspace files: pass rates and average scores there are the AI's own results.

The dashboard also shows evaluator totals, but the Export menu has no evaluator file. To get evaluator data, export each evaluator's results as described above.

Limits

  • The attempts and transcripts exports of a simulation, and the results export of an evaluator, are each limited to 60 downloads per minute for each signed-in person. Above that the download fails with "Too many requests".
  • The simulation and evaluator exports have no row limit. They include every attempt, evaluation or result that matches the filters, so a long range on a busy simulation can take a while to download.
  • The workspace files are limited as described above.

Text that starts with =, +, - or @

Spreadsheets treat a cell that begins with =, +, - or @ as a formula. Learners and designers can type anything, so devlin.ai puts an apostrophe (') in front of any cell whose text begins with one of those characters, including when spaces or line breaks come first. The same is done for text that begins with a tab or a line break.

What you see: a learner message such as "-thanks, that helps" is written to the file as "'-thanks, that helps". The apostrophe is part of the text in the file, so it shows if you open the file in a text editor or load it into another tool. Nothing else in the cell is changed. If you process the files with a script, strip a leading apostrophe that is followed by one of those characters.

AI tools over MCP

An AI tool connected through the MCP server can run two exports. Both return the CSV text directly to the tool, and both work only for a workspace owner or admin.

  • export_conversations lets the tool fetch the transcripts of one simulation or one coach for a time window. It does not cover evaluators.
  • export_evaluations lets the tool fetch the evaluation results of one simulation, or the results of one evaluator, for a time window.

These files are not the same as the ones the app downloads:

  • The window is the last 30 days unless the tool asks for another one or gives explicit start and end dates.
  • The transcript file has the columns conversation_id, started_at, mode, language, learner_id, turn_index, role and content, with one row per turn. role is user for the learner and assistant for the character or coach. It includes completed and uncompleted conversations and has no status or mode filter. Here turn_index is the position of the stored message, so numbers can be skipped, and consecutive character messages are separate rows.
  • The evaluations file of a simulation has conversation_id, created_at, learner_id, pass, score, percentage, one column per criterion holding the rating, feedback and outcome_basis. It reports the AI's original result. Human review overrides are not included.
  • Neither of these two files contains the learner identifier or the learner identity fields. learner_id is the stored learner ID, without the pseudonym the app's attempts file uses in a workspace that has pseudonymized learner IDs.
  • For an evaluator, the tool gets the same file as Export CSV, limited to the time window and without the identity dropdown filter. It includes the learner identifier and identity fields, and the human review columns.
  • An export is refused when the window holds more than 5,000 conversations, evaluations or evaluator results, when a transcript export would hold more than 5,000 messages or turns in total, or when the finished file is too large to return in one reply. Narrow the time window and try again.
  • The formula protection described above applies to these files too.

Other downloads