What an experiment contains
The same prompt version cannot be in both arms. If it were, the same call could
be counted on both sides, which would make the comparison unreliable.
Create an experiment
- Open Experiments in the sidebar and select New experiment.
- Give the experiment a clear name. Add a description that records the change you are testing and the outcome that would make you adopt it.
- Optionally choose a Start of window. Leave it blank to begin collecting calls now, or back-date it to include earlier calls from the versions you selected.
- Build Arm A. Select an agent, then select one or more deployed prompt versions. You can choose another agent and add its versions to the same arm.
- Build Arm B the same way. A version already in Arm A is unavailable in Arm B, and vice versa.
- Select the monitors you want to compare, then select Create experiment.
Choose useful arms
An arm is a population of calls, not a label for an idea. Keep the comparison focused on the prompt change you want to evaluate.- Use Control for the deployed version that represents current behavior and a descriptive label for the alternative, such as Shorter opener.
- Keep each version in one arm only.
- Include only versions whose calls are relevant to the question. Combining unrelated prompts or agents can obscure the effect you are trying to measure.
- Pick monitors that express the desired outcome. For example, a warmer greeting may be compared on a monitor for caller frustration and one for successful appointment booking.
Read the results
The results table has one row per attached monitor. Each arm shows its headline number and, beneath it, how many calls that number is based on.
A verdict appears only once both arms have at least 20 judged calls and the gap
is statistically significant. Until then the verdict reads Needs 20 judged
calls on each side; a gap that could plausibly be noise reads No clear
difference rather than naming a winner.
Rates here divide by judged calls, not by every call. A monitor’s own page
divides by every call its agents handled, so the two numbers can differ. A call
with no verdict is unknown, not unaffected — counting it as unaffected would
understate every rate and make it drift as more calls are judged.
Coverage and evaluating missing calls
A call only carries a monitor’s verdict if that monitor was enabled when the call was processed. An experiment that reaches back before a monitor existed therefore starts at low coverage — the share of eligible calls that have a verdict — and its comparison rests on whatever sample happens to be judged. The banner above the results table shows the weakest arm’s coverage and how many calls are still unjudged. Select Evaluate missing calls to judge them. Evaluation runs in the background, newest calls first, and results appear as they land. It only judges calls that are missing a verdict, so running it twice costs nothing the second time. It stays available after an experiment is concluded — concluding fixes which calls are compared, not how many of them have been judged.Run, conclude, or reopen
A new experiment is Running. Calls from its selected versions that occur after the start of the window belong to the appropriate arm as they arrive. Select Conclude when you no longer want new calls to enter the comparison. This sets the end of the window but keeps the experiment and its selected monitors intact. You can still work with the calls already inside the window. Select Reopen if you want the experiment to begin including new calls again. Deleting an experiment removes the experiment and its arm assignments. It does not delete the agents, prompt versions, calls, or monitors it referenced.Experiments and simulations
Use both tools at different points in a prompt-change workflow:- Use Simulations before deployment to test a candidate prompt against controlled, synthetic caller scenarios.
- Deploy the prompt version once it is ready for real traffic.
- Create an experiment to compare real calls from that version with a control version on the monitors that reflect your goal.
Related
- Agents — where prompt versions and their call history live.
- Monitors — the continuous checks you attach to an experiment.
- Simulations — test a candidate prompt before it reaches real callers.

