Before you launch
Pilot organizations have full Experiments access, including creation by owners and admins, regardless of their paid plan or rollout setting. Turning Pilot off restores the organization’s plan and rollout requirements without deleting its experiments. Members can view experiments but cannot create them. Set up agent version reporting before configuring an experiment. For LiveKit and custom uploads, send an explicit top-levelversion
on every call. Keep the same agent ID across releases and use distinct labels
for prompt or pipeline changes. Confirm calls from both releases appear under
the correct versions; a draft candidate cannot be used as an experiment arm.
Monitor definitions are frozen at launch, including their version, pipeline,
headline stage, threshold, provider, model, labels, and tone. Later monitor edits
do not rewrite experiment history. Provider credentials remain operational: a
missing or revoked current credential appears as an evaluation failure.
Create an experiment
- Open Experiments and select New experiment.
- In Versions, select one agent and two deployed versions. Review the
prompt changes, captured pipeline configuration, and recent call counts.
Missing configuration is shown as unavailable. The diff includes all explicit
version configuration supplied through registration or
versionConfig, including tools and workflows. Older snapshots contain only captured settings. - In Monitors, select the checks to apply to both versions. You can create a monitor or view and edit its configuration in the side panel. Choose a saved group or optionally save your selection as a reusable group.
- In Review, select Check recent traffic to inspect eligible calls, counts for A and B, and identity coverage. A name is not required for this check. It previews historical traffic; it does not add those calls to the new experiment’s results.
- Enter an experiment name and select Launch experiment. The name is required for launch. The experiment starts immediately using server time; it cannot be backdated or scheduled.
Read the results
Calls are first summarized per caller: repeated calls become one bounded mean per outcome. A caller seen in both arms is excluded from inference and reported as cross-arm leakage. The results table shows caller-level means when available and falls back to call-level descriptive means otherwise. Statistical labels are available only when all safeguards hold:- at least 95% of eligible calls have a stable identity;
- cross-arm leakage is below 1% of identifiable callers; and
- both arms contain at least 50 uncontaminated unique callers.
- Collecting — waiting for a predefined checkpoint or minimum sample.
- Statistically clear difference — the interval excludes zero.
- No statistically clear difference — the operator ended the experiment without a clear difference.
- Descriptive only — identity coverage or leakage safeguards are not met.
End or edit
Select End experiment and confirm to close the call-time window at that moment. Later calls no longer join the comparison. Evaluations for earlier calls may still finish, and results remain available. An ended experiment cannot be reopened. Repeated end requests preserve the original end time. After launch you can edit only the name, description, expected allocation, and notification email addresses. Agent versions, cohort, identity strategy, and outcomes stay frozen. Deleting an experiment does not delete its agent, versions, calls, or source monitors.Reuse a group through the API or MCP
GET /v1/experiments/monitor-groups and the list_experiment_monitor_groups
MCP tool return saved groups with their ordered metricIds. Pass
monitorGroupId and those IDs as outcomeMetricIds to create an experiment.
To save a new group at launch, pass newMonitorGroupName and your selected
outcomeMetricIds instead. Do not send both group fields. The server rejects
a group whose members no longer match the selection you reviewed.
The explicit ending endpoint remains POST /v1/experiments/{id}/conclude, and
the corresponding MCP tool is conclude_experiment.
Legacy experiments
Experiments created with the earlier observational workflow remain available as Legacy observational comparisons. Their results are descriptive only and do not receive the new caller-level statistical labels. The UI does not let you change their versions or monitors, reopen them, or start evaluations. Owners and admins can still edit the name, description, and arm labels, end an active comparison, or delete it.Related
- Agent versions — release labels, snapshots, and call attribution.
- Agents — agent identity and call history.
- Monitors — the outcomes available to an experiment.
- Simulations — controlled prompt testing before deployment.

