Skip to main content
An experiment measures calls from exactly two deployed versions of one agent. Your voice platform routes callers consistently to Arm A or Arm B; Zelto never routes traffic or changes a deployed prompt. Use Simulations before deployment. Use Experiments after both versions are deployed and receiving real traffic. Experiments have no automatic end date. They stay active until an operator selects End experiment. Reaching a statistical threshold or finishing an evaluation batch does not end the experiment.

Before you launch

Pilot organizations have full Experiments access, including creation by owners and admins, regardless of their paid plan or rollout setting. Turning Pilot off restores the organization’s plan and rollout requirements without deleting its experiments. Members can view experiments but cannot create them. Set up agent version reporting before configuring an experiment. For LiveKit and custom uploads, send an explicit top-level version on every call. Keep the same agent ID across releases and use distinct labels for prompt or pipeline changes. Confirm calls from both releases appear under the correct versions; a draft candidate cannot be used as an experiment arm. Monitor definitions are frozen at launch, including their version, pipeline, headline stage, threshold, provider, model, labels, and tone. Later monitor edits do not rewrite experiment history. Provider credentials remain operational: a missing or revoked current credential appears as an evaluation failure.
Audio monitors are not supported in Experiments. Their recording pipeline is not compatible with transcript-based experiment evaluation.

Create an experiment

  1. Open Experiments and select New experiment.
  2. In Versions, select one agent and two deployed versions. Review the prompt changes, captured pipeline configuration, and recent call counts. Missing configuration is shown as unavailable. The diff includes all explicit version configuration supplied through registration or versionConfig, including tools and workflows. Older snapshots contain only captured settings.
  3. In Monitors, select the checks to apply to both versions. You can create a monitor or view and edit its configuration in the side panel. Choose a saved group or optionally save your selection as a reusable group.
  4. In Review, select Check recent traffic to inspect eligible calls, counts for A and B, and identity coverage. A name is not required for this check. It previews historical traffic; it does not add those calls to the new experiment’s results.
  5. Enter an experiment name and select Launch experiment. The name is required for launch. The experiment starts immediately using server time; it cannot be backdated or scheduled.
Metadata filtering, caller-identity selection, allocation controls, and notification settings are not part of this creation flow. Your platform remains responsible for routing; Zelto observes the reported versions. If you choose to save a monitor group, it is saved at launch. Reviewing does not save the group. Each experiment keeps a snapshot of its monitor definitions; later source edits or deletions do not change it. Future experiments use the source definitions current at their launch. Keep reporting the actual deployed version on every future call. The two selected version IDs stay fixed: deploying a third version does not replace an arm, and reusing a label for changed code mixes releases into that arm. See the verification checklist. Creating, changing settings, retrying evaluation, concluding, and deleting are owner/admin actions. Members can view experiments and results.

Read the results

Calls are first summarized per caller: repeated calls become one bounded mean per outcome. A caller seen in both arms is excluded from inference and reported as cross-arm leakage. The results table shows caller-level means when available and falls back to call-level descriptive means otherwise. Statistical labels are available only when all safeguards hold:
  • at least 95% of eligible calls have a stable identity;
  • cross-arm leakage is below 1% of identifiable callers; and
  • both arms contain at least 50 uncontaminated unique callers.
At predefined caller-count checkpoints, Zelto displays each arm’s caller-level mean, the B-minus-A effect estimate, and a conservative repeated-look interval. An outcome is Statistically clear only when that interval excludes zero. The result states which arm is numerically higher; it never calls an arm “better,” declares an overall winner, or recommends adoption. Each outcome is a separate 95% analysis. Zelto does not apply a cross-outcome multiplicity adjustment, so interpret every outcome on its own. Possible states are:
  • Collecting — waiting for a predefined checkpoint or minimum sample.
  • Statistically clear difference — the interval excludes zero.
  • No statistically clear difference — the operator ended the experiment without a clear difference.
  • Descriptive only — identity coverage or leakage safeguards are not met.
The call table links each observation to its transcript and shows whether the outcome was evaluated, filtered, or failed. Owners and admins can retry only missing or failed work; completed observations are never re-judged.

End or edit

Select End experiment and confirm to close the call-time window at that moment. Later calls no longer join the comparison. Evaluations for earlier calls may still finish, and results remain available. An ended experiment cannot be reopened. Repeated end requests preserve the original end time. After launch you can edit only the name, description, expected allocation, and notification email addresses. Agent versions, cohort, identity strategy, and outcomes stay frozen. Deleting an experiment does not delete its agent, versions, calls, or source monitors.

Reuse a group through the API or MCP

GET /v1/experiments/monitor-groups and the list_experiment_monitor_groups MCP tool return saved groups with their ordered metricIds. Pass monitorGroupId and those IDs as outcomeMetricIds to create an experiment. To save a new group at launch, pass newMonitorGroupName and your selected outcomeMetricIds instead. Do not send both group fields. The server rejects a group whose members no longer match the selection you reviewed. The explicit ending endpoint remains POST /v1/experiments/{id}/conclude, and the corresponding MCP tool is conclude_experiment.

Legacy experiments

Experiments created with the earlier observational workflow remain available as Legacy observational comparisons. Their results are descriptive only and do not receive the new caller-level statistical labels. The UI does not let you change their versions or monitors, reopen them, or start evaluations. Owners and admins can still edit the name, description, and arm labels, end an active comparison, or delete it.
  • Agent versions — release labels, snapshots, and call attribution.
  • Agents — agent identity and call history.
  • Monitors — the outcomes available to an experiment.
  • Simulations — controlled prompt testing before deployment.