Monitor selected companies
Choose Companies, search for the companies you want, and select at least one. Matching uses the company attributed to each call. All agents making calls for those companies are covered automatically, including future agents. Calls without a company, and calls attributed to an unselected company, are excluded. Agent membership changes do not rewrite historical coverage. You can save a monitor for companies that have no agents or calls yet. It shows No matching calls yet until matching results arrive. This applies to every monitor kind, including audio monitors. Change the selection in the monitor’s configuration. History, rates, exports, and alerts use the current selection; stored results remain intact. Sidebar company filters and call metadata filters narrow that selection further. Historical calls are evaluated only when you explicitly choose the existing historical analysis option. A queued backfill cannot add companies beyond the selection you authorized. Deleting the last selected company leaves a monitor that matches no calls. Disabling the Companies feature does not remove or broaden saved scopes. Model comparisons from an older scope are marked stale; run a new comparison to evaluate the current population.What a monitor is made of
A monitor is an ordered pipeline of stages, evaluated in sequence:
Filters run first and short-circuit: a call that fails any filter is dropped
before later stages run, so it costs nothing. This is the cost control — put
cheap filters in front of expensive judgement whenever they can do the job.
Each
ai_analyze stage produces one typed output: a boolean, a number, an
option from a list, a string, a list of strings, or JSON. You designate one
boolean, number, flow, tool-check, or code stage as the headline, and that
value drives the affected rate. A monitor with no headline is extraction-only:
it still records values per call, but trends volume rather than a rate.
For each AI step, choose Include agent prompt and Include call context
(CRM) independently in the editor. Both are off on new steps; existing steps
keep their settings (older steps without explicit settings include the agent
prompt and exclude CRM context). Steps with the same settings share a model
request; different settings use separate requests, reflected in cost estimates.
Flow checks always include the agent prompt and exclude CRM context. A step
does not receive another context group’s AI results. Real-call previews include
available CRM data only for steps that enable it.
A monitor with only filter stages is free to run — passing the filters is
itself the flag. A monitor built only from tool_check and code stages is
also free: the verdict comes from structured
tool-call events already on the transcript and
from your own code, not from a model.
Tool checks
Use a tool check when the question is whether the agent actually did something — calledsubmit_order, booked the slot, looked the customer up —
not whether it said it did.
The stage matches expected tool names on tool turns (case and punctuation
insensitive: submit_order, submitOrder, and submit-order are the same
tool) and classifies the call from those events:
Two rules that matter in production:
- Confirmation language is never success. An optional claim-pattern regex
over agent turns records that the agent said the action completed. That is
a claim, not evidence the tool ran. The order-placed case — “your order has
been placed” with no
submit_orderevent — is flagged as claimed without a successful call. - Missing tool metadata is not a pass. Calls ingested before tool events were captured, or from a provider that does not emit them, get insufficient evidence (no headline score). They are not treated as successful execution and they are not treated as a proven miss.
Code stages
Use a code stage when the rule is deterministic and specific enough that you would rather write it than describe it to a model. You write a small TypeScript snippet (no imports) that runs once per call and returns a value. It is free — no model call — and it runs the same way in the “try it” preview, in live judging, and when backfilling. The snippet receives one argument,input:
exit(value)short-circuits. A plainreturnrecords the value and lets the pipeline continue.exit(value)records the value and stops — every later stage is skipped, including a following AI stage and its cost. Put a cheap code check in front of an expensive AI stage andexitearly when the answer is already known.- Local results before the first AI stage flow into every model group. A
code stage that
returns before the first AI stage contributes its value as a ground-truth signal. Later code results and other groups’ AI results are not added to those shared signals.
input — an infinite loop
or a runaway allocation is stopped, not the platform. Imports are rejected, and
the snippet’s syntax is checked when you save the monitor.
Creating a monitor
From Monitors → New monitor, search or browse the built-in templates. Every organization gets the same catalog. Templates show the values they produce before you choose one, and selecting a template opens its editable prompt so you can tailor it before continuing. Choose Create custom monitor when you want to start from a plain-language description instead. Describe what you want to watch for in a sentence and choose which agents it applies to. That opens the create workspace: an assistant on the left, and the monitor itself on the right. The assistant can pull real example calls, dry-run your filters against them, and propose a complete pipeline. Every proposal is applied live to the editor on the right, which you can also edit by hand at any point — the assistant builds on whatever is currently on screen.Audio evaluations
The Audio evaluations template category contains monitors that score the recording itself rather than its transcript. Use them to track caller audio quality, response latency, and pronunciation or language accuracy. After choosing a template, set its threshold and agent scope, then create it like any other monitor.Try it before you save
The editor’s Try it step runs the rubric for real against calls you choose, before the monitor exists. Add up to 10 calls and press Run test; each one comes back with the monitor’s verdict, the value of every stage, and the model’s reasoning for it. This is the fastest way to tell a working rubric from a plausible-looking one. A common loop is: run the test, spot a call the monitor got wrong, tell the assistant which one and why, and re-run. The assistant sees your last test results, so “the second call is a false positive” refers to something it can actually act on. Three things to know:- Only calls with a transcript can be tested, so those are the only ones the picker offers. A monitor has nothing to judge without one.
- Test runs use the built-in models. If your monitor uses your own OpenRouter, Vercel AI Gateway, or TypeSafe key, save it first and it will run on your key from then on.
- Running a test is owner/admin-only, and so is simulating a call. Creating a monitor is open to every member — its ongoing per-call spend is the point. A test run evaluates up to ten calls per click, with a model request for each context group reached per call, so it is gated. A member sees the Try it step explaining that rather than a control they can’t submit.
Simulating a call
Sometimes you want to test a rubric for something that hasn’t happened yet — a new guardrail, a rare failure mode, a behaviour you’re trying to prevent rather than measure. Simulate a call generates a synthetic transcript against your agent’s live prompt from a one-line scenario (“a caller gets frustrated because the agent can’t find their booking”), and adds it to the test set alongside your real calls. Simulated calls are scratch material: they are judged like any other case, but they are never saved as conversations and never appear in your call history. For a fuller version of the same idea — reusable scenarios, pass/fail assertions, and prompt-vs-prompt comparison — see Simulations.Judging your recent calls
By default a monitor starts empty and only judges calls that arrive from the moment you save it. For a quiet agent that can mean waiting days to see whether it works. Also judge recent calls backfills instead: pick a window of up to 10 days and the monitor is applied to the calls already in it, so it has a trend immediately. The control shows how many calls fall in the window, and roughly what they cost, as you move the slider — because this is the one choice in the flow that spends model budget per call. It is off unless you turn it on, and it is owner/admin-only, like test runs. Members don’t see it. The backfill runs in the background after you save, judging the newest calls first. It stops at 5,000 calls, so an extremely high-volume org gets its most recent history rather than an unbounded bill. The slider says so before you save when your window is over the cap, and quotes the capped number rather than the one it would have to truncate.Setting your reading order
Monitors used to be listed newest-edited first, everywhere. That meant the order moved whenever anyone edited a monitor, so there was nothing to build muscle memory around when reviewing a call. Drag a monitor by its handle on Monitors — in either the Graph or the List view — to set your own reading order. The order applies wherever monitors are listed for you, including the Monitors panel inside a conversation, which is where it matters most: the panel reads top to bottom while you work a call. You can also reorder from the keyboard: tab to a monitor’s handle, press space to pick it up, move it with the arrow keys, then press space to drop it (escape cancels). Three things worth knowing:- The order is yours alone. Rearranging it changes nothing for your teammates, so you can order monitors around what you personally review first.
- A new monitor never disturbs it. Monitors you have not placed sort after the ones you have, so creating a monitor adds it to the end rather than reshuffling an order you set. Reset order on the Monitors page clears your arrangement and returns you to the default.
- Reordering while filtered only moves what you can see. If you have narrowed the page to one agent, dragging rearranges those monitors among the positions they already hold. Everything the filter is hiding keeps its place.
Reading a monitor
On the main Monitors page, use 1d, 7d, 15d, or 30d for a rolling view through today. Choose Custom to select one UTC calendar day or an exact start and end date; future dates are unavailable. The range filters the headline values, Graph, List, and Compare views, and is stored in the URL so you can bookmark or share it. Overview cards and rows show Active/Paused status and scored coverage for the current monitor version. Historical trend charts use all versions and their labeled transcript-bearing-call denominator. Open Monitors → Reports to browse saved reports and download PDFs. Owners and admins can select monitors in Overview and choose Create report, or use Add to report on a monitor. New drafts retain company, metadata and dates; agent/group filters are not saved and are disclosed before saving. Reports accept up to 20 monitors and 31 calendar days. Open a report, choose the previous week, this week or explicit dates, then use Download PDF in its header. CSV/JSON, evaluation details and delivery history are expandable. Weekly delivery is opt-in.Filtering by call metadata
The Metadata bar on the Monitors page narrows the calls every number is computed over, not which monitors are listed. Add a filter row, pick a property key (dotted paths into the call’s provider metadata, such ascampaign or
model.provider; keys seen on your recent calls are suggested), choose an
operator (equals, does not equal, contains, does not contain, numeric
greater than / less than and their or equal forms, is set, is not set, is true, is false), and enter a value. Up to eight rows combine with
AND. These are the same property filters as the
Conversations list.
With a filter active, every affected rate, call count, and chart on the list,
every column on Compare, and a monitor’s Overview trend, stage graphs,
flow funnel, and Results rows count only the calls whose metadata matches.
The filter lives in the URL as ?meta=key:op:value, so a filtered view is
shareable, it stays applied as you switch between Graph and List, change the
time range, or move between a monitor’s tabs, and it follows you from the list
into a monitor and onto Compare. On Compare the bar appears once two or more
columns are selected.
Opening a monitor from Overview or a report keeps the selected dates and filters
in a reading view with current-version coverage and calls. Report drill-down also
retains its saved timezone. Use the return link to resume the same period, or
Open monitor configuration to leave that scope and access the full detail page:
- Overview — the affected-rate trend plus graphs for each chartable pipeline stage.
- Results — every judged call, with its verdict and per-stage values.
- Benchmark — compare the Fast, Balanced, and Advanced model tiers on an automatic sample of calls this monitor already judged.
- Configuration — the pipeline, agents, and model. Editing bumps the monitor’s version; past results keep the version they were judged under, and are not re-judged.
- Alerts — see below.
Alerts
On the Alerts tab, set a threshold and Zelto notifies you when the monitor crosses it: a comparator (above or below), a percentage, a rolling window, and optionally a number of consecutive days the breach must persist before it fires. Delivery goes to a Slack channel, a list of email addresses, or both. Alerts are evaluated on a daily sweep, so expect a first notification within a day of the breach beginning, plus any sustained-days requirement you set.From a finding to a monitor
A finding is a curated issue with a set of linked calls; a monitor is a continuous check. When a finding turns out to describe something recurring, promote it: from the finding’s actions menu, Promote to monitor creates a monitor seeded with the finding’s existing calls, so the trend starts with the history you already have.Over MCP and the agent
Monitors are fully manageable from the MCP server:create_monitor
(with either a plain-language rubric or a starterKey from
list_starter_monitors), create_pipeline_monitor, list_monitors,
get_monitor_results, update_monitor, set_monitor_enabled,
delete_monitor, promote_finding_to_monitor, and save_monitor_alert.
update_monitor renames, pauses, or relabels a monitor, replaces the agents it
scores (agentIds; pass [] for every agent in the organization, audio
monitors included), or replaces its company selection (companyIds, at least
one). agentIds and companyIds are mutually exclusive; omitting both preserves
the saved scope. Creation tools also accept companyIds, and list_monitors
returns the scope and selected companies. Company scope edits update visible
history without deleting stored results or automatically evaluating old calls.
The tool also changes the sample rate (sampleRate, 1–100). To change a
monitor’s rubric, edit it in the app, or delete it and create a new one.
To change an existing AI step’s context, read its stage ID with get_monitor,
then pass stageContext: [{ stageId: "step-id", useAgentPrompt: false, useCallContext: true }] to update_monitor. Omitted options stay unchanged;
all supplied edits apply atomically. Context changes increment the monitor
version for future evaluations; historical results remain unchanged.
Related
- Agents — what a monitor is assigned to.
- Conversations — where a call’s monitor results appear.
- Findings — curated issues, and where promotion starts.
- Simulations — reusable synthetic scenarios with assertions.
- Slack — where alerts land.
TypeSafe models with your own key
Select TypeSafe (BYOK) in the monitor’s model settings, choosejev-latest
(or enter a TypeSafe model ID), and supply your TypeSafe API key. Zelto encrypts
the key per monitor. Leaving the key blank when editing keeps the saved key only
when the provider is unchanged. Switching providers requires a new key.
TypeSafe supports these AI check outputs:
- Yes / no: a Noul probability of at least 50% becomes
true. - Category: Choice returns one of your unique category options.
- Number: enter 2–10 descriptive rubric levels, one per line, from lowest to highest. The fractional level score maps evenly onto your minimum and maximum (defaults: 0–100). A score of 1.5 across three levels maps to 7.5 on a 0–10 range. Use these scores for judgments, not numeric extraction.

