Skip to main content
Write a test once, run it against any agent, and get a Pass or Fail verdict with the judge’s reasoning. Tests live in a shared workspace library, run against your agent’s current draft, and never touch what’s published, so you can iterate on a prompt and re-run in seconds.

Test types

Tests run as text (no audio is synthesized), but the agent uses its full draft configuration: the knowledge base is consulted, and the agent can invoke its attached tools. In Single Turn and Tool tests, webhook tools send real HTTP requests, so point them at a staging endpoint. Simulation tests mock tools by default. For end-to-end verification with voice, use preview calls.

Create a test

Tests are workspace-level resources, managed under Library → Tests in the console and shareable across every agent in your workspace.
1

Start a new test

Open Library → Tests and click Add test. Give it a name and pick a Type.
2

Describe what to check

Fill in the fields for that type. They are described in the sections below: Single Turn, Tool, and Simulation.
3

Save

Click Create Test. The test is now in your library, ready to attach to agents.

Single Turn tests

A Single Turn test hands your agent a scripted conversation and asks it to produce the next reply:
  1. You script a conversation history of agent and user messages.
  2. The agent generates the next reply using its current draft configuration and the same language model that answers in live conversations.
  3. An LLM judge scores the reply against your Expectation, optionally calibrated by success and failure examples, and returns a verdict with its reasoning.
Under Conversation, click Add message to build the history the agent sees. Each message is either an Agent or User turn, and at least one must be a user turn. Under Judging, write the Expectation: what a correct reply must do. Optionally click Add example to provide success and failure examples. They calibrate the judge but aren’t required.

Tool tests

A Tool test scripts the conversation the same way, but instead of judging the reply it checks which tool the agent called while producing it. Under Require tool execution, pick a tool from the library and choose Should have been called or Should not be called. Leave the tool empty to check that the agent called no tool at all. Integration tools, such as calendar tools, can only be checked in a Simulation test. Under Tool parameters, optionally add the parameter values the call must carry, typed as string, number, or boolean. The result names the tool the agent actually called.

Simulation tests

A Simulation test does not script the user. Instead, a simulated user plays a role you describe, talks to your agent for up to a set number of turns, and an LLM judge scores the finished conversation against your success conditions.

Write the scenario

The scenario is the simulated user’s brief. Click Insert template to start from the four sections the simulator expects: Two rules make scenarios reliable. Reveal one fact at a time, so the agent has to ask for what it needs instead of receiving everything in the first message. And give an explicit ending, otherwise the conversation runs until the turn limit and the judge has to guess whether the user was done. Max turns caps how many user turns the simulation runs. A Conversation script is optional here. Any messages you add are replayed first, then the simulated user takes over from the last message.

Channel

Channel sets the register the agent speaks in during the simulation, the same one it uses on that channel in production: Web voice, Phone inbound, or Phone outbound. On a phone channel, for example, the agent repeats important numbers back in small groups. Phone outbound speaks in the same phone register as Phone inbound and only {{system.channel}} differs, so if the agent should open an outbound call differently, branch on {{system.channel}} in its prompt. Auto, the default, uses Phone inbound when a phone number is assigned to the agent and Web voice otherwise, so pick Phone outbound yourself for an agent that only places calls. As in production, client tools are not available on phone channels, and transfers only happen on phone channels. {{system.channel}} in the agent’s prompt and in the scenario takes the same value. No real call is placed, so {{system.caller_number}} and {{system.dialed_number}} are empty on every channel.

Success conditions

Add one to ten Success conditions, each a plain sentence describing something that must happen in the conversation, for example The agent confirms the refund amount before closing. Give each a short name, or leave it blank and the first words of the description become the name. The judge reads the whole transcript and marks every condition Success, Failure, or Unknown. A run passes only when every condition succeeds. An Unknown verdict means the transcript did not contain enough evidence either way, and the run is flagged Needs review so you read the transcript before trusting the result. Success and failure examples, when you add them, calibrate the judge across all conditions.

Tool mocks

Simulated conversations usually run many times, so by default tool calls never reach your real endpoints: A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Turn on Return an error to make the tool fail with that result instead, for example to test how the agent handles an outage. A webhook tool set to hide errors shows the agent only that the call failed, as in production. When a tool has several entries, the ones with conditions are checked first, in order, and an entry without conditions is the fallback that answers only when none of them match, wherever it sits in the list. Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as (?i) or (?P<name>...) are refused when you save. Write a named group as (?<name>...), and ignore case with (?i:...) around the part it applies to. Forms that JavaScript reads as literal text without Unicode mode are refused too: Unicode escapes such as \p{L} or \u{1F600} (list the characters in a class instead, for example [A-Za-zÀ-ÿ]) and Python escapes such as \A, \z or \N{...} (write \A as ^ and \z as $). The platform checks required tool call parameters itself, so there back-references, emoji and other characters outside the Basic Multilingual Plane, non-ASCII characters under (?i:...), \S inside a class without \s ([\s\S] is fine) and \c are refused as well. Under Mock all, a tool without an entry in the test answers with its own first mock response. A tool with neither fails, and the agent gets the same error it would get if the tool were unavailable in production. Tools whose result the agent never waits for, client tools that don’t wait for an answer and webhook tools in fire and forget mode, need no mock: as in production, the agent only ever hears their acknowledgement, whether the call ran for real, answered from a mock, or was skipped. Such a client tool can’t be mocked at all. A simulation has no client, so a client tool that waits for an answer always answers from its mock entries, under every strategy including Mock none, and needs one. The test form lists the agent’s tools that would fail this way. To let specific tools reach their real endpoint while everything else stays mocked, add them under Call the real endpoint. Only webhook tools and read-only integration tools can go there. Every other tool still needs a mock, so a tool you attach to the agent later fails instead of reaching its endpoint until you mock it. A tool there can still have mock entries: a matching entry answers, and any other call reaches the endpoint. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its Confirm write tool. Mock both, the way production answers them. If you mock only the write tool and the model then calls Confirm write, that call goes unanswered and the run needs review. A run in which a tool call went unanswered never passes: the tool had no mock, or none of its entries matched and neither the tool’s own mock response nor Call the real endpoint could answer. The run fails with Needs review and No mock badges and lists the unanswered calls, even when the judge and every assertion passed, because the conversation left the path it would take in production. A call with no matching entry means either the agent passed arguments you didn’t expect or the entry’s conditions are too narrow. Calls to a tool the test forbids are not listed, they are the agent’s failure and the assertion reports them. The judge is told the call failed because the test prepared no answer for it, and judges the agent as usual without treating the call as having run. Add or widen the mock and run the test again. If you are used to tools without a mock calling their real endpoint, your first runs will show No mock on those calls. Add a mock for each tool the test form lists, or add read-only tools to Call the real endpoint. Mocks and assertions apply to the tools attached to the agent under test. A tool the agent does not have is never mocked and counts as never called: a required call on it fails with This tool is not on the agent, while a forbidden tool or a maximum-only check passes. Client tools on a phone channel are treated the same way. This keeps a shared guard such as “never call issue_refund” passing on agents that cannot call the tool.

Assertions

Under Advanced, add deterministic checks that run alongside the judge:
  • Required tool calls: the agent must call the tool, optionally with parameters that match, between a minimum and a maximum number of matching calls. The default is at least once. Set the maximum to 1 to catch a duplicate booking, or set both the minimum and the maximum to 0 to forbid calls with those arguments.
  • Forbidden tools: the agent must not call the tool.
  • Ended by: who ended the conversation, the agent, the user, a transfer, or any of them. A transfer only happens on a phone channel, so on Web voice a transfer check always fails.
A failed assertion fails the run regardless of the judge’s verdict. Each assertion shows Passed or Failed with a short detail in the result.

Repeat count

Set Repeat to run the same scenario up to 20 times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent’s Tests page shows how many repeats passed out of those that reached a verdict, for example 2 / 3 passed, or 2 / 2 passed · 1 error when one repeat ended in Error. While repeats are still running, the row counts them instead, for example 1 passed · 2 running. Opening the test from that page lets you switch between the finished repeats of the latest batch to read each conversation. The test opens on the first failed repeat. Repeats only apply to Run all. Running the test from its Run tab, or re-running one test from the row menu, runs the conversation once.

Attach tests to agents

A test only runs against agents it’s attached to. Attach from either side:
  • From the library: open the test’s Access tab and toggle it on for each agent.
  • From the Builder: on the agent’s Tests page, click Add tests and pick from the library.
One test can be attached to many agents, and each agent keeps its own last result. Remove from agent detaches the test from that agent only. Delete in the library removes the test from all agents.

Run a single test

Open the test and switch to its Run tab. Pick an agent that has access, then click Run test. Single Turn and Tool tests answer within seconds. A Simulation test runs in the background, shows Queued and then Running, and can take a few minutes. The verdict card shows Pass or Fail with the time the run took, followed by details for the test type: A Simulation run with an Unknown condition fails and carries a Needs review badge. If a run can’t complete, the card shows Error with a message that says whether the agent caused it, for example its custom LLM endpoint stopped responding, or the platform did, for example our language model did not answer. A conversation that broke off or hit its time limit is never judged, so it can’t pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged.

Run every test for an agent

On the agent’s Tests page in the Builder, click Run all. Each row moves through queued → running → Pass/Fail (or Error if a run can’t complete). Simulation tests run once per repeat and the row shows how many repeats passed. The page header summarizes the latest batch, for example 4 passed, 1 failed on last run, with the batch pass rate: green at 100%, amber from 80%, red below. The pass rate counts passed runs out of passed and failed ones, so a run that ended in Error does not lower it. Use the row menu to re-run a single test, edit it, or remove it from the agent.

Tests run against the draft

Tests always exercise the agent’s latest draft configuration, including unpublished changes to the system prompt. That makes the loop fast:
1

Edit the draft

Change the system prompt or first message in Configuration.
2

Re-run

Click Run all.
3

Publish when green

Once results look right, publish the draft. Running tests never publishes anything.

Limits

Going further

Preview calls

Talk to your draft agent live, with voice, tools, and transcripts.

Versions & publishing

How drafts become immutable published versions.

Configuration

System prompt, first message, and everything tests exercise.

Core concepts

Agents, drafts, sessions, and the workspace library model.