> ## Documentation Index
> Fetch the complete documentation index at: https://hanabiaiinc-him188-rename-multi-turn-simulation.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Agent Tests

> Script a conversation or simulate a user, then let an LLM judge score your agent to catch regressions before you publish

Write a test once, run it against any agent, and get a **Pass** or **Fail** verdict with the judge's reasoning. Tests live in a shared workspace library, run against your agent's current draft, and never touch what's published, so you can iterate on a prompt and re-run in seconds.

## Test types

| Type | What it checks | Use it for |
| - | - | - |
| **Single Turn** | The next reply to a scripted conversation, scored by an LLM judge against your expectation. | Wording, tone, and one specific answer. |
| **Tool** | Whether the next reply calls a given tool (or no tool at all), optionally with matching parameters. | Routing decisions and argument extraction. |
| **Simulation** | A whole conversation driven by a simulated user, scored by the judge against your success conditions, plus deterministic checks on what happened. | Flows that only show up over several exchanges, such as taking an order. |

<Note>
  Tests run as text (no audio is synthesized), but the agent uses its full draft
  configuration: the [knowledge base](/agents/build/knowledge-base) is
  consulted, and the agent can invoke its attached tools. In Single Turn and
  Tool tests, [webhook tools](/agents/build/webhook-tools) send real HTTP
  requests, so point them at a staging endpoint. Simulation tests mock tools by
  default. For end-to-end verification with voice, use [preview
  calls](/agents/test/preview-calls).
</Note>

## Create a test

Tests are workspace-level resources, managed under **Library → Tests** in the console and shareable across every agent in your workspace.

<Steps>
  <Step title="Start a new test">
    Open **Library → Tests** and click **Add test**. Give it a name and pick a
    **Type**.
  </Step>

  <Step title="Describe what to check">
    Fill in the fields for that type. They are described in the sections below:
    [Single Turn](#single-turn-tests), [Tool](#tool-tests), and
    [Simulation](#simulation-tests).
  </Step>

  <Step title="Save">
    Click **Create Test**. The test is now in your library, ready to attach to
    agents.
  </Step>
</Steps>

## Single Turn tests

A Single Turn test hands your agent a scripted conversation and asks it to produce the next reply:

1. You script a conversation history of agent and user messages.
2. The agent generates the next reply using its current **draft** configuration and the same language model that answers in live conversations.
3. An LLM judge scores the reply against your **Expectation**, optionally calibrated by success and failure examples, and returns a verdict with its reasoning.

Under **Conversation**, click **Add message** to build the history the agent sees. Each message is either an **Agent** or **User** turn, and at least one must be a user turn. Under **Judging**, write the **Expectation**: what a correct reply must do. Optionally click **Add example** to provide success and failure examples. They calibrate the judge but aren't required.

## Tool tests

A Tool test scripts the conversation the same way, but instead of judging the reply it checks which tool the agent called while producing it.

Under **Require tool execution**, pick a tool from the library and choose **Should have been called** or **Should not be called**. Leave the tool empty to check that the agent called no tool at all. Integration tools, such as calendar tools, can only be checked in a Simulation test. Under **Tool parameters**, optionally add the parameter values the call must carry, typed as string, number, or boolean. The result names the tool the agent actually called.

## Simulation tests

A Simulation test does not script the user. Instead, a simulated user plays a role you describe, talks to your agent for up to a set number of turns, and an LLM judge scores the finished conversation against your success conditions.

### Write the scenario

The scenario is the simulated user's brief. Click **Insert template** to start from the four sections the simulator expects:

| Section | What to write |
| - | - |
| **PERSONA** | Who the user is and how they talk, for example an impatient customer on a lunch break. |
| **GOAL** | What they want out of the conversation. |
| **FACTS** | Details the agent may ask for, such as an order number or a date. Reveal one fact at a time. |
| **ENDING** | When the user hangs up: satisfied, out of patience, or after a fixed number of tries. |

Two rules make scenarios reliable. Reveal one fact at a time, so the agent has to ask for what it needs instead of receiving everything in the first message. And give an explicit ending, otherwise the conversation runs until the turn limit and the judge has to guess whether the user was done.

**Max turns** caps how many user turns the simulation runs. A **Conversation** script is optional here. Any messages you add are replayed first, then the simulated user takes over from the last message.

### Channel

**Channel** sets the register the agent speaks in during the simulation, the same one it uses on that channel in production: **Web voice**, **Phone inbound**, or **Phone outbound**. On a phone channel, for example, the agent repeats important numbers back in small groups. Phone outbound speaks in the same phone register as Phone inbound and only `{{system.channel}}` differs, so if the agent should open an outbound call differently, branch on `{{system.channel}}` in its prompt. **Auto**, the default, uses Phone inbound when a phone number is assigned to the agent and Web voice otherwise, so pick Phone outbound yourself for an agent that only places calls. As in production, client tools are not available on phone channels, and transfers only happen on phone channels. `{{system.channel}}` in the agent's prompt and in the scenario takes the same value. No real call is placed, so `{{system.caller_number}}` and `{{system.dialed_number}}` are empty on every channel.

### Success conditions

Add one to ten **Success conditions**, each a plain sentence describing something that must happen in the conversation, for example `The agent confirms the refund amount before closing`. Give each a short name, or leave it blank and the first words of the description become the name.

The judge reads the whole transcript and marks every condition **Success**, **Failure**, or **Unknown**. A run passes only when every condition succeeds. An **Unknown** verdict means the transcript did not contain enough evidence either way, and the run is flagged **Needs review** so you read the transcript before trusting the result. Success and failure examples, when you add them, calibrate the judge across all conditions.

### Tool mocks

Simulated conversations usually run many times, so by default tool calls never reach your real endpoints:

| Strategy | Behaviour |
| - | - |
| **Mock all** | Every tool answers from a mock. This is the default. A tool without any mock fails, see below. |
| **Mock selected** | Only the tools you list answer from a mock. Choose whether the rest return an error or call their real endpoint. |
| **Mock none** | Every tool with a real endpoint calls it. Webhooks may create real side effects, so reserve this for a safe test environment. |

A mock entry is the tool, the result it returns (JSON or plain text), an optional HTTP status for webhook tools, and optional parameter conditions so the mock only applies when the agent calls the tool with matching arguments (exact match, regular expression, or any value). Turn on **Return an error** to make the tool fail with that result instead, for example to test how the agent handles an outage. A webhook tool set to hide errors shows the agent only that the call failed, as in production. When a tool has several entries, the ones with conditions are checked first, in order, and an entry without conditions is the fallback that answers only when none of them match, wherever it sits in the list. Regular expressions, in mock conditions and in required tool call parameters alike, use JavaScript syntax without Unicode mode. Python forms such as `(?i)` or `(?P<name>...)` are refused when you save. Write a named group as `(?<name>...)`, and ignore case with `(?i:...)` around the part it applies to. Forms that JavaScript reads as literal text without Unicode mode are refused too: Unicode escapes such as `\p{L}` or `\u{1F600}` (list the characters in a class instead, for example `[A-Za-zÀ-ÿ]`) and Python escapes such as `\A`, `\z` or `\N{...}` (write `\A` as `^` and `\z` as `$`). The platform checks required tool call parameters itself, so there back-references, emoji and other characters outside the Basic Multilingual Plane, non-ASCII characters under `(?i:...)`, `\S` inside a class without `\s` (`[\s\S]` is fine) and `\c` are refused as well.

Under **Mock all**, a tool without an entry in the test answers with its own first mock response. A tool with neither fails, and the agent gets the same error it would get if the tool were unavailable in production. Tools whose result the agent never waits for, client tools that don't wait for an answer and webhook tools in fire and forget mode, need no mock: as in production, the agent only ever hears their acknowledgement, whether the call ran for real, answered from a mock, or was skipped. Such a client tool can't be mocked at all. A simulation has no client, so a client tool that waits for an answer always answers from its mock entries, under every strategy including **Mock none**, and needs one. The test form lists the agent's tools that would fail this way. To let specific tools reach their real endpoint while everything else stays mocked, add them under **Call the real endpoint**. Only webhook tools and read-only integration tools can go there. Every other tool still needs a mock, so a tool you attach to the agent later fails instead of reaching its endpoint until you mock it. A tool there can still have mock entries: a matching entry answers, and any other call reaches the endpoint. Integration tools, such as calendar tools, are mocked like any other tool. An integration that writes only after the caller confirms, such as creating a calendar event, answers the first call with a confirmation request and commits through its **Confirm write** tool. Mock both, the way production answers them. If you mock only the write tool and the model then calls Confirm write, that call goes unanswered and the run needs review.

A run in which a tool call went unanswered never passes: the tool had no mock, or none of its entries matched and neither the tool's own mock response nor **Call the real endpoint** could answer. The run fails with **Needs review** and **No mock** badges and lists the unanswered calls, even when the judge and every assertion passed, because the conversation left the path it would take in production. A call with no matching entry means either the agent passed arguments you didn't expect or the entry's conditions are too narrow. Calls to a tool the test forbids are not listed, they are the agent's failure and the assertion reports them. The judge is told the call failed because the test prepared no answer for it, and judges the agent as usual without treating the call as having run. Add or widen the mock and run the test again.

If you are used to tools without a mock calling their real endpoint, your first runs will show **No mock** on those calls. Add a mock for each tool the test form lists, or add read-only tools to **Call the real endpoint**.

Mocks and assertions apply to the tools attached to the agent under test. A tool the agent does not have is never mocked and counts as never called: a required call on it fails with *This tool is not on the agent*, while a forbidden tool or a maximum-only check passes. Client tools on a phone channel are treated the same way. This keeps a shared guard such as "never call issue\_refund" passing on agents that cannot call the tool.

### Assertions

Under **Advanced**, add deterministic checks that run alongside the judge:

* **Required tool calls**: the agent must call the tool, optionally with parameters that match, between a minimum and a maximum number of matching calls. The default is at least once. Set the maximum to 1 to catch a duplicate booking, or set both the minimum and the maximum to 0 to forbid calls with those arguments.
* **Forbidden tools**: the agent must not call the tool.
* **Ended by**: who ended the conversation, the agent, the user, a transfer, or any of them. A transfer only happens on a phone channel, so on Web voice a transfer check always fails.

A failed assertion fails the run regardless of the judge's verdict. Each assertion shows **Passed** or **Failed** with a short detail in the result.

### Repeat count

Set **Repeat** to run the same scenario up to 20 times in a batch. Simulated conversations vary from run to run, so repeats surface flaky behaviour that a single run would miss. The agent's Tests page shows how many repeats passed out of those that reached a verdict, for example `2 / 3 passed`, or `2 / 2 passed · 1 error` when one repeat ended in **Error**. While repeats are still running, the row counts them instead, for example `1 passed · 2 running`. Opening the test from that page lets you switch between the finished repeats of the latest batch to read each conversation. The test opens on the first failed repeat.

Repeats only apply to **Run all**. Running the test from its **Run** tab, or re-running one test from the row menu, runs the conversation once.

## Attach tests to agents

A test only runs against agents it's attached to. Attach from either side:

* **From the library**: open the test's **Access** tab and toggle it on for each agent.
* **From the Builder**: on the agent's **Tests** page, click **Add tests** and pick from the library.

One test can be attached to many agents, and each agent keeps its own last result. **Remove from agent** detaches the test from that agent only. **Delete** in the library removes the test from all agents.

## Run a single test

Open the test and switch to its **Run** tab. Pick an agent that has access, then click **Run test**. Single Turn and Tool tests answer within seconds. A Simulation test runs in the background, shows **Queued** and then **Running**, and can take a few minutes.

The verdict card shows **Pass** or **Fail** with the time the run took, followed by details for the test type:

| Type | What you see |
| - | - |
| **Single Turn** | The full reply the agent generated and why the judge passed or failed it. |
| **Tool** | Which tool the agent called, if any, against what the test expected. |
| **Simulation** | The judge's summary, every success condition with its verdict and rationale, every assertion with its outcome, who ended the conversation, after how many turns and on which channel, the LLM cost of the agent, the simulated user, and the judge with their token breakdown, and the full transcript with each tool call marked by where its answer came from: **Mock** for an entry in the test, **Tool mock** for the tool's own mock response, **Real** when it reached the real endpoint, **No mock** or **No matching mock** when it failed for lack of one. |

A Simulation run with an **Unknown** condition fails and carries a **Needs review** badge. If a run can't complete, the card shows **Error** with a message that says whether the agent caused it, for example its custom LLM endpoint stopped responding, or the platform did, for example our language model did not answer. A conversation that broke off or hit its time limit is never judged, so it can't pass by accident. The transcript is still shown below the error, as it is when the conversation finished but could not be judged.

## Run every test for an agent

On the agent's **Tests** page in the Builder, click **Run all**. Each row moves through **queued → running → Pass/Fail** (or **Error** if a run can't complete). Simulation tests run once per repeat and the row shows how many repeats passed. The page header summarizes the latest batch, for example `4 passed, 1 failed on last run`, with the batch **pass rate**: green at 100%, amber from 80%, red below. The pass rate counts passed runs out of passed and failed ones, so a run that ended in **Error** does not lower it. Use the row menu to re-run a single test, edit it, or remove it from the agent.

## Tests run against the draft

Tests always exercise the agent's latest **draft** configuration, including unpublished changes to the system prompt. That makes the loop fast:

<Steps>
  <Step title="Edit the draft">
    Change the system prompt or first message in
    [Configuration](/agents/build/configuration).
  </Step>

  <Step title="Re-run">Click **Run all**.</Step>

  <Step title="Publish when green">
    Once results look right, [publish the
    draft](/agents/deploy/versions-publishing). Running tests never publishes
    anything.
  </Step>
</Steps>

## Limits

| Field | Limit |
| - | - |
| Test name | 200 characters |
| Conversation message | 2,000 characters each, at least one message (optional for Simulation) |
| Conversation (Simulation) | 50 messages |
| Expectation | 400 characters |
| Success / failure example | 400 characters each |
| Scenario (Simulation) | 10,000 characters |
| Success conditions (Simulation) | 1 to 10, description 500 characters each |
| Max turns (Simulation) | 50 |
| Repeat count (Simulation) | 20 |
| Dynamic variables | 50 per test |

## Going further

<CardGroup cols={2}>
  <Card title="Preview calls" icon="phone" href="/agents/test/preview-calls">
    Talk to your draft agent live, with voice, tools, and transcripts.
  </Card>

  <Card title="Versions & publishing" icon="rocket" href="/agents/deploy/versions-publishing">
    How drafts become immutable published versions.
  </Card>

  <Card title="Configuration" icon="sliders" href="/agents/build/configuration">
    System prompt, first message, and everything tests exercise.
  </Card>

  <Card title="Core concepts" icon="lightbulb" href="/agents/concepts">
    Agents, drafts, sessions, and the workspace library model.
  </Card>
</CardGroup>
