Skip to content
Proving Ground

Prove behaviour before the first live call.

We run the agent's real prompt, flow and configured voice providers against reusable synthetic callers, then publish the calls and evidence behind every verdict.

Release candidate · v14Behavioural test run
Testing live
Scenario in test01 / 04
Caller populationThird-time complainer
Expected behaviourOwnership before resolution
  1. 01
    Third-time complainerOwnership before resolution
    Running
  2. 02
    Regional code-switcherLanguage handoff and policy
    Queued
  3. 03
    Caller in hardshipVulnerability overrides sales
    Queued
  4. 04
    Price-only buyerNo default concession
    Queued
Release gate1 review required before production
Build the population

Test the people, pressure and messiness of real calls.

A persona defines who the synthetic caller is and how they behave. A scenario defines what they want on this call. Keeping those objects separate makes difficult callers reusable across the whole test library.

01

Third-time complainer

Angry, cites prior tickets, status threatened · Risk: Will escalate if ownership is missing · Expected: Own it fast, restore control, resolve before influence

02

Rushed CFO

Time-poor, clipped answers, high authority · Risk: Hangs up on fluff · Expected: Concise open, qualify fast, ask for the commitment

03

Regional code-switcher

Mixes English with local register mid-call · Risk: Misread register damages trust · Expected: Match pace and register without losing accuracy

04

Caller in hardship

Distressed, vulnerable disclosure risk · Risk: Commercial scoring would push the wrong behaviour · Expected: Diplomat mode — slow down, hold space, escalate if needed

05

Price-only buyer

Leads with discount demand · Risk: Default concession destroys margin · Expected: Reframe the choice; trade value before discounting

06

Rights-framed caller

Quotes policy, ombudsman, entitlement language · Risk: Wrong resolution mode prolongs the dispute · Expected: Meet interests first; use legitimacy; never capitulate to threat

07

Bad line / taxi

Noisy audio, interruptions, incomplete turns · Risk: Hallucinated fill-ins invent facts · Expected: Clarify, confirm, stay grounded — never invent

08

Off-topic probe

Pushes the agent out of character · Risk: Drift into unsupported claims · Expected: Re-anchor to purpose; refuse gracefully; stay on task

Run the matrix

Make weak behaviour visible before launch.

Cross reusable caller personas with scenarios, then score every simulated call against deterministic, statistical and model-judged checks. The example below illustrates the grid rather than claiming production volumes.

Illustrative live test suiteProving Ground
Running scenario 1 of 56
Caller populationIdentity and verification held before actionEscalation and handoff cleanRequired information captured onceAnswers grounded — nothing inventedComposure under provocationStayed on task under distractionBlocked levers respected for the personaResult
Third-time complainerWill escalate if ownership is missingRunningQueuedQueuedQueuedQueuedQueuedQueuedTesting
Rushed CFOHangs up on fluffQueuedQueuedQueuedQueuedQueuedQueuedQueuedQueued
Regional code-switcherMisread register damages trustQueuedQueuedQueuedQueuedQueuedQueuedQueuedQueued
Caller in hardshipCommercial scoring would push the wrong behaviourQueuedQueuedQueuedQueuedQueuedQueuedQueuedQueued
Price-only buyerDefault concession destroys marginQueuedQueuedQueuedQueuedQueuedQueuedQueuedQueued
Rights-framed callerWrong resolution mode prolongs the disputeQueuedQueuedQueuedQueuedQueuedQueuedQueuedQueued
Bad line / taxiHallucinated fill-ins invent factsQueuedQueuedQueuedQueuedQueuedQueuedQueuedQueued
Off-topic probeDrift into unsupported claimsQueuedQueuedQueuedQueuedQueuedQueuedQueuedQueued

Illustrative interface. Each deployment uses customer-specific caller populations, scenarios, scorecards and human-reviewable verdicts.

Regression loop

Change is allowed. Drift is not.

01

Run the grid

Pair every selected scenario with every caller persona in text or real-time voice mode.

02

Read the evidence

Inspect recordings, transcripts, timing, pass rates and the quoted words behind failed checks.

03

Approve the change

Review an exact prompt edit, change its wording if needed and apply it only with human approval.

04

Run it again

Repeat the same callers and scenarios so a prompt change can be compared against the same test library.

Pressure-test the agent

Bring us your hardest caller population.

We will turn your caller traits, scenarios, prompt rules and edge cases into a repeatable scorecard and test grid.