Third-time complainer
Angry, cites prior tickets, status threatened · Risk: Will escalate if ownership is missing · Expected: Own it fast, restore control, resolve before influence
We run the agent's real prompt, flow and configured voice providers against reusable synthetic callers, then publish the calls and evidence behind every verdict.
A persona defines who the synthetic caller is and how they behave. A scenario defines what they want on this call. Keeping those objects separate makes difficult callers reusable across the whole test library.
Angry, cites prior tickets, status threatened · Risk: Will escalate if ownership is missing · Expected: Own it fast, restore control, resolve before influence
Time-poor, clipped answers, high authority · Risk: Hangs up on fluff · Expected: Concise open, qualify fast, ask for the commitment
Mixes English with local register mid-call · Risk: Misread register damages trust · Expected: Match pace and register without losing accuracy
Distressed, vulnerable disclosure risk · Risk: Commercial scoring would push the wrong behaviour · Expected: Diplomat mode — slow down, hold space, escalate if needed
Leads with discount demand · Risk: Default concession destroys margin · Expected: Reframe the choice; trade value before discounting
Quotes policy, ombudsman, entitlement language · Risk: Wrong resolution mode prolongs the dispute · Expected: Meet interests first; use legitimacy; never capitulate to threat
Noisy audio, interruptions, incomplete turns · Risk: Hallucinated fill-ins invent facts · Expected: Clarify, confirm, stay grounded — never invent
Pushes the agent out of character · Risk: Drift into unsupported claims · Expected: Re-anchor to purpose; refuse gracefully; stay on task
Cross reusable caller personas with scenarios, then score every simulated call against deterministic, statistical and model-judged checks. The example below illustrates the grid rather than claiming production volumes.
| Caller population | Identity and verification held before action | Escalation and handoff clean | Required information captured once | Answers grounded — nothing invented | Composure under provocation | Stayed on task under distraction | Blocked levers respected for the persona | Result |
|---|---|---|---|---|---|---|---|---|
| Third-time complainerWill escalate if ownership is missing | Running | Queued | Queued | Queued | Queued | Queued | Queued | Testing |
| Rushed CFOHangs up on fluff | Queued | Queued | Queued | Queued | Queued | Queued | Queued | Queued |
| Regional code-switcherMisread register damages trust | Queued | Queued | Queued | Queued | Queued | Queued | Queued | Queued |
| Caller in hardshipCommercial scoring would push the wrong behaviour | Queued | Queued | Queued | Queued | Queued | Queued | Queued | Queued |
| Price-only buyerDefault concession destroys margin | Queued | Queued | Queued | Queued | Queued | Queued | Queued | Queued |
| Rights-framed callerWrong resolution mode prolongs the dispute | Queued | Queued | Queued | Queued | Queued | Queued | Queued | Queued |
| Bad line / taxiHallucinated fill-ins invent facts | Queued | Queued | Queued | Queued | Queued | Queued | Queued | Queued |
| Off-topic probeDrift into unsupported claims | Queued | Queued | Queued | Queued | Queued | Queued | Queued | Queued |
Illustrative interface. Each deployment uses customer-specific caller populations, scenarios, scorecards and human-reviewable verdicts.
Pair every selected scenario with every caller persona in text or real-time voice mode.
Inspect recordings, transcripts, timing, pass rates and the quoted words behind failed checks.
Review an exact prompt edit, change its wording if needed and apply it only with human approval.
Repeat the same callers and scenarios so a prompt change can be compared against the same test library.
We will turn your caller traits, scenarios, prompt rules and edge cases into a repeatable scorecard and test grid.