

9 min read time
Topics

Aryan Vijay Bhosale

Harshit Rajgarhia

Akhil Pothanapalli

Asif Shaik

Abhishek Mukherji

Dinesh Manocha
Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text.
Full-duplex voice agent research has proliferated across academia and industry. Benchmarks like τ-Voice showed that voice agents retain only 30–45% of text-agent capability on identical grounded tasks under realistic audio, and EVA-Bench showed that no system is simultaneously good at task accuracy and conversational experience. The Full-Duplex-Bench series progressed from static turn-taking probes through overlap handling and multi-turn evaluation to real disfluent speech with chained tool calls, while parallel lines isolate interruption and post-interruption recovery, and a recent survey maps the architectural landscape.
While the difficulty and framing of these tasks and evaluations were justified at the time, the voice agents of today are far more capable and deserve benchmarks that keep up with their rapid development. While prior benchmarks pursued coverage across domains in their own ways, they failed to question the base premise on which voice agents are built: their ability to integrate seamlessly into daily life. Existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation.
Introducing DuplexWorld
To tackle these issues, we introduce DuplexWorld comprising six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.
6 Conversational Worlds
World | What decides correctness | The hard part | Best Pass@1 |
|---|---|---|---|
Banking | financial-crime statute | Translating a symptom into an operation when several plausible, sympathetically framed requests are reporting-threshold tripwires. | 0.326 Voice Think Fast |
Logistics | contract topology over physical goods | Deciding who is entitled to redirect goods the agent can move in the record but cannot see. | 0.533 Realtime-2.1 |
Healthcare | lawful disclosure | Refusing without leaking in the act of refusing. A fully verified caller may be entitled to nothing. | 0.519 Voice Think Fast |
Insurance | the platform’s own limited remit | Taking first notice of a loss the agent is structurally forbidden from deciding. | 0.556 Voice Think Fast |
Travel | the platform’s own limited remit | Doing nothing, correctly. The modal right answer is a factual correction or a refusal held with nowhere to escalate to. | 0.674 Realtime-2.1 |
Pathfinding | perceptual grounding | Grounding an instruction against a heading neither party can state. | 0.533 Voice Think Fast |
The six worlds. Best single-world Pass@1 across the five systems tested, realistic channel.
11 conversational types
Every scenario instantiates one of eleven conversation types — a domain-independent shape, filled in by the world’s premise, records and risk tiers. The caller never names the operation; they describe a symptom, and inferring what is being asked for is part of the task. We believe that the introduction of newer conversational worlds should be accompanied by a wider gamut of tasks that can help voice agents keep up with users.

Three of the nine enterprise shapes worked through, with the two-line sketches the benchmark itself carries. The other six are named.
Extending conversations beyond enterprise customer service
A true test of analytical ability of agents has always been maze-solving or navigation. In our Pathfing world, we want to see if a voice copilot can guide a pedestrian on foot, one street corner at a time, to a named entrance on a 8 × 8 grid of streets inspired by the New York City map. We systematically increase difficulty by introducing roadblocks to test re-routing and multi-destination voice agent assistance through our all-day assistance task.
![]() | ![]() | ![]() |
Single intent | Rerouting | All-day assistance |
Single intent is a base route of five blocks and two turns, nothing closed. Rerouting seals every route the copilot can see with four closures absent from its map, leaving exactly one seven-block survivor. All-day assistance runs two legs of five and six blocks, the second destination revealed only on arrival at the first. Single intent is shared with the enterprise worlds; the other two arise only here.
To understand more about our harness, experimental setup and to experience DuplexWorld more vividly, head over to our project page.
Metrics
Agentic capability GS · Pass@1 · Pass³ · ϱ⁺ · π⁻ | Conversational dynamics TT · CP · SEL | Naturalness FAI · DNSMOS · UTMOS · NISQA |
|---|---|---|
A canonical hash of the record store against a gold replay; in Pathfinding, the walker’s final junction and side of street. The reward multiplies exactly the binary factors each scenario names: goal state, a gold-action match, and a judged assertion where what is said or refused is the point. Pathfinding has no gold action list, so an efficiency conjunct η ≥ 0.75 replaces those. Effort is measured against each task’s own reference workload. | Every floor transfer’s offset on an on-time curve chosen by the transfer’s kind, with a single unanswered user turn zeroing the conversation; one topic-free judge pass over the transcript; and the fraction of injected distractors correctly ignored, against gold labels written at injection time. | One judge pass over the transcript with the agent’s instructions, role and tool schemas, five binary dimensions scored as their minimum; plus three no-reference mean-opinion-score predictors run over the isolated agent channel. |
The harness elicits more than the suite scores: backchannel placement, the prosody of overlap and whether a mid-utterance revision is fluent are all elicited and none is scored. A system could be graceless at every overlap and lose nothing on the agentic pillar. That is why the pillars are never composed into one number, and why what follows is a table rather than a ranking.
System | GS | Pass@1 | Pass@3 | Pass³ | TT | CP | SEL | FAI | DNSMOS | UTMOS | NISQA |
|---|---|---|---|---|---|---|---|---|---|---|---|
Voice Think Fast | 0.779 ±.031 | 0.490 ±.039 | 0.694 ±.068 | 0.266 ±.057 | 0.635 ±.015 | 1.814 ±.051 | 0.630 ±.024 | 1.880 ±.052 | 3.127 ±.007 | 3.687 ±.010 | 3.093 ±.009 |
Realtime-2.1 | 0.619 ±.037 | 0.433 ±.035 | 0.659 ±.068 | 0.212 ±.048 | 0.653 ±.018 | 1.543 ±.048 | 0.414 ±.024 | 2.025 ±.057 | 3.350 ±.004 | 4.095 ±.010 | 3.600 ±.010 |
3.1-Flash-Live | 0.726 ±.029 | 0.398 ±.039 | 0.639 ±.073 | 0.165 ±.054 | 0.388 ±.012 | 1.511 ±.043 | 0.742 ±.024 | 1.720 ±.051 | 3.378 ±.008 | 3.402 ±.010 | 3.477 ±.012 |
Realtime-2.1-mini | 0.405 ±.038 | 0.188 ±.028 | 0.358 ±.065 | 0.056 ±.028 | 0.521 ±.025 | 1.307 ±.041 | 0.468 ±.023 | 1.631 ±.054 | 3.334 ±.009 | 4.022 ±.033 | 3.611 ±.022 |
Nova 2 Sonic | 0.263 ±.028 | 0.011 ±.007 | 0.019 ±.019 | 0.006 ±.009 | 0.566 ±.021 | 1.019 ±.012 | 0.977 ±.006 | 1.635 ±.054 | 3.172 ±.015 | 2.556 ±.017 | 2.562 ±.017 |
The suite, pooled over the six worlds, realistic channel. n = 135 conversations per enterprise cell and 45 per Pathfinding cell. Subscripts are 95% bootstrap half-widths. CP and FAI are on 1 to 3 and the MOS predictors on 1 to 5; everything else is on 0 to 1. Best per column in magenta. The effort pair is reported separately.
State versus reward
Read the first column against the second. Across the 25 enterprise cells the mean gap between the goal-state check and the reward is +0.255, and every cell is positive. A do-nothing agent passes goal state wherever the gold terminal state equals the seeded state, so a benchmark that headlines goal state is partly reporting how many of its tasks are no-ops. The reward is lost in the remaining conjuncts: acting through the sanctioned sequence, and saying what was done.
Performance across different types of metrics

Each system’s rank on one metric from each pillar. The crossings are the finding.
GPT Realtime-2.1 holds the best turn-taking at 0.653 and Grok Voice Think Fast the best conversation progression at 1.814 on a 1-to-3 scale, while Gemini-3.1-Flash-Live is third on the reward and last on turn-taking at 0.388. We can see how exceptional agentic performance does not necessarily imply that agents can maintain conversational dynamics and naturalness. We hope that this encourages the development of more capable speech-to-speech systems in the future.
Performance of voice agents across worlds

Pass@1 for each system in each of the six worlds. The bar is the interval the system spans; each marker is one world, keyed above. Markers are nudged vertically where two worlds share a score.
We see that while the relative ranking of voice agents remains roughly unchanged, the subject matter of a world alone swings an agent’s performance by up to 0.474.
Reliability Analysis

Reliability over k attempts, pooled over the six worlds. Dashed is Pass@k, at least one pass in k. Solid is Passᵏ, all k of k.
We can see how the best voice agents score <50% (Grok Voice Think Fast) on pass@1 and <20% on pass^5 agentic capability on DuplexWorld. This means that even today, voice agents lack the ability to be effective and robust on-call assistants.
Conclusion
If these voice agents must integrate reliably into enterprise and consumer AI workflows, we would need the development of most effective and reliable speech-to-speech agents. We’re glad to see the advent of speech-to-speech models from mere conversational systems to agents that are capable of handling complex tool calling and hope that we can see the emergence of much more powerful models soon.
What DuplexWorld changes about evaluating voice agents
6 worlds, 11 conversation types | 156 scenarios authored, 144 scored | 3,825 conversations, 387 simulated hours |
0.490 best pooled Pass@1, of 1.000 | 0.266 best Pass³, all three of three | 17.1% of spoken credentials misheard |
The larger lesson from DuplexWorld might that identifying “best voice agent” may be the wrong objective. A system that excels at completing tasks may be weaker at managing a conversation, while one that sounds natural may struggle when the situation changes or the same task has to be completed reliably more than once. That makes voice-agent selection a question of fit: a bank, a healthcare provider, and a navigation service may need very different combinations of accuracy, conversational skill, and consistency. Better evaluation should make those tradeoffs visible so that organizations can choose agents based on what the job actually demands.
Learn more about Centific’s applied AI research here.
Are your ready to get
modular
AI solutions delivered?
Connect data, models, and people — in one enterprise-ready platform.
Latest Insights
Connect with Centific
Updates from the frontier of AI data.
Receive updates on platform improvements, new workflows, evaluation capabilities, data quality enhancements, and best practices for enterprise AI teams.




