Abstract image

Research insight

Research insight

Can voice agents help you get through the day?

Can voice agents help you get through the day?

Explore how DuplexWorld evaluates speech-to-speech AI agents across task completion, conversational dynamics, and naturalness, revealing why strong performance in one area does not guarantee strength in another.

Explore how DuplexWorld evaluates speech-to-speech AI agents across task completion, conversational dynamics, and naturalness, revealing why strong performance in one area does not guarantee strength in another.

9 min read time

Table of contents

Share

Summarize

Topics

Voice AI
AI Evaluation
AI Agents
AI Benchmarks
Voice AI
AI Evaluation
AI Agents
AI Benchmarks

Author(s)

Author(s)

Centific logo

Aryan Vijay Bhosale

Harshit Rajgarhia

Harshit Rajgarhia

Centifc logo

Akhil Pothanapalli

Centifc logo

Asif Shaik

Abhishek Mukherji

Abhishek Mukherji

University of Maryland logo

Dinesh Manocha

Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text.

Full-duplex voice agent research has proliferated across academia and industry. Benchmarks like τ-Voice showed that voice agents retain only 30–45% of text-agent capability on identical grounded tasks under realistic audio, and EVA-Bench showed that no system is simultaneously good at task accuracy and conversational experience. The Full-Duplex-Bench series progressed from static turn-taking probes through overlap handling and multi-turn evaluation to real disfluent speech with chained tool calls, while parallel lines isolate interruption and post-interruption recovery, and a recent survey maps the architectural landscape.

While the difficulty and framing of these tasks and evaluations were justified at the time, the voice agents of today are far more capable and deserve benchmarks that keep up with their rapid development. While prior benchmarks pursued coverage across domains in their own ways, they failed to question the base premise on which voice agents are built: their ability to integrate seamlessly into daily life. Existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation.

Introducing DuplexWorld

To tackle these issues, we introduce DuplexWorld comprising six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.

6 Conversational Worlds


World

What decides correctness

The hard part

Best Pass@1

Banking

financial-crime statute

Translating a symptom into an operation when several plausible, sympathetically framed requests are reporting-threshold tripwires.

0.326  Voice Think Fast

Logistics

contract topology over physical goods

Deciding who is entitled to redirect goods the agent can move in the record but cannot see.

0.533  Realtime-2.1 

Healthcare

lawful disclosure

Refusing without leaking in the act of refusing. A fully verified caller may be entitled to nothing.

0.519  Voice Think Fast

Insurance

the platform’s own limited remit

Taking first notice of a loss the agent is structurally forbidden from deciding.

0.556  Voice Think Fast

Travel

the platform’s own limited remit

Doing nothing, correctly. The modal right answer is a factual correction or a refusal held with nowhere to escalate to.

0.674  Realtime-2.1

Pathfinding

perceptual grounding

Grounding an instruction against a heading neither party can state.

0.533  Voice Think Fast

The six worlds. Best single-world Pass@1 across the five systems tested, realistic channel.

11 conversational types

Every scenario instantiates one of eleven conversation types — a domain-independent shape, filled in by the world’s premise, records and risk tiers. The caller never names the operation; they describe a symptom, and inferring what is being asked for is part of the task. We believe that the introduction of newer conversational worlds should be accompanied by a wider gamut of tasks that can help voice agents keep up with users.


Three of the nine enterprise shapes

Three of the nine enterprise shapes worked through, with the two-line sketches the benchmark itself carries. The other six are named.

Extending conversations beyond enterprise customer service

A true test of analytical ability of agents has always been maze-solving or navigation. In our Pathfing world, we want to see if a voice copilot can guide a pedestrian on foot, one street corner at a time, to a named entrance on a 8 × 8 grid of streets inspired by the New York City map. We systematically increase difficulty by introducing roadblocks to test re-routing and multi-destination voice agent assistance through our all-day assistance task.


Single intentReroutingAll-day assistance

Single intent

Rerouting

All-day assistance

Single intent is a base route of five blocks and two turns, nothing closed. Rerouting seals every route the copilot can see with four closures absent from its map, leaving exactly one seven-block survivor. All-day assistance runs two legs of five and six blocks, the second destination revealed only on arrival at the first. Single intent is shared with the enterprise worlds; the other two arise only here.

To understand more about our harness, experimental setup and to experience DuplexWorld more vividly, head over to our project page

Metrics


Agentic capability 

GS · Pass@1 · Pass³ · ϱ⁺ · π⁻ 

Conversational dynamics 

TT · CP · SEL 

Naturalness 

FAI · DNSMOS · UTMOS · NISQA 

A canonical hash of the record store against a gold replay; in Pathfinding, the walker’s final junction and side of street. The reward multiplies exactly the binary factors each scenario names: goal state, a gold-action match, and a judged assertion where what is said or refused is the point. Pathfinding has no gold action list, so an efficiency conjunct η ≥ 0.75 replaces those. Effort is measured against each task’s own reference workload.

Every floor transfer’s offset on an on-time curve chosen by the transfer’s kind, with a single unanswered user turn zeroing the conversation; one topic-free judge pass over the transcript; and the fraction of injected distractors correctly ignored, against gold labels written at injection time.

One judge pass over the transcript with the agent’s instructions, role and tool schemas, five binary dimensions scored as their minimum; plus three no-reference mean-opinion-score predictors run over the isolated agent channel.

The harness elicits more than the suite scores: backchannel placement, the prosody of overlap and whether a mid-utterance revision is fluent are all elicited and none is scored. A system could be graceless at every overlap and lose nothing on the agentic pillar. That is why the pillars are never composed into one number, and why what follows is a table rather than a ranking.


System 

GS 

Pass@1 

Pass@3 

Pass³ 

TT 

CP 

SEL 

FAI 

DNSMOS 

UTMOS 

NISQA 

Voice Think Fast

0.779

±.031

0.490

±.039

0.694

±.068

0.266

±.057

0.635

±.015

1.814

±.051

0.630

±.024

1.880

±.052

3.127

±.007

3.687

±.010

3.093

±.009

Realtime-2.1

0.619

±.037

0.433

±.035

0.659

±.068

0.212

±.048

0.653

±.018

1.543

±.048

0.414

±.024

2.025

±.057

3.350

±.004

4.095

±.010

3.600

±.010

3.1-Flash-Live

0.726

±.029

0.398

±.039

0.639

±.073

0.165

±.054

0.388

±.012

1.511

±.043

0.742

±.024

1.720

±.051

3.378

±.008

3.402

±.010

3.477

±.012

Realtime-2.1-mini

0.405

±.038

0.188

±.028

0.358

±.065

0.056

±.028

0.521

±.025

1.307

±.041

0.468

±.023

1.631

±.054

3.334

±.009

4.022

±.033

3.611

±.022

Nova 2 Sonic

0.263

±.028

0.011

±.007

0.019

±.019

0.006

±.009

0.566

±.021

1.019

±.012

0.977

±.006

1.635

±.054

3.172

±.015

2.556

±.017

2.562

±.017

The suite, pooled over the six worlds, realistic channel. n = 135 conversations per enterprise cell and 45 per Pathfinding cell. Subscripts are 95% bootstrap half-widths. CP and FAI are on 1 to 3 and the MOS predictors on 1 to 5; everything else is on 0 to 1. Best per column in magenta. The effort pair is reported separately.

State versus reward

Read the first column against the second. Across the 25 enterprise cells the mean gap between the goal-state check and the reward is +0.255, and every cell is positive. A do-nothing agent passes goal state wherever the gold terminal state equals the seeded state, so a benchmark that headlines goal state is partly reporting how many of its tasks are no-ops. The reward is lost in the remaining conjuncts: acting through the sanctioned sequence, and saying what was done.

Performance across different types of metrics

Performance across different types of metrics

Each system’s rank on one metric from each pillar. The crossings are the finding.

GPT Realtime-2.1 holds the best turn-taking at 0.653 and Grok Voice Think Fast the best conversation progression at 1.814 on a 1-to-3 scale, while Gemini-3.1-Flash-Live is third on the reward and last on turn-taking at 0.388. We can see how exceptional agentic performance does not necessarily imply that agents can maintain conversational dynamics and naturalness. We hope that this encourages the development of more capable speech-to-speech systems in the future.

Performance of voice agents across worlds

Performance of voice agents across worlds

Pass@1 for each system in each of the six worlds. The bar is the interval the system spans; each marker is one world, keyed above. Markers are nudged vertically where two worlds share a score.

We see that while the relative ranking of voice agents remains roughly unchanged, the subject matter of a world alone swings an agent’s performance by up to 0.474.

Reliability Analysis

Reliability Analysis

Reliability over k attempts, pooled over the six worlds. Dashed is Pass@k, at least one pass in k. Solid is Passᵏ, all k of k.


We can see how the best voice agents score <50% (Grok Voice Think Fast) on pass@1 and <20% on pass^5 agentic capability on DuplexWorld. This means that even today, voice agents lack the ability to be effective and robust on-call assistants.

Conclusion

If these voice agents must integrate reliably into enterprise and consumer AI workflows, we would need the development of most effective and reliable speech-to-speech agents. We’re glad to see the advent of speech-to-speech models from mere conversational systems to agents that are capable of handling complex tool calling and hope that we can see the emergence of much more powerful models soon.

What DuplexWorld changes about evaluating voice agents


6

worlds, 11 conversation types

156

scenarios authored, 144 scored

3,825

conversations, 387 simulated hours

0.490

best pooled Pass@1, of 1.000

0.266

best Pass³, all three of three

17.1%

of spoken credentials misheard 

The larger lesson from DuplexWorld might that identifying “best voice agent” may be the wrong objective. A system that excels at completing tasks may be weaker at managing a conversation, while one that sounds natural may struggle when the situation changes or the same task has to be completed reliably more than once. That makes voice-agent selection a question of fit: a bank, a healthcare provider, and a navigation service may need very different combinations of accuracy, conversational skill, and consistency. Better evaluation should make those tradeoffs visible so that organizations can choose agents based on what the job actually demands.


Learn more about Centific’s applied AI research here.  

Are your ready to get

modular

AI solutions delivered?

Centific offers a plugin-based architecture built to scale your AI with your business, supporting end-to-end reliability and security. Streamline and accelerate deployment—whether on the cloud or at the edge—with a leading frontier AI data foundry.

Centific offers a plugin-based architecture built to scale your AI with your business, supporting end-to-end reliability and security. Streamline and accelerate deployment—whether on the cloud or at the edge—with a leading frontier AI data foundry.

Connect data, models, and people — in one enterprise-ready platform.

Latest Insights

Ideas, insights, and

Ideas, insights, and

research from our team

research from our team

From original research to field-tested perspectives—how leading organizations build, evaluate, and scale AI with confidence.

From original research to field-tested perspectives—how leading organizations build, evaluate, and scale AI with confidence.

Connect with Centific

Stay ahead of what’s next

Stay ahead

Updates from the frontier of AI data.

Receive updates on platform improvements, new workflows, evaluation capabilities, data quality enhancements, and best practices for enterprise AI teams.

By proceeding, you agree to our Terms of Use and Privacy Policy