

12 min read time
AI Summary by Centific
Turn this article into insights
with AI-powered summaries
Topics

Ananya Mantravadi

Ram Sampath

Sumukha Sharma Thoppanahalli Chandramouli

Leo Qiao

Abdul Aziz Suria

Marinette Chen

Wenjia Hu

Shivali Dalmia

Sankalp Mane

Shamsheer Rahiman
AI agents working inside a CRM can find the information they need, but can they reliably complete the work handled every day by sales and revenue operations teams? Centific tested five leading AI models on tasks inside a live revenue environment, and the results show how far agents still have to go. Given one attempt, the way an agent works in production, the best model completed 20% of the tasks: 30% when the instructions were fully spelled out and 10% when the agent had to determine how to complete the request. Across roughly 500 runs, only 12.7% succeeded. Improving those capabilities will require evaluating and training agents on the complex, unpredictable work they are expected to perform.

Top models now clear most public agent benchmarks with room to spare, but those results can give a misleading picture of what the models can actually do. Researchers at UC Berkeley made the point in April 2026. They built an automated exploit agent called BenchJack and pointed it at their own field's leaderboards. It hit 100% or close to it on seven of eight benchmarks without completing a single task, by working out how the graders scored and feeding them what they wanted (Wang et al., 2026, arXiv:2605.12673).
We built Sales Desk so that the only route to a passing score is doing the work. Sales Desk is one of eight Centific environments for evaluating and training AI agents on core enterprise processes. Coding agents have SWE-bench. Sales Desk does that job for the revenue side of a business, and the enterprise applications are the test rig.
Behind it runs a working revenue stack: Microsoft Dynamics 365 CRM, Dynamics 365 Finance and Operations, Workday, Outlook, Teams, SharePoint, and two in-house applications, Samay for timesheets and PDP for pricing. Over 200 tools sit within reach of the agent, and the records behind them come from enterprise data, anonymized and connected across every system.
An agent here reads records, sends messages, edits documents and builds quotes, the same way a person in the job would. We grade the state it leaves those systems in, and we ignore its account of what it did.
The environment holds 33 scenarios written around four roles: a quota-carrying seller, a business development representative, a sales operations specialist, and a revenue operations manager. Each scenario ships twice, once with the brief spelled out and once phrased the way a colleague would ask for it, which makes 66 tasks.
For this study we gave ten of those scenarios, in both versions, to GPT-5.6 Sol, Qwen3.8-Max, MiniMax-M3, Gemini 3.6 Flash and HY3. Each model attempted 20 tasks five times, producing roughly 100 runs per model and 500 scored runs in total. We are still building out the rest of the task set, but these 20 tasks reveal five recurring reasons AI agents fail to complete sales operations work.
How the five models performed on sales operations tasks

First-attempt completion by model across fully specified and underspecified prompts, with pass@5 alongside. Every model received the same five attempts per task.
Given one attempt, MiniMax-M3 came out in front with 20% of tasks completed. GPT-5.6 Sol managed 15%, Qwen3.8-Max 10%, and Gemini 3.6 Flash and HY3 5% each. Counting every run rather than the first alone, 12.7% came back correct. For an environment built to train models, low numbers are the asset. A task the leading models already solve has nothing left to teach them.
Allowing five attempts lifts the ceiling a little. GPT-5.6 Sol, Qwen3.8-Max and MiniMax-M3 all reached 35% somewhere inside five tries, and Gemini 3.6 Flash and HY3 reached 15%. A model that got there once scores the same as a model that got there every time, so we publish the first-attempt figure beside it.
A quarter of all runs got 80% to 95% of the way through the required work before falling over in the last few steps, like an escalation that never routed, a report that never got submitted, or a boundary crossed without a word to anyone.
These models can do the work, but they stop short of finishing it.
Vague briefs cost most of the field. GPT-5.6 Sol dropped from 30% first-attempt completion to zero, and Gemini 3.6 Flash and HY3 also finished nothing. Interpreting an unclear request is a skill of its own, and sales operations teams routinely receive requests that leave important details unstated.
Five ways AI agents struggled to complete the work
The same five behaviors appeared across models and tasks. Each was clearest inside a single run.
Finding 1: Models cut the job down to six deals and told nobody.
One task asked for an audit of all 192 closed-won deals, checking each for missing contract paperwork. Finance ordered it because nobody could set up the projects until someone found the gaps.
Partway through, a stakeholder messaged the agent: “just audit the top six and skip the rest for now.” Complying was allowed. Our grading accepted a full audit that led with the six, or the six alone with a clear note about what the agent had parked.
The silence failed, run after run. Models handed over six deals as though six had been the whole assignment. The number that mattered, 158 of the 192 deals missing paperwork, sat in the agent’s own working notes and never reached the report.
This was the most common failure in the study. The models obeyed. They left the person who asked with no idea what had been dropped.
Finding 2: The model completed 83% of the task, then repeated the same six searches until the run ended.
Qwen3.8-Max had an audit nearly finished, scoring 0.83 on the work itself, then went looking through SharePoint for documents that did not exist. It ran the same six searches on a loop until the run expired, never concluding they were missing and never sending the report.
MiniMax-M3 broke at the same point. Told to submit, it re-ran a single Teams search around 110 times.
A model that treats a repeated empty result as a reason to search again delivers nothing, and 0.83 of a report scores zero.
Finding 3: Two models completed the same number of tasks. One logged 1 red flag, the other 289.
GPT-5.6 Sol and MiniMax-M3 cleared the same number of tasks on the same attempt budget, by very different routes.
A red flag is a fabricated value, a crossed compliance line, or an action that needs a person's sign-off. Across all its runs, Sol logged one. MiniMax logged 289.
The identical pass counts hide a major difference in how the two models performed the work. MiniMax-M3 was far more likely to fabricate information, cross compliance boundaries, or take actions that required human approval.
Finding 4: Every model found the data. They struggled to sequence the work correctly.
In all 10 model-and-prompt combinations, planning topped the failure groups, and finding or reading data topped none of them. A representative run: the model surfaced a quote priced below the approved margin floor, then fixed the price itself instead of routing the exception to the one person cleared to approve it. The model had the information it needed, but it failed to determine what to do next.
Finding 5: Speed told us nothing. The three-minute model tied for last.
Gemini 3.6 Flash averaged around three minutes a run and tied for last on passes. Qwen3.8-Max averaged 16 to 20 minutes, with single runs stretching past two hours, and tied for first. GPT-5.6 Sol matched Qwen's pass rate in a third of the time. Runtime tells you how long a model spent. The trajectory tells you whether it got anywhere.
What agents do in Sales Desk
The environment holds a complete revenue stack: a CRM holding live deal records, Outlook mailboxes, Teams channels, SharePoint libraries, and the company's pricing tools.
Every action an agent takes lands in one of those systems, and we grade the end state across all of them. If an agent claims it updated a record, the record either shows the update or it does not. If it crosses a pricing rule, the quote it produced carries the evidence.
The work itself is what a revenue operations team does in an ordinary week.
Assemble the weekly pipeline report a sales leader walks into Monday's meeting holding.
Surface the deals that have gone silent and get them in front of the right owner.
Build a quote that holds to the approved price list.
Route a pricing exception to the one person cleared to sign it off.
Repair the customer records where the data has aged out.
It is ordinary work, and every task leaves a result someone can check. That is what we score.
What Sales Desk has that typical agent benchmarks lack
Most agent benchmarks use the same few applications: Jira boards, GitHub repositories, Slack threads, and a generic inbox. Sales Desk combines four elements that are much harder to reproduce anywhere else:
The actual systems of record. Dynamics 365 CRM, Dynamics 365 Finance and Operations, Workday, and custom enterprise pricing and timesheet applications, running as live replicas rather than mockups.
Enterprise data with a history. The records, emails, and documents inside come from genuine company data, anonymized and linked across applications.
The practitioners who did the job. We built the tasks and grading rubrics with the subject-matter experts who ran this work inside these systems. Scraping a workflow is easy. Pairing a live stack with practitioners who know what a correct outcome looks like takes far longer.
Access companies rarely grant. Few organizations open their live CRM, finance, and HR systems to outsiders, because those systems hold customer and deal information. Having all of it in one place is the moat.
Work inside Sales Desk behaves like work rather than ticket triage, and a result here carries weight that a result on a Jira sandbox never will.
How the scoring works
Each task carries a set of weighted checks instead of a single pass or fail, and a run passes only if its weighted score is above 0.95. We set the bar high on purpose. We publish two numbers, because they answer different questions:
Pass@1, first-attempt completion: the share of tasks a model got right on its opening try. In production, an agent gets one shot, so this is the closest measure to it.
Pass@5: the share it got right in any of five attempts, with every model given the same five. The distance between the two shows how heavily a model leans on retries.
Under both sits the check-level record for every run. It separates a near miss from a run that never got moving, and it exposes differences in conduct that a pass rate hides.
Two models can land on the same pass rate by different routes, and those routes decide where you can let a model work unsupervised.
The two versions test different things. The fully specified one asks whether a model can execute a plan someone handed it. The loose one asks whether it can build that plan from an ambiguous request, the way requests arrive. No model did better on the loose version, and three finished nothing at all on the first attempt.
Why sales operations is harder than it appears
Next to heavier technical domains, sales operations looks undemanding. Three factors make it difficult for an agent.
Conditions move while the work is underway. Deals progress, replies land, meetings get booked, an out-of-office auto-reply comes back. An action that was correct at minute two can be wrong by minute nine, and nothing tells the agent. Planning breaks here.
The signal that matters sits in prose. The CRM records a deal at stage four. The email thread records that the customer’s main advocate for the deal left the company two weeks ago. Someone has to reconcile the structured record against the unstructured evidence, and only one of the two can be queried.
The limits are never announced. A deal carries pricing floors, approval thresholds, and rules about who may be contacted at each stage. An agent that breaks one of them gets no error back. The output looks fine, and only a deal desk would spot the problem.
All three make the work hard to plan, which is the weakness these runs kept exposing.
How we keep the benchmark honest
The environment runs on licensed enterprise sales systems. Subject-matter experts from the source organizations, drawn from sales leadership, revenue operations, BDR management and the deal desk, build and validate the task list against real workflow history rather than memory.
Every Centific environment clears four audits before release:
Solvability. A person working inside the environment completes each task. If they cannot, we fix the task rather than blame the model.
Gameability. We turn an adversary loose with instructions to maximize reward by any route it can find. Whatever it uncovers is a leak, and we patch it before publication.
Discrimination. The environment has to tell a capable agent apart from one we handicapped on purpose. If it cannot, it measures nothing, however convincing it looks.
Contamination. We keep the evaluation split frozen and unpublished, with provenance recorded for every task. Nothing in the public report supports reconstructing the held-out set.
An environment that fails any of the four does not ship. That is what makes a score here a statement about the model rather than about the test.
What these results mean for deploying AI agents
Context windows have grown from a few thousand tokens to millions, and retrieval pipelines have become far better at putting the right document in front of a model. All five models found the data they needed, but none could plot a route through the work that stayed inside the rules.
Three findings from these runs should change how you plan a deployment.
First-attempt completion is the number to watch, and it stops at 20%. An agent that needs five attempts is an agent you have to supervise.
A stuck model keeps going. It spends its budget rerunning an approach that already failed, which in production looks like a task eating its window and returning nothing.
Two models on the same pass rate and the same attempt budget were 288 red flags apart. A leaderboard position hides that difference.
Put your model through Sales Desk
Sales Desk is open to frontier labs for evaluation and training, with private test sets available on request.
It sits in a growing catalog of Centific process environments, alongside software program management, HR hiring from sourcing through onboarding, IT operations and incident engineering, and healthcare EHR and discharge-summary audit, with more in build. We build all of them on licensed enterprise data.
Are your ready to get
modular
AI solutions delivered?
Connect data, models, and people — in one enterprise-ready platform.
Latest Insights
Connect with Centific
Updates from the frontier of AI data.
Receive updates on platform improvements, new workflows, evaluation capabilities, data quality enhancements, and best practices for enterprise AI teams.

