RL Environment
RL environment for enterprise customer-support agents that trains with zero human annotation. A Dialogue Agent and an adversarial, GRPO-trained Customer Agent co-evolve in a multi-turn banking workspace, where the Customer Agent conceals intent to induce misroutes and the Dialogue Agent retrains on its own failure transcripts to route calls correctly. Tasks span transactions, disputes, card issues, transfers, and fraud escalations.

Screenshots

Industry
Environment specs
Persona / role
Problem
Enterprise customer-support AI must handle messy, unpredictable users, but training it to be robust requires large amounts of labelled conversation data. Real support logs are private, expensive to annotate, and quickly go stale as customer behavior changes. So agents look good on benchmarks but then misroute real customers — sending a fraud dispute to the wrong tool, or escalating a call that should have been resolved — confidently and at the worst possible moment.
Solution
We built RL-ADA, which trains agents without any human labels by using "world feedback": reward taken directly from what happens in each conversation. Two agents compete and improve together. A support agent is rewarded for routing customers to the right tool; an adversarial "customer" agent is rewarded for sounding realistic while tricking it into the wrong one. Whichever is losing gets retrained on its own failures, then they rematch, with an automated judge scoring every episode.
Impact
Across five cycles in a banking test, with zero labelled data, routing errors dropped to zero and the strict success rate doubled. The adversarial agent even invented its own trick — hiding intent inside dense, realistic detail — giving enterprises a self-improving way to stress-test agents before deployment.
Security
Disciplined security and privacy practices aligned with global standards to protect sensitive data, intellectual property, and model assets throughout the AI lifecycle.
Centific applies rigorous security, access control, and auditability standards to safeguard enterprise data, human workflows, and AI systems at scale.
Connect with Centific
Updates from the frontier of AI data.
Receive updates on platform improvements, new workflows, evaluation capabilities, data quality enhancements, and best practices for enterprise AI teams.