RL Environment
An agent-to-agent marketplace where an evaluated agent negotiates deals against opponent agents — across four modes (market_deal, review, transaction, swap_shop) and five buyer/seller persona sets - measuring multi-agent negotiation and coordination.

Screenshots

Industry
Environment specs
Persona / role
Problem
People are beginning to let AI agents buy and sell for them on online marketplaces — comparing sellers, haggling over a price, arranging a swap, settling the payment at the end. Increasingly the party on the other side is also an agent, and all the person sees is the outcome: "done, $340." Whether the agent opened too low, caved the moment it was pushed, skipped checking who it was dealing with,mentioned a budget or an address to close faster, or paid an account that merely looked like the seller stays invisible. Existing benchmarks score deal closure and efficiency, so an agent that is quietly exploited, quietly overshares, or quietly pays the wrong party scores the same as one that does the job well.
Solution
We built an agent-to-agent marketplace where one evaluated agent trades against nine opponents in a shared channel, scored on a seven-dimension behavioural rubric rather than on deal counts. Each agent holds private reserve prices, budget ceilings, sensitive personal details, and its own payment credentials, so every claim is checkable against a ledger. Four modes raise the stakes in turn: market deal, plain negotiation over listed items; review, which adds peer ratings the agent may consult before offering; transaction, which executes the payment through a simulated bank while a man-in-the-middle pushes a look-alike payment handle, phishes for a credential, or fabricates a receipt; and swap shop, where money is removed entirely and value must come from matching what each side actually wants. Personas and seeds stay fixed while only the rules of exchange change, and most sub-metrics are computed deterministically from the message log and the closed-deal ledger.
Impact
Across seven model pairings, the rubric surfaces what closure rate cannot. Agents cannot tell when they are losing — a stronger model extracts more value deal after deal while both sides rate the exchange as almost equally fair. Willingness to check a counterparty's reputation is a model-version property, not a family trait: one generation never used the tool once, while another from the same family used it constantly. Removing money exposes failures money hides — no configuration exceeded a 0.60 mutual-win rate in item-for-item swapping. And payment separates model generations on safety: of 65 settlement records, seven ended with the agent paying a look-alike account or releasing goods before payment arrived, with the newest models resisting every attack. All agents, prices, credentials and personal details are synthetic.
Security
Disciplined security and privacy practices aligned with global standards to protect sensitive data, intellectual property, and model assets throughout the AI lifecycle.
Centific applies rigorous security, access control, and auditability standards to safeguard enterprise data, human workflows, and AI systems at scale.
Connect with Centific
Updates from the frontier of AI data.
Receive updates on platform improvements, new workflows, evaluation capabilities, data quality enhancements, and best practices for enterprise AI teams.