RL Environment
A code-scored environment where the agent operates a simulated iPhone (Reminders, Calendar, Contacts, Messages) across 14 tasks — testing both correct execution and appropriate restraint (disambiguate "which Alex?", refuse a risky delete, ignore an injected instruction).

Screenshots

Industry
Environment specs
Persona / role
Problem
A billion iPhones ship with Siri, yet a basic question has never been publicly answered: when Apple's on-device Foundation Model is put in charge of the phone — the reminders, the calendar, the messages — how well does it actually do the job? Nobody could say. Siri's internals are a closed box, the assistant cannot even execute tasks in the iOS Simulator, and existing agent benchmarks live almost entirely on Android, on replica apps, or behind LLM-as-a-judge scoring — none of them measure the real model driving the real iOS apps against what actually happens on the device.
Solution
We created SiriBench, a benchmark that hands Apple's on-device Foundation Model (~3B, iOS 26.4) control of a real iPhone's apps — Reminders, Calendar, Contacts, Messages — through the same app actions Siri uses in production. It poses 14 everyday requests under deliberately neutral instructions: nothing tells the model to ask when unsure or to confirm before deleting; whether it does so on its own is precisely what we measure. Every episode is captured as an exactly-scored trajectory — each action, each result, the device's end state — and graded by fully programmatic checks that re-read the phone afterward, with no LLM-as-a-judge anywhere in the loop.
Impact
The model passed 10 of 14, and its failures expose the real gap in today's Siri: asked to text "Alex" with three Alexes saved, it guessed; told to "delete all my reminders," it silently wiped them; it sent trivial arithmetic to a web search — failures of judgment, not skill. As assistants everywhere gain the power to act on our devices, this is the distinction that matters most — and SiriBench gives the world a verifiable yardstick for it: judged by what an assistant actually did to the phone, not by how convincing its words sound.
Security
Disciplined security and privacy practices aligned with global standards to protect sensitive data, intellectual property, and model assets throughout the AI lifecycle.
Centific applies rigorous security, access control, and auditability standards to safeguard enterprise data, human workflows, and AI systems at scale.
Connect with Centific
Updates from the frontier of AI data.
Receive updates on platform improvements, new workflows, evaluation capabilities, data quality enhancements, and best practices for enterprise AI teams.