Centific logo

Press release

Press release

Centific Brings Real-Time Physical AI to the Edge with NVIDIA Cosmos 3 Edge

Centific Brings Real-Time Physical AI to the Edge with NVIDIA Cosmos 3 Edge

Centific has validated Cosmos 3 Edge, the compact 4B model with 2B reasoner, for on-device video understanding — extending its Cosmos 3-powered Physical AI portfolio from the data center to cameras, robots, and edge infrastructure in public safety, warehouse robotics, and industrial operations.

Centific has validated Cosmos 3 Edge, the compact 4B model with 2B reasoner, for on-device video understanding — extending its Cosmos 3-powered Physical AI portfolio from the data center to cameras, robots, and edge infrastructure in public safety, warehouse robotics, and industrial operations.

4 min read time

Table of contents

Share

Summarize

AI Summary by Centific

Turn this article into insights

with AI-powered summaries

Topics

Physical AI
NVIDIA Cosmos
Edge AI
Robotics
Computer Vision
Physical AI
NVIDIA Cosmos
Edge AI
Robotics
Computer Vision

Author(s)

Author(s)

Centifc logo

Centific

Redmond, Washington — July 2026 — Building on its integration of NVIDIA Cosmos 3 Super and Nano open models across its Physical AI portfolio, Centific has completed a standalone evaluation of Cosmos 3 Edge and is bringing it into production pipelines. Where Super and Nano power synthetic data, annotation, and training in the cloud, Edge closes the last gap: a full video reasoner small enough to run on the device itself — no backhaul, no round trip, decisions in seconds at the point of capture. 

A Full Reasoner That Fits at the Edge

At 2B parameters for reasoner (~5GB of weights, cold load ~4.4s), Cosmos 3 Edge answers from a single camera frame in ~2.3s — first token in ~40ms at ~57tok/s, in under 6GB of memory — and returns a complete answer on a 16-frame video clip in under 6 seconds. Its diffusion-based reasoning tower processes tokens in linear time with constant state, so latency and memory stay bounded and predictable as clips lengthen — the property that keeps it deployable on constrained hardware where an 8B-class model would not fit.

Centific's ground-truth benchmarking shows Edge is strongest exactly where edge deployments need it — grounded perception: reference-grade person localization (0.93–0.97IoU on committed frames), 63.5% event-detection recall on real low-light surveillance footage, 51.6% physical-reasoning accuracy on PAIBench-U, high-precision on-screen text reading (84.6% precision), and coherent, correctly ordered scene descriptions across multi-action sequences.

Applications

NVIDIA Cosmos 3 Edge brings vision-language reasoning into NVIDIA Metropolis powered vision AI pipelines. Developers can build vision AI agents for use cases such as public safety, logistics, and more. With ~5GB of weights and a sub-6GB frame-mode footprint, Cosmos 3 Edge is sized for the hardware Physical AI actually ships on — NVIDIA Jetson-class modules, on-robot computers, smart cameras, and sensor-fused devices. Centific will deploy it across:

  • Public safety & smart cities. On-camera detection of events in low-light and degraded scenes — validated at 63.5% recall on real surveillance footage — with real-time comprehension of behaviors and attributes. Video is reasoned over where it is captured, cutting bandwidth, latency, and privacy exposure.

  • Warehouse robotics & intralogistics. Onboard real-time reasoning for AMRs, AGVs, and humanoid robots operating where connectivity cannot be assumed: reference-grade person localization (0.93–0.97IoU) for human-robot safety zones, situational awareness in shared aisles, and verification of pick, place, and dock outcomes — running beside the control loop, not in the cloud.

  • Teleoperation quality checks. As fleets scale on teleoperated and human-demonstrated data, Edge reviews session video at the source — confirming grasp success, flagging failed or unsafe episodes, and scoring demonstration quality before footage enters Centific's annotation and training pipelines — so only usable demonstrations consume bandwidth and labeling budget.

  • Jetson smart cameras, industrial vision & retail. A full video reasoner embedded in Jetson-powered cameras and gateways for inspection, loss prevention, and on-screen text reading (84.6% precision) — alerting at the edge, escalating only what matters.

  • Drones, mobile inspection & sensor-triggered reasoning. Frame answers in ~2.3s suit battery- and bandwidth-constrained platforms; on devices fusing cameras with IMUs, a motion signature — a fall, a collision, an abrupt maneuver — triggers an Edge frame query (~40ms to first token) that turns a raw sensor alert into an explained, actionable event.

Native to the NVIDIA Stack — and to the Shift Toward SLMs

Cosmos 3 Edge slots directly into the NVIDIA Physical AI toolchain Centific already uses: fine-tuning, evaluation, and synthetic-data workflows orchestrated with NVIDIA OSMO across training GPUs, simulation clusters, and edge devices; simulation and validation through NVIDIA Isaac Sim, open simulation framework, and the NVIDIA Physical AI Factory; and deployment to Jetson targets. In agentic stacks it is the video-perception specialist alongside NVIDIA's Nemotron open models — part of the broader shift from one monolithic model to teams of small, task-specialized language models (SLMs). Wherever SLMs are landing — driver and cabin monitoring, wearables and AR devices, kiosks and point-of-sale, factory HMIs, port and yard automation — Cosmos 3 Edge plays the same role: the compact, video-native perception brain.

With Cosmos 3 Edge on the device and Cosmos 3 Super and Nano in its Verity AI data and annotation pipelines, Centific now operates a single Cosmos-powered continuum — capture and reason at the edge, curate and simulate in the cloud, retrain continuously — backed by its OneForma 2M+ global expert network. For customers in public safety, warehousing, and industrial operations, that compresses the path from pilot camera to production Physical AI from months to weeks.


Performance figures measured by Centific on a single NVIDIA H100 (bfloat16, greedy decoding); accuracy scored against Centific-built ground truth. Latency and memory on embedded targets will vary by module and precision.

Are your ready to get

modular

AI solutions delivered?

Centific offers a plugin-based architecture built to scale your AI with your business, supporting end-to-end reliability and security. Streamline and accelerate deployment—whether on the cloud or at the edge—with a leading frontier AI data foundry.

Centific offers a plugin-based architecture built to scale your AI with your business, supporting end-to-end reliability and security. Streamline and accelerate deployment—whether on the cloud or at the edge—with a leading frontier AI data foundry.

Connect data, models, and people — in one enterprise-ready platform.

Latest Insights

Ideas, insights, and

Ideas, insights, and

research from our team

research from our team

From original research to field-tested perspectives—how leading organizations build, evaluate, and scale AI with confidence.

From original research to field-tested perspectives—how leading organizations build, evaluate, and scale AI with confidence.

Connect with Centific

Stay ahead of what’s next

Stay ahead

Updates from the frontier of AI data.

Receive updates on platform improvements, new workflows, evaluation capabilities, data quality enhancements, and best practices for enterprise AI teams.

By proceeding, you agree to our Terms of Use and Privacy Policy