Video AgentsAll tags
Agentic VMS

Build a video agent by describing it

Custom video detection used to mean an ML team, thousands of labelled images, and six to eight weeks. Iris made it a ten-minute conversation. What that changes.

By Daniel Reyes · 2 min read
Iris launch graphic
Iris launched on April 9, 2025 and was demonstrated on stage at Google Cloud Next the following day. Image: Spot AI

For most of the history of computer vision, wanting a new detector was a project. You collected footage, labelled thousands of frames, trained a model, tested it, found it brittle, labelled more. Spot AI's own account of the old way: "dedicated AI/ML teams with advanced degrees, thousands of annotated images, and 6-8 weeks of complex development."

Iris, launched in April 2025, replaced that with a conversation. A user describes what they want to catch, shows the system about twenty example images, corrects it with positive and negative feedback, and has a working agent in around ten minutes. Then they wire it to an action: stop a machine, lock a door, send a clip.

What ChatGPT did for text, Iris does for video, making powerful AI capabilities accessible without requiring technical expertise.

That is Rish Gupta at launch. He demonstrated it on stage at Google Cloud Next the next day.

Why this is the important feature

It is tempting to file Iris under convenience. It is closer to a change in who gets to decide what a camera looks for.

Every site has its own definition of wrong. A car wash cares whether a brush is spinning. A pistachio processor cares whether a specific door is propped open. A laundromat cares whether someone is still inside after closing. None of those are in a vendor's catalogue, and none of them justify an ML engagement. A tool that lets the operations manager build the detector, in their own words, turns a thousand small problems into addressable ones.

The company describes three things Iris does. Train: teach it what matters, "no coding, machine learning expertise, or AI engineering required." Report: turn what it sees into structured reports. Understand: ask it questions about what is happening across facilities in plain language.

What it does not solve

Describing a thing is not the same as seeing it reliably. Twenty images do not cover night, rain, and a forklift blocking the view, which is why the same company runs a verifier model behind its agents and why customers still ask for pilots before they buy. And a builder that is easy to use is easy to misuse: it is now trivial to create an agent that fires constantly, and alert fatigue is the fastest way to get a system switched off.

Still, the direction is set. Two thousand agents were built on the platform in 2025, the company said in its year-end note, and the most used was not exotic. It was missing-PPE detection, the kind of thing a safety manager describes in one sentence.