Benchmarks for agents that act, not just describe
Describing a scene is a solved demo. Doing the right thing in it, repeatedly, over a shift, is not. The metrics buyers are actually writing into their evaluations.

Every vendor demo ends the same way. The model looks at a clip and says what is in it. A forklift. A person in a red vest. A door left open. The audience nods. Then the system goes into a building with forty cameras and a night shift, and the question changes from "can it see" to "can it be trusted to act, every hour, without wearing people out."
The benchmarks that answer the second question are different from the ones that win the demo. Here are the ones showing up in real evaluations.
1. Alerts a human would thank you for
The number that matters is not detections. It is the ratio of alerts that deserved a person's attention to alerts that did not. One retailer's internal review, shared at a Spot AI planning session this summer, put the problem in a single figure: more than 65,000 alerts, and not enough asset-protection staff to process them. Spot AI's own first newsletter conceded that "over 90% of security alerts are harmless: think errant shopping carts or idling delivery trucks."
An advisor with a long career in remote guarding put the metric more precisely in a call with the company: measure mean time to first action, and the ratio of valid dispatches to customer complaints. Raw alarm counts, he argued, tell you nothing. An agent that finds ten real events and buries them in a thousand false ones has a recall score and no users.
2. Latency you can act inside
Spot AI advertises under one second from event to alert for its security agent. The reason that number is on the marketing page is that buyers ask for it. A partner integrating the platform into its own dispatch workflow reported that a clip fetch taking around 82 seconds was too slow for real-time response, and asked for a handful of still frames delivered instantly instead. The lesson generalises: the benchmark is not model inference time. It is the time until the person who has to do something has what they need.
3. Governance of the errors
The most telling document in this category is a manufacturer's evaluation checklist. A global animal-health company sent Spot AI ten numbered requirements before agreeing to a demo. Three of them were about being wrong on purpose:
- "False positive governance — How errors are measured, tuned, and managed over time"
- "Human-in-the-loop — How low-confidence or high-risk exceptions are routed to human reviewers"
- "Auditability — Whether alerts can be traced to source frame, timestamp, model version, reviewer, and disposition"
None of that is a computer-vision benchmark. All of it is what separates a tool that survives a compliance review from one that does not. Spot AI's engineers describe a proposer-verifier design in which a second model checks each candidate event before it becomes an alert, and the company has discussed letting partners choose whether they would rather tolerate more false positives or more false negatives. That choice is a benchmark too: the right error rate depends on what the alert triggers.
4. The outcome the budget owner recognises
The last benchmark is the only one a finance team reads. Not incidents detected, but a loss line that moved. Spot AI's published customer results are framed this way for a reason: a grocery chain with 27 percent fewer incidents across twelve stores, an apparel retailer whose cash shrink fell from 6 percent to 1 percent, a manufacturer reporting a 15 percent improvement in safety compliance. Whatever you think of vendor-reported numbers, the shape is right. An agent that acts is measured by what changed after it acted.
What to ask for
Put the four together and you have a scorecard a buyer can run in an eight-week pilot: the share of alerts that earned attention, the time to a usable alert, the audit trail behind each one, and a baseline loss figure taken before day one. Vision accuracy is the entry fee. It is not the score.