Technology
How to Test an AI Detection Claim Before You Buy: A Pilot Protocol for Security Buyers
A vendor tells you their detection is 99 percent accurate. Ask what the denominator is and the conversation usually stops.
That is not a gotcha. It is the whole problem. An accuracy percentage with no stated test set, no stated confidence threshold, and no stated error breakdown is not a measurable claim, and it cannot be compared against another vendor's percentage even when the two numbers look like they are describing the same thing. This piece gives you the two numbers that actually define a detection claim and a pilot protocol concrete enough to paste into an RFP.
We sell in this category. The protocol below is written so it can be run against us.
Precision and Recall Are the Pair That Define a Claim
Two error types exist and a single accuracy figure hides the trade between them.
Precision answers: of everything the system flagged, what fraction was real? Low precision means false positives. It means alarms that waste guard attention until people stop responding to them.
Recall answers: of everything that actually happened, what fraction did the system flag? Low recall means false negatives. It means the event you bought the system for went past the camera and nothing fired.
These pull against each other. Any detection system has a confidence threshold, and moving it trades one error type for the other. Lower the threshold and recall goes up while precision falls, so you catch more real events and eat more false alarms. Raise it and precision goes up while recall falls, so the alert queue gets quiet and misses start accumulating where nobody sees them.
This is why a lone accuracy number is not just incomplete, it is structurally misleading. A vendor can produce a very high number by tuning hard in either direction and reporting only the metric that flatters the result. Neither number means anything without the other, and neither means anything without the threshold they were measured at.
Nothing in the paragraphs above is a vendor claim or an industry statistic. These are the definitions of the metrics.
There Is No FRVT for General Security Analytics
For face recognition, an independent public benchmark exists. NIST runs the Face Recognition Technology Evaluation and Face Analysis Technology Evaluation programs, which are open to developer submission and, in NIST's own description, "an ongoing activity" that "runs continuously," with participants able to submit as often as every four calendar months (nist.gov, Face Technology Evaluations FRTE/FATE program page, retrieved 2026-08-13). The methodology is documented publicly, including in NIST IR 8491 (nvlpubs.nist.gov, retrieved 2026-08-13).
That is the shape of a real third-party evaluation. Fixed test data the vendor does not control. Published metric definitions. Open, repeated submission. Results anyone can read.
No equivalent exists for general security video analytics. There is no neutral public body publishing precision and recall for weapon detection, person detection, loitering, or vehicle classification on a fixed dataset across vendors.
The consequence for a buyer is direct. Because no neutral benchmark exists, the only number you can trust is the one produced on your own site, on your own scenes, over a defined window. Every published figure in this category, including any figure we might publish, was generated by the party selling the product on data that party selected.
We are also not going to quote you an industry-wide false alarm rate. The widely circulated figures on that point trace back to vendors selling in this same category rather than to any neutral source, which makes them marketing rather than evidence. Watch for that shape generally. A precise-sounding statistic attached to an impressive-sounding organization is a claim about provenance, not proof of it, and the fastest way to check is to find the number in that organization's own published material.
The Pilot Protocol
A pilot is the substitute for the benchmark that does not exist. Most pilots fail to produce a usable number because they were never designed to. Here is what makes one count.
Define the site and the scenes before anything is installed. Name the specific camera positions and what each one is watching. A loading dock at shift change, a rear perimeter gate overnight, a lobby during business hours. Detection performance is scene-dependent, so a number produced at the easy position tells you nothing about the hard one.
Define the window, and make it long enough to include the bad days. A two-week pilot in clear weather is a demo. Run it across the lighting and weather conditions the system will actually live in, including night, rain, and low sun angle. Thirty days is a reasonable floor. Ninety is better if the deployment is large.
Log ground truth independently. This is the step that gets skipped and the step that decides whether the pilot means anything. Someone has to establish what actually happened, separately from what the system reported. Reviewing recorded video on a sampled schedule is the usual method. If the only record of what happened is the system's own alert log, you have not measured recall, you have measured the system against itself.
Count both error types. False positives are easy to count because they arrive in the alert queue. False negatives are invisible by construction, which is exactly why the independent ground truth above is mandatory. A pilot that only counts false alarms measures precision and is blind to recall.
Fix the confidence threshold and record it. If the vendor retunes mid-pilot, the numbers before and after are not comparable and the pilot restarts. Agree the threshold in writing at the start.
Agree the walk-away number before the trial begins. Decide, in advance, what result ends the evaluation. Something like: fewer than three false alarms per camera per day, and at least 90 percent of staged test events detected. Set your own thresholds against your own operational tolerance. Setting them afterward means negotiating with a number you have already grown attached to.
Settle who counts and who pays. If the vendor tallies their own results, the tally is a sales document. Name the person on your side who owns the count, and put the cost of the trial in the contract rather than accepting a free pilot that quietly obligates you.
Three Questions Before You Sign Anything
| What the vendor says | What to ask instead | What a real answer looks like |
|---|---|---|
| "99 percent accurate" | Is that precision or recall, and what is the other one? | Both figures, stated together, at a named confidence threshold. |
| "Tested extensively" | Measured on what data, collected where, and who selected it? | A described dataset with scene conditions, event counts, and who controlled it. |
| "Industry-leading detection" | Compared against whom, on what shared test set? | An honest "no shared test set exists," followed by a pilot offer. |
| "Under X seconds detection" | Is that latency or accuracy? They are different questions. | Latency measured from event visible to alert delivered, stated separately from precision and recall. |
| "Certified" or "independently validated" | By which body, against which published methodology? | A named program with a public methodology document you can open yourself. |
The last two rows catch the most common conflations. A vendor who answers the first question with a shrug has told you the number was never measured in a way that survives the question.
Latency and Accuracy Are Different Questions
Our own published figure is gun detection in under 3 seconds. That is a latency specification. It describes the interval from a firearm being visible on camera to an alert reaching responders, and it is a hardware and pipeline claim about processing on the edge rather than routing frames to a distant cloud and back.
Latency says nothing about whether the detection was correct. A system can be fast and wrong. Vendors conflate the two constantly, usually by putting a speed number and an accuracy number in the same sentence so the reader carries the confidence from one to the other.
So hold us to the same protocol. Run the pilot, log ground truth independently, count both error types, and measure our latency separately from our precision and recall. We have written before about why the recovery window matters more than the failure rate and about what a 0.77 percent failure rate actually means in operation, and the principle is the same one. A number is only useful when you know what it was measured against.
Iron Gate's Position
We do not publish a precision or recall figure for general detection, because we have no neutral test set to publish it against and a self-generated number would carry exactly the problem this piece describes.
What we will do is design the pilot with you, agree the walk-away threshold in writing before it starts, and accept an independent count. If a competing vendor declines those terms, that is information worth more than any percentage on their datasheet.
Common Questions
How do I test a security camera AI detection claim before buying?
Run a pilot on your own site with defined camera scenes, a window long enough to include night and bad weather, independently logged ground truth, both error types counted, a fixed confidence threshold, and a walk-away number agreed in writing before the trial starts.
What does "99 percent accuracy" mean for gun detection or person detection?
By itself, nothing measurable. Ask whether the figure is precision or recall, what the other one is, what data it was measured on, and at what confidence threshold. Without those, two vendors' percentages cannot be compared.
What is the difference between precision and recall in video analytics?
Precision is the share of alerts that were real, so low precision means false alarms. Recall is the share of real events that were detected, so low recall means misses. Moving the confidence threshold trades one for the other, which is why a single accuracy number can hide either failure.
Is there an independent benchmark for security camera AI?
Not for general security analytics. NIST runs open, ongoing evaluations for face recognition and face analysis (retrieved 2026-08-13), but no neutral public body publishes comparable cross-vendor results for weapon, person, or vehicle detection. Your own pilot is the substitute.
How many false alarms per day is acceptable from an AI camera?
That is an operational decision rather than a technical constant, and it depends on who responds and what else they are doing. Set the number yourself before the pilot, based on the attention your team actually has, and treat it as a walk-away threshold rather than a preference.
What should a security camera pilot or proof of concept include?
Named camera positions and scenes, a defined start and end date, independent ground-truth logging, counts of both false positives and false negatives, a fixed and recorded confidence threshold, a named owner of the count on your side, and an agreed walk-away result. Put all of it in the contract.
Sources
- NIST Face Technology Evaluations (FRTE/FATE) program page, nist.gov, retrieved 2026-08-13
- NIST IR 8491, Face Analysis Technology Evaluation (FATE), nvlpubs.nist.gov, retrieved 2026-08-13. Cited as the model for a rigorous third-party evaluation. It is face-specific and is not a benchmark for general security analytics.
Ready to Talk Security?
Our engineering team can walk you through the right solution for your environment.
Book a Security Assessment