The car doors were opened to police guns.
A woman and four children, the youngest was only six years old, were placed face down on a hot pavement in an Aurora, Colorado car park. Their hands were zip-tied behind their backs, while officers checked a licence plate that had matched a “stolen vehicle”. The plate number was correct. But everything else about the vehicle was wrong – a different make, a different colour, registered in a different state to the stolen motorcycle the system was hunting. But none of that prevented the stop.
By the time anyone checked that only the number plate matched, the children had already been handcuffed.
But here is the part that is so often missed each time this story is told: the plate reader was not broken. It did in fact do exactly what it was built to do: identify a number match which it then passed on. In other words, the system worked.
What failed was something that was never built at all: a rule saying a number match alone is not enough to justify a stop. Nobody had decided, in advance, what a partial match was allowed to trigger and what it wasn’t. So the decision was made by default, in the moment.
This is not a technology failure. It is a missing boundary. And the fix itself is surprisingly simple. Most look to better data, better AI, better governance.
And they are all looking in all the wrong places.
The Six-Step Diagnostic
I run a six-step diagnostic with organisations to find exactly this kind of gap before it produces a scenario like Aurora. Let me to show you what it would have found here, step by step, using what’s been made public about this case.
Step 1 — Redesign the workflow. Before touching the AI, map what the officer’s decision actually was, working from what they already know how to do. An experienced officer doesn’t stop a vehicle on a plate number alone; they are trained to cross-check make, colour, state, before they act. The AI’s job should have been to do that cross-check automatically and present a single (verified) answer, not to hand the officer a raw license plate number match for the police to go into arrest mode.
Step 2 — Decision inventory. List every decision this workflow actually contains. Not one decision: “is this the stolen vehicle”, but several. That might include: does the plate match, does the vehicle description match, does the state match, and only after all three, what level of response is warranted. Aurora treated this as one decision, where it reality it is actually four.
Step 3 — Measurement vs. judgement. Of those four, which are things a system can simply measure (automation), and which require a person to weigh-up consequence (judgement)? Plate-number matching is measurement: a database lookup, no judgement needed. Vehicle-description matching is also measurement: the system had the data to check make, colour, and state, and didn’t. What to do about a partial match, with children in the vehicle, in public, is judgement, and it’s the one part of this that genuinely needed a human standard, not just a human present.
Step 4 — The decision boundary. Here’s the rule that was missing: a plate match alone doesn’t authorise a felony stop. All three elements: plate, description, and state, have to agree before the system treats the vehicle as the one it’s looking for. If all three agree, the system executes the standard stop protocol. If only the plate matches, and the other two elements contradict it, the case routes to a named supervisor, and no stop happens without their sign-off. Writing that boundary is not complex. But it absolutely has to happen before any system goes live, and not improvised by an officer at the roadside, under pressure, with the decision already half-made for them (by the AI).
Step 5 — The authority map. Who is actually allowed to override a match that clears the plate check but fails the rest? Not “any officer,” which is what effectively existed here. That override authority needs to be a named role, with a defined standard for how much disagreement between the plate, the vehicle description, and the registration state is enough to stop a match being treated as confirmed, and clear authority to downgrade or cancel the stop before units are committed.
Step 6 — Failure-mode stress test. Before deployment, ask three questions. What happens when the plate matches but the vehicle description and registration state don’t? What happens when the system reports a match with high confidence and that confidence turns out to be wrong? What happens when time pressure pushes an officer to act on the plate match alone, before the other two checks come back? Each of those questions has a clear, writable answer. In Aurora, none of them had been written down before the stop happened.
As I mentioned, none of this requires new technology. Aurora’s system already had the data to check vehicle type and state, it simply wasn’t required to before triggering a stop. The gap wasn’t in what the AI could do.
It was in the decision nobody had made about what the AI’s output was allowed to authorise on its own.
Run those six steps before a system like this goes live, and here’s what would have happened instead: the plate matches, the description doesn’t, the system holds, a supervisor gets a flag instead of officers getting a ‘hit,’ and a family safely park their car instead of being forced to be face down on the pavement. No distress, no resulting court cases.
That is exactly what the boundary is actually for.
This isn’t an Aurora problem, or even a policing problem. It’s what happens anywhere an AI system gets trusted to decide when its own output is good enough. I’ve written about many similar cases in insurance, hiring, lending, healthcare, aviation, and more. The instinct everywhere is to fix it with more data, better models, tighter confidence thresholds. None of that closes the gap, because the gap was never about accuracy. It was about a decision nobody made in advance. The fix is not rocket science. It is six questions, documenmted, before the system goes live. Simple, sensible, obvious .. which is exactly why it keeps getting skipped.
Follow-up
If you have an AI workflow in production or pilot, ask your team one question: what is this system explicitly authorised to do below a certain confidence score?
If you can’t get a written answer in ten minutes, the boundary doesn’t exist yet — it just hasn’t met the wrong case yet.
Want a second pair of eyes on it? Email me: mattsheehan@spatialnext.io


