The AI Decision Spine Model: The Maturity Path from Insight to Authority

In June 2026, the ACLU filed suit on behalf of Robert Dillon, a Florida man arrested after a facial recognition system returned a 93 percent confidence match against a photo purported to be him. Ninety-three percent is pretty decisive. Dillon was arrested and spent a night in jail.

For a crime he didn’t commit.

Nothing about this failure was a modelling problem. The facial recognition system worked as designed, it returned candidates, ranked by similarity. What was missing was a decision no one had made in advance: at what point does a confidence score become an action, and who is accountable for drawing that line?

That gap – not the models accuracy, nor the data quality – is what the AI Decision Spine Model is built to close.

Six stages, one fork

The diagram above walks through each of the AI maturity stages leading to the AI Decision Spine Model. Most organisations move through the first three stages without much friction.

  • Stage 0 to Stage 1 is digital transformation — consolidating data into one connected digital environment.
  • Stage 1 to Stage 2 adding an AI assistant (LLM) to that environment.
  • Stage 2 to Stage 3 moving to the rapidly advancing AI and data layers: grounding, automation, and AI-augmented human judgment. This layer includes simulation and world models.
  • Stage 3.5 is the trap, this is the rubber-stamp state; a human is nominally “in the loop” approving the AI’s output, but with no real standard, time, or standing to say no. It’s more dangerous than having no human at all, because it looks like oversight but functions as theatre. In Aurora: a plate-reader system returned an exact number match to a stolen vehicle — but nobody had a standard requiring the vehicle itself to be checked too, so a family was stopped and held at gunpoint, mistaken for a motorcycle in a different state.
  • Stage 3 to Stage 4 is the entry point to the AI Decision Spine Model. Stage 4 is centred on Decision Authority, it has 2 parts. Firstly, the spine; this defines the decision boundary (think of this as the law) written in advance, before any live case, setting out which decisions an AI’s output may settle on its own, which need a named human’s approval, and which are never delegated at all. Secondly, the guardrail; this applies the boundary to each live case (think of this as the judge): the mechanism that checks a specific decision against the spine and routes it – automate, escalate, or block – rather than trusting the AI’s confidence score to decide that for itself.
  • Stage 4 to Stage 5 is Decision Execution. Guardrail enforcement splits this between automation, where a decision has cleared the spine cleanly enough that the system can act on its own, no human required and augmentation, where the guardrail routes the decision to a named human instead, because it falls short of the automation bar and needs their judgement before it becomes an action.

The Importance of Stages 4 and 5 – and dangers of Stage 3.5

Stage 3.5 “Human in the loop” (HITL) is usually offered as the answer to AI risk. In reality it is most often the mechanism that hides it.

The academic term for what happens to that human is the “moral crumple zone,” coined by researcher Madeleine Clare Elish in a 2019 paper on accountability in automated systems. Just as a car’s crumple zone is engineered to absorb the physical force of a collision, a human placed nominally “in the loop” of an automated system can end up engineered – whether anyone intends it or not – to absorb the moral and legal force of its failures. The system gets to claim it has oversight. The human gets the blame when the oversight turns out to have been theoretical. Elish’s own case studies are aviation automation accidents and driverless-car incidents where a human was technically in control and functionally unable to be.

The clearest documented example is Uber’s autonomous test vehicle in Tempe, Arizona, in March 2018, which struck and killed a pedestrian, Elaine Herzberg. There was a human in the loop: a safety driver, present specifically to intervene if the system failed. The National Transportation Safety Board’s final report found the vehicle’s sensors detected Herzberg six seconds before impact and never braked, because Uber’s engineers had suppressed emergency braking to cut down on false-alarm stops, that was a design choice that quietly handed the actual decision back to the human. The NTSB named the underlying failure directly: Uber had built “inadequate mechanisms for addressing operators’ automation complacency.” The driver wasn’t reckless. She was doing exactly what a human put in that role for that many hours, watching a system that rarely needs intervening on, predictably does — she stopped actively monitoring it. The role was designed to fail this way; it just took time for it to.

This isn’t a one-off human failing. It’s what happens by default whenever “human in the loop” is treated as a design decision rather than a designed decision. A 2024 Harvard Business School study of 228 evaluators found reviewers were 19 percentage points more likely to defer to an AI’s recommendation when it came with an explanation, and a further 5 points more likely when that explanation was narrative rather than a bare score — meaning the more persuasive an AI system sounds, the less real scrutiny a human gives it. And the underlying constraint isn’t fixable by training or willpower: Caltech researchers Zheng and Meister found human conscious thought processes information at roughly 10 bits per second, against sensory intake running around a billion bits per second. Ask a person to meaningfully review AI output arriving faster than that gap allows, and there are only two outcomes: the organisation slows down and loses the speed it built the AI system to gain, or the human review becomes a formality that exists to be passed. Almost every organisation, under that pressure, ends up with the second without ever deciding to.

That’s the HITL cop-out: naming a human as the control without asking whether that human has the standard, the time, or the standing to actually exercise it. It satisfies a checklist. It survives an audit, until the audit asks the right question. And it has now caused fatalities (Uber/Herzberg), wrongful arrests at gunpoint (Aurora), and a teenager handcuffed over a bag of Doritos (Baltimore County) — different industries, different systems, the same structural gap. Stage 4 exists specifically to close it: not by adding another human to the loop, but by deciding, in writing, before the live case arrives, exactly what that human is there to do and what authority they actually have to do it.

The regulators just caught up

Two governments have now written this exact distinction into law, arriving at it independently, from liability law rather than systems design.

China’s Cyberspace Administration, NDRC, and MIIT jointly issued the Implementation Opinions on Intelligent Agents, effective 15 July 2026, the first national policy treating AI agents as their own regulated category. Article 6 requires that before an agent is deployed, its decisions need to be sorted into the same three tiers as the spine: human-only, approval-required, autonomous.

That is Stage 4, legislated.

Article 7 goes further, and it’s the more useful part. Chinese legal commentary has summarised the liability principle as “look at control, not code”, responsibility falling on whoever controlled the agent’s behaviour, not whoever built it. But control can only be judged after the fact if there is a record of which tier a decision fell into and whether approval actually happened. A tiering policy with no such trail is, in the regulation’s own framing, unreconstructable, so paper, not proof. That’s Article 7 legislating against exactly the HITL cop-out: a declared boundary that nobody can show actually held.

Illinois answered the same pressure differently. On 6 July 2026, Governor JB Pritzker signed the Artificial Intelligence Safety Measures Act, making Illinois the first US state to mandate annual independent third-party audits of frontier AI developers with audit fees structured so they can’t be tied to the findings, just to keep the auditor honest.

So we have different subjects, and different mechanisms. One tiers and traces an agent’s own behaviour, the other checks a developer’s safety practices from the outside. But there is a shared demand: a policy is not evidence. Only a record, audited or traceable, closes that gap.

Which is, in miniature, the entire case for Stage 4.

Follow-up

If you have an AI workflow in production or pilot, ask your team one question: “What is this system explicitly authorised to do below a certain confidence score?”

If you can’t find that answer in writing within ten minutes, the boundary doesn’t exist yet — it just hasn’t been caught out by the wrong case yet.

Want a second pair of eyes on it? Email me: mattsheehan@spatialnext.io


Sources

Leave a Reply

Discover more from SpatialNext

Subscribe now to keep reading and get access to the full archive.

Continue reading