VAR Slowed It Down. So Did I. Neither of Us Could Tell .. The Ref Still Sent Balogun Off

He never really had a choice.

Folarin Balogun collided with a Bosnia-Herzegovina defender. The on-field referee, in real time, did not send him off. Then VAR intervened. The video assistant referees in the booth pulled up slow-motion replays and still frames of the point of contact, these are the kinds of image that make any collision look worse than they actually were and advised the referee to initiate a review and go to the monitor. Once he stood in front of those pictures, there was only ever going to be one outcome. A red card.

USMNT were reduced to ten men, at the World Cup, on the back of images the review protocol itself says shouldn’t be used to judge a challenge like that one.

And that protocol (rule):

Slow motion = for finding where contact happened.

Normal speed = for judging how serious/dangerous it was.

Balogun’s red card was about seriousness, in other words was it a dangerous challenge or just an accidental coming-together. That is supposed to be judged at normal speed. Instead, the VAR judged it using slow motion. A shoulder brushing a shin looks like a stamp when you slow it down enough. The referee did what any referee would do when shown decisive-looking images by the people whose job is to catch his errors; he agreed with them.

So where was the failure?

This was not a judgement failure, it was a sensing failure, made by humans, using the wrong instrument for the job. As we will discuss in a moment on the Tielemans penalty in the Belgium game, low-motion footage watched by a person is a terrible way to measure contact and force. This is like comparing a linesman eyeballing an offside line versus what we now have in semi-automated AI offside technology. More on that topic in a moment.

The bottom line, the failure here was not because a human made a bad call. This was more a measurement question – how much contact, how much force – and what was handed to human perception squinting at distorted footage, instead of to something built to measure it.

This is the third VAR controversy inside 24 hours, and each one breaks in a different place.

Harry Kane went down in the box against DR Congo. No penalty, and no VAR review. The referee went with his own on field judgement. One might argue Kane embellished in his fall a little, but the question remains would he have stayed on his feet if the keeper hadn’t got there? That’s not something replay resolves. An AI model trained on trajectory, momentum, and contact geometry could probably give a better-calibrated answer to that than a referee’s gut in real time. So, a model could generate a probability, but not a fact. Is that a useful machine delivered input to a referee, to help him make a better decision?

Youri Tielemans went down against Senegal, deep into extra time. VAR reviewed for seven minutes, the longest delay in the tournament so far, before recommending that the referee to the monitor, who then awarded the penalty that won the game. With less than 1 minute left in extra time, this decision won the game for Belgium. I watched the challenge in real-time and in slow motion. Did Camara get the ball first or Tielemans foot? Did Tielemans dive or was the force enough to bring him down?

My eyes could not tell.

The layer nobody’s watching

In my Tuesday article, I discussed the referee – Alireza Faghani – standing in front of a perfect replay and still getting the Mbappรฉ call wrong. That’s a story about the referee, the game’s final arbiter, the one with the whistle and the final authority to act on what he sees.

Balogun is a different story. It’s not about whether the referee exercised good judgment. It’s about whether he was ever given the chance to. The VAR booth is a sensing system wearing a judgment system’s clothes โ€” three officials staring at slowed footage, trying to see something more clearly than they could in real time. Instead, the tool worked against them. They did their best with the wrong tool, and shared the same with the final decision-maker – the on-field referee.

In this conversation, I am using VAR as a way โ€” and I’ll admit not a perfect one, as David Basri rightly pushed back on in my last article โ€” to discuss the complexities of the evolving human-machine relationship, and the need to deliberately design the human layer, not just assume one exists because a person is technically present.

Both halves of that need designing; not just the judgment calls. The black-and-white decisions need a machine built to actually measure them correctly, and the judgment calls need a human given real authority, training, and protection to make them. Skip either half and the system fails, just in different ways.

In football we have the black-and-white decisions and the judgement calls. For offside, or whether the ball crossed the goal-line, these are decisions we can trust a machine to make. Where things become genuinely difficult is the non-black-and-white ..

The judgement calls.

Adding the machine

Our discussion here is centred on judgement, and that’s one part of the bigger design picture. But there’s one critical thing to add to this conversation specifically around VAR: slowed-down footage of an incident, reviewed by humans in a booth, then by a human on the field, driving a final decision; that chain is fundamentally flawed. I believe there’s a case for adding more technology in the booth, to augment every human in that loop, not replace them.

Contact-force estimation, biomechanical tracking, something closer to what already exists for offside lines and goal-line calls, would remove Balogun’s specific failure. A machine built to measure force and trajectory doesn’t distort the way slowed footage watched by a tired human eye distorts.

I believe more AI in the booth, not less, is the fix for that exact problem.

Let’s be honest, this does not fix the whole problem. Even perfect contact data still needs someone to decide the threshold – how much force, sustained for how long, in what part of the challenge, counts as a red card rather than a coincidence. That’s not a measurement. It’s a judgment about what the measurement means, and it’s a human decision no sensor can make. Better sensing in the booth doesn’t remove the human from the loop. It removes the bad version of the human from the loop .. the one squinting at distorted footage and mistaking a perception problem for a judgment call.

That leaves the real judgment, the threshold, exactly where it belongs.

Where this stops being about football

Football has more scrutiny on this relationship than any other domain on earth. A billion opinions within the hour. A governing body actively redesigning the protocol, the pitch-side monitor, the offside technology, in public, season after season. And in 24 hours, that system still surfaced three distinct ways for the human-machine relationship to fail; a sensing tool used for the wrong job, a genuinely irreducible ambiguity, and a question no replay could ever resolve because the event itself never happened.

If that’s what happens under maximum scrutiny, take a moment to consider what’s happening everywhere there is none.

This is where I want to close this article, because though I am a passionate football fan .. this is the world I’m actually focused on: not the boardroom decision you can revisit next quarter, but the physical-world decisions made under a running clock, that can’t be undone. An incident commander reading a model’s evacuation sequence as the fire turns. A grid operator acting on a load-shed recommendation with seconds to decide. A claims adjuster ratifying a total-loss call they half-doubt. A flood-response team routing resources off a prediction of where the water goes next.

Most of what organisations call governance in these settings is the audit, the ethics policy, the documented sign-off, the same layer VAR already has in its protocols and training and review boards. That’s necessary. It is not the same as designing the human layer, and Balogun is the proof: a fully credentialed, protocol-trained booth still failed, not because nobody was watching, but because the tool they were watching with was wrong for the question being asked.

Which brings me back to the question I left open with Kane and Tielemans. Suppose the model is good; well-calibrated, trained on real trajectory and momentum data, genuinely better than a gut read under pressure. It hands the person in the loop a number. What authority does that person actually retain over it? Can they override a 90% and be backed for it, or does overriding it look, on paper, like ignoring the evidence? If nobody has answered that before the number ever reaches them, you haven’t built a human layer. You’ve built a more precise-looking version of the dashboard Sarah couldn’t push back on.

That is the hard part, and it’s why I don’t think this gets solved by better sensors alone, or by braver individuals alone. It gets solved, if it gets solved, by designing both halves deliberately: machines built to measure what can actually be measured, and humans given real, protected, trained authority over what can’t.

Critically before the system goes live, not after the first person pays the price for a missing design.

Football is showing us this in real time, argued about by a billion people, and it’s still this hard. Everywhere else, it’s harder, and almost nobody is watching.

Has your organisation designed the human layer .. or just assumed it’s there?

If you’re not sure, that’s usually your answer.

I write Decision Layer Weekly for people working on exactly this gap: Subscribe here. And if it’s live for you right now, reach me directly: mattsheehan@spatialnext.io


Matt Sheehan

Matt is a geographer and AI strategist with 25 years at the intersection of geospatial intelligence and decision-making. He maps the architecture connecting three layers most organisations haven’t yet seen together: the sensing layer the geospatial industry has built, the causal reasoning layer now arriving, and the human decision layer nobody is designing. The third layer is where most deployments fail. And where Matt has his primary focus.

Leave a Reply

Discover more from SpatialNext

Subscribe now to keep reading and get access to the full archive.

Continue reading