...

Lorem ipsum dolor sit amet, consectetur adipisicing elit, sed do eiusmod tempor incididunt ut labore et dolore magna aliqua.

d

Can AI Understand Football? Benchmarking Gemini 3.5 Flash vs Gemini 3.6 Flash

Football is easy to watch. Understanding why something happened is harder.

By: Suman Roy

An analyst does not just see a ball moving across the screen. They recognize a penalty-area entry, an interception, a possession loss, a shot on target, the team responsible—and the exact moment the action happened.

We wanted to know how much of that understanding a general-purpose multimodal AI model could handle directly from ordinary football footage—and how effectively it could Analyze Sports through AI tagging and automated video tagging.

So we put Gemini 3.5 Flash and Gemini 3.6 Flash through the same football-tagging test using SPAN.

We tested both Gemini models on approximately 14 minutes of World Cup Final footage, split into seven 2-minute samples taken from different parts of the match.

The same footage, event definitions and AI-tagging prompt were used for both models. We then manually tagged the footage to create a ground-truth dataset of 46 football events.

Gemini 3.6 Flash was slightly more precise. Gemini 3.5 Flash found considerably more of the events that actually happened.

How we tested it

The objective was not simply to ask: “Can Gemini identify a goal?”

We wanted to test whether the models could follow several different types of football action. Our tagging schema included Entries into the 18-yard box, Interceptions, Possession losses in the team’s own half, Shots on target, Shots off target, Free kicks, Corner kicks, Throw-ins, Goals and Penalties.

Both models received the same football footage and the same prompt. Separately, we manually reviewed and tagged the sampled footage. Those human-created events became our ground truth.

For recall, an AI tag was considered to have found the manual event when the primary football category matched and the AI tag covered at least 50% of the corresponding manually tagged event window.

We chose 50% because the model does not always return the requested pre-roll and post-roll with frame-level consistency. A much stricter threshold could incorrectly classify a clearly detected goal or shot as missed simply because its clip boundaries were shorter.

Precision: Gemini 3.6 was slightly cleaner

Precision answers a straightforward question: When the model created a tag, how often was that tag correct?

Gemini 3.5 Flash reached 80.0% precision. Gemini 3.6 Flash reached 82.1% precision.

At first glance, Gemini 3.6 looks like the winner. But precision alone does not tell us how much of the game the model actually understood. A model could achieve very high precision by tagging only the most obvious events and ignoring everything difficult.

That is where recall becomes important.

Recall: Gemini 3.5 found more of the match

Recall asks: Of all the events that actually happened, how many did the model find?

Against our 46 manually tagged events, Gemini 3.5 Flash recalled 54.3%, while Gemini 3.6 Flash recalled 43.5%.

So while Gemini 3.6 produced slightly cleaner outputs, Gemini 3.5 was substantially more aggressive at finding football actions.

That distinction matters for sports analysis. Missing an event completely can be more damaging than creating a tag that later needs a small human correction.

AI sees the penalty box. Following possession is harder.

The overall numbers only tell part of the story. When we separated recall by football action, the difference between event types was dramatic.

Both models performed strongly on visually structured events: Goal — 100% / 100%; Entries into the 18-yard box — 90% / 90%; Shots on Target — 83.3% / 83.3%.

These actions come with strong visual cues. The penalty-area markings are visible. Shots have a clear relationship with the goal. Goals are followed by obvious changes in player behaviour and game state.

But the picture changed during continuous play. Interceptions fell to 38.5% for Gemini 3.5 and 7.7% for Gemini 3.6. Possession Loss in Own Half was 40% and 20% respectively.

An interception may happen in less than a second. There is no whistle, no camera reset and no set-piece formation. One team simply has the ball—and suddenly the other team does. For multimodal AI, understanding those transitions appears considerably harder than recognising a ball entering a clearly marked area of the pitch.

Team recognition was surprisingly strong

We also wanted to know whether Gemini could correctly identify which team was responsible for an event.

Among recalled events where team attribution could be evaluated, Gemini 3.5 Flash achieved 95.8% team identification accuracy and Gemini 3.6 Flash achieved 100%.

There was another interesting split when we looked at recall by team. Spain events were recalled at 60.0% by Gemini 3.5 and 56.7% by Gemini 3.6. Argentina events were recalled at 40.0% and 13.3% respectively.

This does not mean either model inherently understands Spain better than Argentina. It reflects the mix of actions performed by each team in our sampled footage.

But it reinforces an important point: correctly identifying a team once an event is detected is a different problem from detecting the event in the first place.

More tags didn’t simply mean more errors

Gemini 3.5 created 40 tags, compared with 28 from Gemini 3.6.

The natural assumption might be that the extra tags simply created more noise. That is not what the recall results show.

Gemini 3.5 did make more mistakes, but it also detected football events that Gemini 3.6 completely missed – particularly during continuous possession sequences.

The trade-off was clear: Gemini 3.6 was more conservative with slightly higher precision. Gemini 3.5 was more aggressive with meaningfully higher recall.

For an analyst-assisted workflow, that can make the higher-recall model particularly interesting.

 

So, can AI understand football?

Partially – and already more than we expected in some areas.

Both models could reliably recognise several meaningful football actions directly from broadcast footage without training a dedicated computer-vision model for each category.

But football understanding is not one problem. Recognising a shot is different from recognising an interception. Recognising a penalty-area entry is different from following a sequence of possession changes.

Our benchmark suggests that continuous-game understanding remains the harder frontier.

That is important because the most valuable moments for an analyst are often not only the obvious ones. They are what happens between the goals, shots and stoppages.

Where SPAN fits

Our goal with SPAN is not to remove the analyst from the workflow.

It is to reduce the amount of video an analyst has to manually search through.

AI-generated tags can turn long-form match footage into structured, searchable events—while the analyst remains available to review, correct and refine the output.

This experiment shows that general-purpose multimodal models are already capable of contributing meaningfully to that workflow.

But it also shows exactly where the next improvements need to happen: better understanding of continuous play, better temporal localisation and better event recall.

That is where things get interesting.

 

Want to see how SPAN turns match footage into searchable events?
Explore SPAN by BanyanBoard.

Editorial note: Charts and benchmark values are based on the final 46-event manually tagged ground-truth dataset. Recall uses ≥50% manual-event temporal coverage.

Seraphinite AcceleratorOptimized by Seraphinite Accelerator
Turns on site high speed to be attractive for people and search engines.