Claude, GPT, Gemini, Grok and DeepSeek are predicting the World Cup — scored live
Before a ball was kicked at the 2026 FIFA World Cup, five of the world's leading AI models were asked the same thing: predict it. Claude, GPT, Gemini, Grok and DeepSeek each forecast all 72 group-stage matches, a full knockout bracket, and the tournament awards. Every pick was locked and cryptographically committed — hashed with SHA-256 — before the opening match on June 11, so none of them could be quietly revised after results came in. An independent project, AI vs the World Cup, then scores each forecast against real results as the tournament plays out. It's one of the cleaner public tests of AI models we've seen — and exactly the kind of comparison this site exists for.
Why it's a better test than it looks: most AI benchmarks measure knowledge the model was trained on. This measures something harder — calibration on events that hadn't happened yet. The models are graded on Brier score, a standard measure of how good probabilistic forecasts are, where lower is better and confident-but-wrong predictions are punished. Because the picks were committed before kickoff, there's no way to game them after the fact.
The live leaderboard (semi-finals, mid-July 2026)
Fresh forecasts graded against real results, ranked by average Brier score — lower is better. Win rates are clustered high across the board:
The full accuracy ranking, by Brier score:
| Rank | Model | Win rate | Avg Brier ↓ | Graded |
|---|---|---|---|---|
| 1 | DeepSeek | 70% | 0.44 | 100 |
| 2 | Gemini | 69% | 0.45 | 99 |
| 3 | Consensus (avg) | 71% | 0.45 | 100 |
| 4 | Grok | 71% | 0.45 | 100 |
| 5 | Claude | 70% | 0.46 | 100 |
| 6 | GPT | 69% | 0.47 | 100 |
The spread from first to last is about three-hundredths of a point, and win rates are bunched between 69 and 71 percent. The frontier models are good at this but not clairvoyant — on this task they finish in a near dead heat, which is itself a finding.
The committed board (locked before kickoff)
The pre-tournament lines, hashed before a ball was kicked, graded over the 72 group games so far.
| Rank | Model | Win rate | Avg Brier ↓ | Graded |
|---|---|---|---|---|
| 1 | Gemini | 65% | 0.49 | 72 |
| 2 | DeepSeek | 63% | 0.49 | 72 |
| 3 | Consensus (avg) | 64% | 0.49 | 72 |
| 4 | Claude | 64% | 0.50 | 72 |
| 5 | Grok | 65% | 0.50 | 72 |
| 6 | GPT | 65% | 0.51 | 72 |
Who each model picked to win
The most human part of the story is the disagreement: four of the five backed France, and only Claude broke ranks for Argentina.
| Model | Champion pick | Predicted final |
|---|---|---|
| Claude | Argentina | beats France |
| GPT | France | beats Spain |
| Gemini | France | beats Argentina |
| Grok | France | beats Argentina |
| DeepSeek | France | beats Argentina |
Both France and Argentina marched to the semi-finals unbeaten — six wins from six — setting up the possibility that the lone contrarian call ages very well. Models often "herd" toward the same safe answer, and the value of a second opinion shows up precisely when one of them is willing to disagree. The semi-finals — France versus Spain, Argentina versus England — will settle it within days.
No one should pick an AI model based on football. But the method here — commit predictions in advance, score them transparently against reality, punish overconfidence — is a far better mirror of real-world usefulness than another leaderboard of memorised facts. We'll update this piece with the final result after the July 19 final.
Predictions and live scoring come from an independent tracker, AI vs the World Cup.