It was World Cup season, and as always, half the world was busy betting on the winners. Which raises the real question: how do you beat the odds, and who's actually best at it?

Humans are the default. We watch, we argue, we back our favourites ... but confidence and accuracy aren't the same thing.

Now take LLMs, our loyal companions and decision making assistants. They're trained on a huge slice of world knowledge, so it's fair to wonder whether some genuine football understanding got swept up in there too. Can a model turn all that world knowledge into an informed prediction?

Then there's the wilder question. Remember 2010, when Paul the Octopus went eight-for-eight predicting World Cup matches? Do animals know more than we give them credit for? There was only one way to find out, so we included a few pets in the experiment as well.

We built a small prediction tournament. Humans brought their opinions. Large language models brought their training data. Pets brought pure instinct. May the best predictor win.

First, the fine print. The section below is collapsed, but it is worth opening: it covers the prompt every model received and the protocol we followed with the pets.

Meet the oracles

No amount of statistics beats watching an animal decide the fate of a nation by walking toward a bowl. Here are some of our contestants.

The scoreboard

Let's see the results. Here is who called the tournament best, ranked by correct picks. Filter by tribe to see how each one stacked up.

Before you scan the table, one thing needs explaining: there are two ways to be good at this, and they disagree with each other.

What the rank is built on
Correct picks
One point per match called right. Nothing else counts, and there is no credit for confidence, elegance, or a good excuse.
Shown alongside
Accuracy
Correct picks as a share of the matches you actually predicted. It says how often you were right, and says nothing at all about how often you showed up.

We rank on correct picks, because accuracy is the easier number to game. Sit out the genuine coin flips, predict only the matches nobody could get wrong, and you can walk away with a beautiful percentage and a handful of calls. Total correct picks has no such loophole: the only way up the table is to keep turning up and keep being right.

In fairness, nobody here was actually gaming anything. The pets who missed rounds weren't ducking the hard fixtures; their owners just had jobs, calendars, and the occasional match that kicked off during a nap. Which is exactly why accuracy sits in the table too: without it, you would never notice that Nia went 6 out of 8. On accuracy alone that is a top-ten finish. On points it is 26th, and a career cut tragically short by scheduling.

RankNameTribeAccuracyCorrectPredictions
1GLM 5.2
LLM
81%
1316
1Qwen
LLM
81%
1316
3Mr. L
Human
75%
1216
3RankBaseline
LLM
75%
1216
3Opus4.8
LLM
75%
1216
3Kimi K2.7
LLM
75%
1216
7Mr. A
Human
69%
1116
7Nemotron
LLM
69%
1116
7Minimax
LLM
69%
1116
7Gemma
LLM
69%
1116
11Mr. B
Human
63%
1016
11Pip
Petcat
63%
1016
11OddsBaseline
LLM
63%
1016
11Gpt5.5
LLM
63%
1016
11Gemini3.1
LLM
63%
1016
11Deepseek
LLM
63%
1016
11Ms. G
Human
67%
1015
19Miki
Petcat
56%
916
19Gams
LLM
56%
916
19Kiss
Petdog
64%
914
19Mr. Y
Human
60%
915
23Ms. M
Human
50%
816
23Mochi
Petcat
50%
816
25Holger
Petdog
70%
710
26Nia
Petdog
75%
68
26Gofio
Petcat
75%
68
26Monti
Petdog
40%
615
26Halee
Petdog
38%
616
30Belka
Petdog
31%
516
30Kure
Petchicken ensamble
63%
58
32Frida
Petcat
25%
416
33Kivi
Petcat
43%
37

Click a column to sort. Ranked by correct picks by default; accuracy normalizes for the fact that some players (mostly pets) sat out a few rounds.

Two models took the top: GLM 5.2 and Qwen tied for first, each calling 13 of 16 matches right (81%). The highest scoring human, Mr. L, came in just behind them in a four-way tie for third, and the highest scoring pet, Pip, appears in a tie for 11th. The whole upper half of the table is machines and humans, with the LLMs everywhere near the top and filling most of it. Not a runaway, but a clear showing.

81%
Top score (GLM 5.2 & Qwen, LLMs)
69%
Avg LLM accuracy
64%
Avg human accuracy
53%
Avg pet accuracy

Three groups, three different distributions

The leaderboard shows who came out on top. When we plot every contestant's accuracy on one line, grouped by tribe, the three groups' distributions start to look very different. Each dot represents one participant, the labelled diamond marks the group average, the whiskers reach one standard deviation either side (how spread out each group is), and the dashed line at 50% shows what we would expect from a pure coin flip.

Pets:
one playertribe average±1 SD (spread)coin flip (50%)
Pets
LLMs
Humans
Pets: randomness, beautifully displayed
The pets finished almost exactly where we would expect random guesses to land: around 50%. Their scores are spread widely, from 25% to 75%, a standard deviation of 17 points, by far the largest of any group. That is exactly the point: with only 7 to 16 games each, even a row of identical coin-flippers would scatter this widely, so the spread is the randomness, not a sign of hidden talent. Some got lucky, some did not, and the distribution shows that very honestly.
LLMs: consistent and hard to surprise
The models clustered between 56% and 81%, a standard deviation of 8 points, the narrowest spread of any group, alongside the highest average overall (69%). That is what pattern-based prediction looks like here: they mostly saw the same signals, made similar choices, and ended up in the same good-but-not-magical range, and the two that edged ahead, GLM 5.2 and Qwen, did it on the strength of a single extra pick.
Humans: informed, and closer to the models than you'd expect
The humans landed in almost the same band as the LLMs, from 50% to 75%, with a standard deviation of about 9 points, barely wider than the models', and not the wild scatter you might expect from gut calls. Their average, 64%, sits just below the models' 69%. That gap is visible on the page, but it is small, and it rests on only six people, which raises the obvious question: is it a real difference, or just noise?
Are LLMs genuinely better than humans?

To find out, we decided to bootstrap the differences among the two groups. Draw eleven models and six humans at random from the ones we have and get their predictions. From each of the subgroups calculate their accuracy and then substract it. By repeating this a hundred thousand times we get a bootstrap distribution of the group differences.

LLMs ahead in 92% of rerun contestsHumans ahead in 8%

Typical gap +5.4 points, but the middle 95% of rerun contests runs from -2.1 to +13.4 (grey lines), and that range includes zero.

The typical rerun favours the models by about 5 points, but the middle 95% runs from 2 points behind to 13 ahead. Zero sits inside that range, so we cannot call the models genuinely better.

The bigger problem is not that trivial. Resampling reshuffles the same six people, so it shows how fragile the result is without telling us whether those six people are representative of the broader population. We could have recruted six footbal experts and the models would probably have lost. Recruit six people who have never watched a match and even the pets might have looked good.

No amount of resampling fixes that. What we can honestly claim is narrow: among these entrants, on these sixteen matches, the models edged ahead. The separation the numbers genuinely do support is the other one: the pets show what randomness looks like, while the models and the humans both show what informed prediction looks like, landing in strikingly similar territory.

Are bigger models better?

The LLMs were the tribe we could actually put numbers on, not just numbers from. The most obvious suspect for what might make a model a better forecaster is simply how big it is: a bigger model has soaked up more of the internet, so presumably it knows more football. Does that show up in the accuracy at all?

Fit on:

Correlation: r = 0.14 (very weak), fitted on 10 models, including three whose size is only an estimate (hollow dots).

Across the whole field the line is close to flat: raw size, on its own, barely tracks accuracy. The biggest models didn't run away with it. Our largest entry, GPT-5.5, landed mid-table, while the two models that tied for first, GLM 5.2 and Qwen, sit nowhere near the top of the size axis.

There is a catch worth being upfront about. The three hollow dots are the closed models (Opus 4.8, GPT-5.5, Gemini 3.1), and none of their labs publish a parameter count, so those positions are rough estimates we pieced together from public reporting and analyst guesses, not numbers to bank on. Since they are also large and only middling scorers, they visibly flatten the line. So the chart has a switch: flip it to Published sizes only and those three grey out, leaving a fit drawn purely on the eight open-weight models whose size is documented. Do that and the tilt turns weakly positive (r ≈ 0.45).

Which number is the real one? Honestly, neither is worth much: eight noisy points is not enough for a correlation the statistics would stand behind, and the other view leans on sizes we had to guess. That is rather the point of showing both. Whichever way you read it, bigger did not reliably mean better.

Our comments

After the matches were over and the table stopped changing, a few patterns were too funny not to mention.

Reverse Frida would have finished near the very top. Frida had the worst accuracy of anyone who predicted a full slate, getting only 4 of her 16 picks right (25%). But take every one of her calls and choose the opposite, and she goes 12 for 16, a 75% hit rate, enough to tie for third on accuracy, behind only the two models that won it. A perfectly unreliable predictor is just a reliable one wearing a disguise.

Survivorship bias is the reason animal “oracles” become famous. We remember the lucky ones and forget the rest.

Paul the Octopus became a legend after going 8 for 8 at the 2010 World Cup, even though a perfect streak of eight binary picks has about a 1 in 256 chance of happening by luck alone (Wikipedia).

Punxsutawney Phil is the less glamorous version of the same story. He has been predicting the weather since 1887, and his accuracy is usually reported somewhere around 35% to 40%, which is worse than a coin flip (Wikipedia).

Experimental setup. While we did strive for the most objective experiment setup, a few questions arose during the experiment.

How do we reduce the influence of habit? What if a pet always heads for the left bowl simply because that's where it usually gets fed? We could move the feeding spot, but it's not obvious that this actually fixes the problem. And since we wanted the bowls placed 30 to 50 cm apart, we also wondered: could we cancel out habit altogether by putting them as close together as possible instead?

Should we run each pet more than once? A single trip to a bowl could be pure chance. But if we repeat it, does the pet start learning from the previous round instead of predicting fresh each time?

Final verdict

So, who was best at predicting football? Two language models, GLM 5.2 and Qwen, edged out the field, but only just, and not by a margin the statistics will vouch for. The fairest reading is that the informed predictors, human and machine alike, clustered together well above the pets, who gave us the cleanest picture of what random guessing looks like once you keep every contestant and not only the winners.

Finally, a big thank you to the pets, who had no idea they were part of a competition but still became the stars of it, and an even bigger one to their owners. Coaxing an animal toward a pair of food bowls before every kickoff is more work than it sounds, and none of this would exist without you.

A special thank you goes to the owners who kept their pets predicting through all 16 matches: Belka, Frida, Halee, Miki, Mochi and Pip. Their full slates are the reason we can draw honest conclusions instead of just collecting anecdotes. We know it took real time and effort across the whole tournament, and we are genuinely grateful for it.

And for the next tournament, we already know whose predictions we'll be quietly betting on: reverse Frida.