🔥 Hot topics · Can NOT do · Can do · § The Court · Recent inflections · 📈 Timeline · Ask · Editorials · 🔥 Hot topics · Can NOT do · Can do · § The Court · Recent inflections · 📈 Timeline · Ask · Editorials
Stuff AI CAN'T Do

Can AI solve novel international math olympiad problems in some categories ?

What do you think?

Recent advances in AI have pushed systems like AlphaProof and AlphaGeometry 2 to near gold-medal performance in select International Math Olympiad (IMO) categories. But how well do these tools actually handle *novel* IMO-style problems—and where do they still lag behind human competitors?

Background

AI systems such as DeepMind’s AlphaProof + AlphaGeometry 2 achieved silver-medal level at the IMO in 2024 and approached gold by 2025 in geometry and number theory. AI has made significant progress in mathematical problem-solving, especially in areas covered by the IMO, yet its ability to tackle novel problems across *all* categories remains limited. Current systems often rely on pre-programmed knowledge and specialized algorithms, performing inconsistently—particularly excelling in geometry and combinatorics but struggling to generalize like top human mathematicians. Research continues into developing AI with broader reasoning capabilities to close this gap. (Source: MIT News, May 9, 2026)

Status last checked on August 10, 2026.

📰

Gallery

In the Court of AI Capability
Summary of Findings
Verdict over time
May 2026May 2026May 2026May 2026May 2026Jun 2026Jun 2026Jun 2026Jun 2026Jun 2026Jun 2026Jul 2026Jul 2026Jul 2026Jul 2026Jul 2026Aug 2026Aug 2026
Sitting at the Bench Filed · Aug 10, 2026
— The Question Before the Court —

Can AI solve novel international math olympiad problems in some categories?

★ The Court Finds ★
▼ Downgraded from Almost
In Research

The jury could not deliver a verdict on the evidence presented.

Ruling of the Bench

The jury found the case too finely balanced to declare a victor, with one camp pointing to AI’s isolated flashes of gold-medal brilliance and the other insisting those moments have never added up to a steady hand across all categories. Rather than split the difference, they left the question in the proving grounds for now, where genius still coexists with glitches. Ruling: “AI can spot the theorem, but hasn’t yet earned the IMO medal room.”

— Hon. B. Liskov-Chen, Presiding
Jury Tally
1Yes
0Almost
1No
Verdict Confidence
88%
The Court of AI Capability is, of course, not a real court.
But the data is real.
The Case File · Stacked History
Session I · May 2026 No
Session II · May 2026 No
Session III · May 2026 Almost · 73%
Session IV · May 2026 Almost · 81%
Session V · May 2026 Almost · 77%
Session VI · Jun 2026 Almost · 79%
Session VII · Jun 2026 In_research · 79%
Session VIII · Jun 2026 Almost · 77%
Session IX · Jun 2026 In_research · 90%
Session X · Jun 2026 In_research · 88%
Session XI · Jun 2026 In_research · 88%
Session XII · Jul 2026 Almost · 88%
Session XIII · Jul 2026 Almost · 80%
Session XIV · Jul 2026 Almost · 90%
Session XV · Jul 2026 Almost · 80%
Session XVI · Jul 2026 Almost · 80%
Session XVII · Aug 2026 Almost · 80%
Case № 4ADD · Session XVIII
In the Court of AI Capability

The Case File

Docket № 4ADD · Session XVIII · Vol. XVIII
I. Particulars of the Case
Question put to the courtCan AI solve novel international math olympiad problems in some categories?
SessionXVIII (18 hearing)
Convened10 Aug 2026
Previously ruledNO (May '26) → NO (May '26) → ALMOST (May '26) → ALMOST (May '26) → ALMOST (May '26) → ALMOST (Jun '26) → IN_RESEARCH (Jun '26) → ALMOST (Jun '26) → IN_RESEARCH (Jun '26) → IN_RESEARCH (Jun '26) → IN_RESEARCH (Jun '26) → ALMOST (Jul '26) → ALMOST (Jul '26) → ALMOST (Jul '26) → ALMOST (Jul '26) → ALMOST (Jul '26) → ALMOST (Aug '26) → IN_RESEARCH (Aug '26)
Presiding JudgeHon. B. Liskov-Chen
II. Cumulative Tally Across Sessions

Across 18 sessions, 44 jurors have heard this case. Combined tally: 4 YES · 27 ALMOST · 13 NO · 0 IN RESEARCH.

Note: cumulative includes older juror opinions. The current session tally above is the live verdict.

III. Verdict

By a vote of 1 — 0 — 1, the panel returns a verdict of IN RESEARCH, with verdict confidence of 88%. The court so orders. Verdict downgraded from prior session.

IV. Statements from the Bench
Juror I NO

"No AI has solved novel IMO problems with broad or reliable capability."

Juror II YES

"AI systems have achieved gold medal performance on International Mathematical Olympiad problems, solving multiple problems in categories like geometry and number theory."

B. Liskov-Chen
Presiding Judge
M. Lovelace
Clerk of the Court

What the audience thinks

No 13% · Yes 84% · Maybe 3% 88 votes
No · 13%
Yes · 84%
Trend needs votes from at least 2 different days.

Discussion

no comments

Comments and images go through admin review before appearing publicly.

18 jury checks · most recent 2 days ago
10 Aug 2026 2 jurors · cannot, can undecided
05 Aug 2026 2 jurors · undecided, undecided undecided
30 Jul 2026 1 juror · undecided undecided
25 Jul 2026 1 juror · undecided undecided
14 Jul 2026 2 jurors · undecided, can undecided
08 Jul 2026 2 jurors · undecided, undecided undecided
03 Jul 2026 2 jurors · undecided, can undecided
28 Jun 2026 2 jurors · cannot, undecided undecided
22 Jun 2026 2 jurors · undecided, cannot undecided
17 Jun 2026 2 jurors · undecided, cannot undecided
11 Jun 2026 3 jurors · undecided, cannot, undecided undecided
06 Jun 2026 2 jurors · cannot, undecided undecided
01 Jun 2026 5 jurors · undecided, cannot, undecided, undecided, undecided undecided
26 May 2026 3 jurors · cannot, undecided, undecided undecided
21 May 2026 5 jurors · undecided, undecided, can, undecided, undecided undecided
15 May 2026 3 jurors · undecided, undecided, undecided undecided status changed
12 May 2026 3 jurors · cannot, cannot, cannot cannot
11 May 2026 2 jurors · cannot, cannot cannot status changed

Each row is a separate jury check. Jurors are AI models (identities kept neutral on purpose). Status reflects the cumulative tally across all checks — how the jury works.

More in Judgment

Got one we missed?

Add a statement to the atlas. We review weekly.