Skip to content

US vs China AI race

For each public test: the best US model, the best Chinese model, and the gap between them. A model counts as American or Chinese based on where Epoch AI lists the company that made it.

Epoch Capabilities Index

Epoch Capabilities Index: best model over time

Higher is better. Each dot is a model, placed on the day it came out. The line steps up when a model sets a new record.

Best in US

167.33

Claude Opus 5.5

Anthropic, released 22 Sep 2026

Epoch AI, Epoch Capabilities Index

Best in China

157.45

Kimi K3

Moonshot AI, released 16 Jul 2026

Epoch AI, Epoch Capabilities Index

On Epoch Capabilities Index, the best US model, Claude Opus 5.5 at 167.33, is 9.88 points ahead of the best Chinese model, Kimi K3 at 157.45 (n = 65 scores). Released Sep 2026 (US) and Jul 2026 (China).

The scale doesn't start at zero, so gaps look bigger than they are.

Tap a dot to see the model, its exact score and the source. Arrow keys work too.

Show every score as a table
Epoch Capabilities Index: every US and Chinese score
ModelCountryTestScoreReleasedSource
GPT-6.1 SolOpenAIUSEpoch Capabilities Index166.09Epoch AI, Epoch Capabilities Index
Claude Sonnet 5.5AnthropicUSEpoch Capabilities Index165.03Epoch AI, Epoch Capabilities Index
Claude Opus 5.5AnthropicUSEpoch Capabilities Index167.33Epoch AI, Epoch Capabilities Index
GPT-6 SolOpenAIUSEpoch Capabilities Index162.72Epoch AI, Epoch Capabilities Index
DeepSeek V4.1 FlashDeepSeekChinaEpoch Capabilities Index154.9Epoch AI, Epoch Capabilities Index
GPT-6 AstraOpenAIUSEpoch Capabilities Index166.45Epoch AI, Epoch Capabilities Index
Muse Spark 1.3Meta AIUSEpoch Capabilities Index156.75Epoch AI, Epoch Capabilities Index
Claude Fable 5.1AnthropicUSEpoch Capabilities Index164.7Epoch AI, Epoch Capabilities Index
Qwen3.8 Max (0902)Alibaba (Qwen)ChinaEpoch Capabilities Index155.05Epoch AI, Epoch Capabilities Index
GLM-5.3-FlashZ.ai (Zhipu AI)ChinaEpoch Capabilities Index151.88Epoch AI, Epoch Capabilities Index
GLM-5.3Z.ai (Zhipu AI)ChinaEpoch Capabilities Index155.61Epoch AI, Epoch Capabilities Index
DeepSeek V4 Pro 0813DeepSeekChinaEpoch Capabilities Index155.31Epoch AI, Epoch Capabilities Index
Gemini 3.7 FlashGoogle DeepMindUSEpoch Capabilities Index157.27Epoch AI, Epoch Capabilities Index
Grok 4.6xAIUSEpoch Capabilities Index156.44Epoch AI, Epoch Capabilities Index
Qwen 3.8 MaxAlibaba (Qwen)ChinaEpoch Capabilities Index156.41Epoch AI, Epoch Capabilities Index
DeepSeek V4 Flash 0731DeepSeekChinaEpoch Capabilities Index154.49Epoch AI, Epoch Capabilities Index
Claude Opus 5AnthropicUSEpoch Capabilities Index162.78Epoch AI, Epoch Capabilities Index
Kimi K3Moonshot AIChinaEpoch Capabilities Index157.45Epoch AI, Epoch Capabilities Index
Inkling-SmallThinking Machines LabUSEpoch Capabilities Index150.15Epoch AI, Epoch Capabilities Index
GPT-5.6 SolOpenAIUSEpoch Capabilities Index161.66Epoch AI, Epoch Capabilities Index
GPT-5.6 TerraOpenAIUSEpoch Capabilities Index159.62Epoch AI, Epoch Capabilities Index
GLM-5.2Z.ai (Zhipu AI)ChinaEpoch Capabilities Index151.78Epoch AI, Epoch Capabilities Index
Claude Fable 5AnthropicUSEpoch Capabilities Index162.06Epoch AI, Epoch Capabilities Index
Nemotron 3 UltraNvidiaUSEpoch Capabilities Index146.17Epoch AI, Epoch Capabilities Index
MiniMax-M3MiniMaxChinaEpoch Capabilities Index146.95Epoch AI, Epoch Capabilities Index
Claude Opus 4.8AnthropicUSEpoch Capabilities Index158.21Epoch AI, Epoch Capabilities Index
Qwen3.7-MaxAlibaba (Qwen)ChinaEpoch Capabilities Index153.68Epoch AI, Epoch Capabilities Index
GPT-5.5OpenAIUSEpoch Capabilities Index159.1Epoch AI, Epoch Capabilities Index
GPT-5.5 ProOpenAIUSEpoch Capabilities Index162.07Epoch AI, Epoch Capabilities Index
Kimi K2.6Moonshot AIChinaEpoch Capabilities Index151.05Epoch AI, Epoch Capabilities Index
GLM-5.1Z.ai (Zhipu AI)ChinaEpoch Capabilities Index149.84Epoch AI, Epoch Capabilities Index
GPT-5.4 ProOpenAIUSEpoch Capabilities Index158.93Epoch AI, Epoch Capabilities Index
GPT-5.3 CodexOpenAIUSEpoch Capabilities Index156.77Epoch AI, Epoch Capabilities Index
Kimi K2.5Moonshot AIChinaEpoch Capabilities Index148.03Epoch AI, Epoch Capabilities Index
GPT-5.2 ProOpenAIUSEpoch Capabilities Index155.4Epoch AI, Epoch Capabilities Index
DeepSeek-V3.2DeepSeekChinaEpoch Capabilities Index146.27Epoch AI, Epoch Capabilities Index
Gemini 3 ProGoogle DeepMindUSEpoch Capabilities Index152.92Epoch AI, Epoch Capabilities Index
Kimi K2 ThinkingMoonshot AIChinaEpoch Capabilities Index146.01Epoch AI, Epoch Capabilities Index
GPT-5 ProOpenAIUSEpoch Capabilities Index150.28Epoch AI, Epoch Capabilities Index
DeepSeek-V3.2-ExpDeepSeekChinaEpoch Capabilities Index145Epoch AI, Epoch Capabilities Index
GPT-5OpenAIUSEpoch Capabilities Index150Epoch AI, Epoch Capabilities Index
Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba (Qwen)ChinaEpoch Capabilities Index143.85Epoch AI, Epoch Capabilities Index
o3-proOpenAIUSEpoch Capabilities Index147.42Epoch AI, Epoch Capabilities Index
DeepSeek-R1 (May 2025)DeepSeekChinaEpoch Capabilities Index141.29Epoch AI, Epoch Capabilities Index
Qwen3-235B-A22BAlibaba (Qwen)ChinaEpoch Capabilities Index139.35Epoch AI, Epoch Capabilities Index
o3OpenAIUSEpoch Capabilities Index146.86Epoch AI, Epoch Capabilities Index
Gemini 2.5 Pro (Mar 2025)Google DeepMindUSEpoch Capabilities Index144.16Epoch AI, Epoch Capabilities Index
DeepSeek-R1DeepSeekChinaEpoch Capabilities Index138.97Epoch AI, Epoch Capabilities Index
DeepSeek-V3DeepSeekChinaEpoch Capabilities Index132.34Epoch AI, Epoch Capabilities Index
o1OpenAIUSEpoch Capabilities Index141.91Epoch AI, Epoch Capabilities Index
Qwen2.5-72BAlibaba (Qwen)ChinaEpoch Capabilities Index129Epoch AI, Epoch Capabilities Index
Qwen2.5-32BAlibaba (Qwen)ChinaEpoch Capabilities Index128.52Epoch AI, Epoch Capabilities Index
o1-miniOpenAIUSEpoch Capabilities Index135.82Epoch AI, Epoch Capabilities Index
Llama 3.1-405BMeta AIUSEpoch Capabilities Index128.75Epoch AI, Epoch Capabilities Index
Claude 3.5 SonnetAnthropicUSEpoch Capabilities Index130Epoch AI, Epoch Capabilities Index
Qwen2-72BAlibaba (Qwen)ChinaEpoch Capabilities Index125.28Epoch AI, Epoch Capabilities Index
GPT-4o (May 2024)OpenAIUSEpoch Capabilities Index128.97Epoch AI, Epoch Capabilities Index
DeepSeek-V2 (MoE-236B, May 2024)DeepSeekChinaEpoch Capabilities Index124.77Epoch AI, Epoch Capabilities Index
GPT-4 Turbo (Apr 2024)OpenAIUSEpoch Capabilities Index127.25Epoch AI, Epoch Capabilities Index
Claude 3 OpusAnthropicUSEpoch Capabilities Index126.91Epoch AI, Epoch Capabilities Index
GPT-4 Turbo (Nov 2023)OpenAIUSEpoch Capabilities Index126.46Epoch AI, Epoch Capabilities Index
Yi-34B01.AIChinaEpoch Capabilities Index117.39Epoch AI, Epoch Capabilities Index
Qwen-14BAlibaba (Qwen)ChinaEpoch Capabilities Index113.03Epoch AI, Epoch Capabilities Index
GPT-4 (Mar 2023)OpenAIUSEpoch Capabilities Index125.89Epoch AI, Epoch Capabilities Index
LLaMA-65BMeta AIUSEpoch Capabilities Index110.17Epoch AI, Epoch Capabilities Index

Source: Epoch AI (CC BY 4.0). Each score links to its page in the table.

FrontierMath-Tiers-1-3-v2-Private

FrontierMath-Tiers-1-3-v2-Private: best model over time

Higher is better. Each dot is a model, placed on the day it came out. The line steps up when a model sets a new record.

Best in US

93.7%

GPT-6 Astra

OpenAI, released 3 Sep 2026

Epoch AI, Benchmarking Hub

Best in China

74.7%

Qwen 3.8 Max

Alibaba (Qwen), released 2 Aug 2026

Epoch AI, Benchmarking Hub

On FrontierMath-Tiers-1-3-v2-Private, the best US model, GPT-6 Astra at 93.7%, is 19 percentage points ahead of the best Chinese model, Qwen 3.8 Max at 74.7% (n = 35 scores). Released Sep 2026 (US) and Aug 2026 (China).

Tap a dot to see the model, its exact score and the source. Arrow keys work too.

Show every score as a table
FrontierMath-Tiers-1-3-v2-Private: every US and Chinese score
ModelCountryTestScoreReleasedSource
GPT-6.1 SolOpenAIUSFrontierMath-Tiers-1-3-v2-Private93.7Epoch AI, Benchmarking Hub
Claude Sonnet 5.5AnthropicUSFrontierMath-Tiers-1-3-v2-Private88.8Epoch AI, Benchmarking Hub
Claude Opus 5.5AnthropicUSFrontierMath-Tiers-1-3-v2-Private91.2Epoch AI, Benchmarking Hub
GPT-6 SolOpenAIUSFrontierMath-Tiers-1-3-v2-Private89.8Epoch AI, Benchmarking Hub
GPT-6 AstraOpenAIUSFrontierMath-Tiers-1-3-v2-Private93.7Epoch AI, Benchmarking Hub
Muse Spark 1.3Meta AIUSFrontierMath-Tiers-1-3-v2-Private74.4Epoch AI, Benchmarking Hub
Claude Fable 5.1AnthropicUSFrontierMath-Tiers-1-3-v2-Private90.2Epoch AI, Benchmarking Hub
Qwen3.8 Max (0902)Alibaba (Qwen)ChinaFrontierMath-Tiers-1-3-v2-Private65.6Epoch AI, Benchmarking Hub
GLM-5.3-FlashZ.ai (Zhipu AI)ChinaFrontierMath-Tiers-1-3-v2-Private55.8Epoch AI, Benchmarking Hub
GLM-5.3Z.ai (Zhipu AI)ChinaFrontierMath-Tiers-1-3-v2-Private68.8Epoch AI, Benchmarking Hub
DeepSeek V4 Pro 0813DeepSeekChinaFrontierMath-Tiers-1-3-v2-Private64.6Epoch AI, Benchmarking Hub
Gemini 3.7 FlashGoogle DeepMindUSFrontierMath-Tiers-1-3-v2-Private71.6Epoch AI, Benchmarking Hub
Grok 4.6xAIUSFrontierMath-Tiers-1-3-v2-Private66Epoch AI, Benchmarking Hub
Qwen 3.8 MaxAlibaba (Qwen)ChinaFrontierMath-Tiers-1-3-v2-Private74.7Epoch AI, Benchmarking Hub
DeepSeek V4 Flash 0731DeepSeekChinaFrontierMath-Tiers-1-3-v2-Private57.5Epoch AI, Benchmarking Hub
Claude Opus 5AnthropicUSFrontierMath-Tiers-1-3-v2-Private85.6Epoch AI, Benchmarking Hub
Kimi K3Moonshot AIChinaFrontierMath-Tiers-1-3-v2-Private72.2Epoch AI, Benchmarking Hub
Inkling-SmallThinking Machines LabUSFrontierMath-Tiers-1-3-v2-Private46.3Epoch AI, Benchmarking Hub
GPT-5.6 SolOpenAIUSFrontierMath-Tiers-1-3-v2-Private89.1Epoch AI, Benchmarking Hub
GPT-5.6 TerraOpenAIUSFrontierMath-Tiers-1-3-v2-Private86Epoch AI, Benchmarking Hub
GLM-5.2Z.ai (Zhipu AI)ChinaFrontierMath-Tiers-1-3-v2-Private59.2Epoch AI, Benchmarking Hub
Claude Fable 5AnthropicUSFrontierMath-Tiers-1-3-v2-Private87Epoch AI, Benchmarking Hub
Claude Opus 4.8AnthropicUSFrontierMath-Tiers-1-3-v2-Private80Epoch AI, Benchmarking Hub
Qwen3.7-MaxAlibaba (Qwen)ChinaFrontierMath-Tiers-1-3-v2-Private64.6Epoch AI, Benchmarking Hub
GPT-5.5OpenAIUSFrontierMath-Tiers-1-3-v2-Private85.3Epoch AI, Benchmarking Hub
GPT-5.5 ProOpenAIUSFrontierMath-Tiers-1-3-v2-Private87.7Epoch AI, Benchmarking Hub
Kimi K2.6Moonshot AIChinaFrontierMath-Tiers-1-3-v2-Private57.2Epoch AI, Benchmarking Hub
GLM-5.1Z.ai (Zhipu AI)ChinaFrontierMath-Tiers-1-3-v2-Private36.8Epoch AI, Benchmarking Hub
GPT-5.4 ProOpenAIUSFrontierMath-Tiers-1-3-v2-Private82.5Epoch AI, Benchmarking Hub
GPT-5.2 ProOpenAIUSFrontierMath-Tiers-1-3-v2-Private74Epoch AI, Benchmarking Hub
GPT-5 ProOpenAIUSFrontierMath-Tiers-1-3-v2-Private55.8Epoch AI, Benchmarking Hub
GPT-5OpenAIUSFrontierMath-Tiers-1-3-v2-Private55.4Epoch AI, Benchmarking Hub
o3OpenAIUSFrontierMath-Tiers-1-3-v2-Private33.3Epoch AI, Benchmarking Hub
o1OpenAIUSFrontierMath-Tiers-1-3-v2-Private14.7Epoch AI, Benchmarking Hub
GPT-4 Turbo (Apr 2024)OpenAIUSFrontierMath-Tiers-1-3-v2-Private0.7Epoch AI, Benchmarking Hub

Source: Epoch AI (CC BY 4.0). Each score links to its page in the table.

GPQA diamond

GPQA diamond: best model over time

Higher is better. Each dot is a model, placed on the day it came out. The line steps up when a model sets a new record.

Best in US

94.4%

GPT-6 Astra

OpenAI, released 3 Sep 2026

Epoch AI, Benchmarking Hub

Best in China

90.8%

Kimi K3

Moonshot AI, released 16 Jul 2026

Epoch AI, Benchmarking Hub

On GPQA diamond, the best US model, GPT-6 Astra at 94.4%, is 3.6 percentage points ahead of the best Chinese model, Kimi K3 at 90.8% (n = 54 scores). Released Sep 2026 (US) and Jul 2026 (China).

Tap a dot to see the model, its exact score and the source. Arrow keys work too.

Show every score as a table
GPQA diamond: every US and Chinese score
ModelCountryTestScoreReleasedSource
GPT-6.1 SolOpenAIUSGPQA diamond93.9Epoch AI, Benchmarking Hub
Claude Sonnet 5.5AnthropicUSGPQA diamond94.1Epoch AI, Benchmarking Hub
Claude Opus 5.5AnthropicUSGPQA diamond87.5Epoch AI, Benchmarking Hub
GPT-6 SolOpenAIUSGPQA diamond92.4Epoch AI, Benchmarking Hub
GPT-6 AstraOpenAIUSGPQA diamond94.4Epoch AI, Benchmarking Hub
Qwen3.8 Max (0902)Alibaba (Qwen)ChinaGPQA diamond89.7Epoch AI, Benchmarking Hub
GLM-5.3-FlashZ.ai (Zhipu AI)ChinaGPQA diamond86.9Epoch AI, Benchmarking Hub
GLM-5.3Z.ai (Zhipu AI)ChinaGPQA diamond87.9Epoch AI, Benchmarking Hub
DeepSeek V4 Pro 0813DeepSeekChinaGPQA diamond88.9Epoch AI, Benchmarking Hub
Gemini 3.7 FlashGoogle DeepMindUSGPQA diamond93.1Epoch AI, Benchmarking Hub
Grok 4.6xAIUSGPQA diamond92Epoch AI, Benchmarking Hub
Qwen 3.8 MaxAlibaba (Qwen)ChinaGPQA diamond90.2Epoch AI, Benchmarking Hub
DeepSeek V4 Flash 0731DeepSeekChinaGPQA diamond88Epoch AI, Benchmarking Hub
Claude Opus 5AnthropicUSGPQA diamond91.8Epoch AI, Benchmarking Hub
Kimi K3Moonshot AIChinaGPQA diamond90.8Epoch AI, Benchmarking Hub
Inkling-SmallThinking Machines LabUSGPQA diamond84.7Epoch AI, Benchmarking Hub
GPT-5.6 SolOpenAIUSGPQA diamond91.3Epoch AI, Benchmarking Hub
GPT-5.6 TerraOpenAIUSGPQA diamond91.1Epoch AI, Benchmarking Hub
GLM-5.2Z.ai (Zhipu AI)ChinaGPQA diamond89.1Epoch AI, Benchmarking Hub
Claude Fable 5AnthropicUSGPQA diamond81.1Epoch AI, Benchmarking Hub
Nemotron 3 UltraNvidiaUSGPQA diamond80.5Epoch AI, Benchmarking Hub
MiniMax-M3MiniMaxChinaGPQA diamond87.9Epoch AI, Benchmarking Hub
Claude Opus 4.8AnthropicUSGPQA diamond88Epoch AI, Benchmarking Hub
Qwen3.7-MaxAlibaba (Qwen)ChinaGPQA diamond87.9Epoch AI, Benchmarking Hub
GPT-5.5OpenAIUSGPQA diamond92Epoch AI, Benchmarking Hub
GPT-5.5 ProOpenAIUSGPQA diamond91.9Epoch AI, Benchmarking Hub
Kimi K2.6Moonshot AIChinaGPQA diamond87.7Epoch AI, Benchmarking Hub
GLM-5.1Z.ai (Zhipu AI)ChinaGPQA diamond86.5Epoch AI, Benchmarking Hub
GPT-5.4 ProOpenAIUSGPQA diamond92.8Epoch AI, Benchmarking Hub
Kimi K2.5Moonshot AIChinaGPQA diamond83.5Epoch AI, Benchmarking Hub
DeepSeek-V3.2DeepSeekChinaGPQA diamond77.9Epoch AI, Benchmarking Hub
Gemini 3 ProGoogle DeepMindUSGPQA diamond90.2Epoch AI, Benchmarking Hub
Kimi K2 ThinkingMoonshot AIChinaGPQA diamond79Epoch AI, Benchmarking Hub
GPT-5OpenAIUSGPQA diamond81.6Epoch AI, Benchmarking Hub
Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba (Qwen)ChinaGPQA diamond73.4Epoch AI, Benchmarking Hub
DeepSeek-R1 (May 2025)DeepSeekChinaGPQA diamond68.4Epoch AI, Benchmarking Hub
Qwen3-235B-A22BAlibaba (Qwen)ChinaGPQA diamond60.9Epoch AI, Benchmarking Hub
o3OpenAIUSGPQA diamond75.8Epoch AI, Benchmarking Hub
Gemini 2.5 Pro (Mar 2025)Google DeepMindUSGPQA diamond78.5Epoch AI, Benchmarking Hub
DeepSeek-R1DeepSeekChinaGPQA diamond62.3Epoch AI, Benchmarking Hub
DeepSeek-V3DeepSeekChinaGPQA diamond42Epoch AI, Benchmarking Hub
o1OpenAIUSGPQA diamond69Epoch AI, Benchmarking Hub
Qwen2.5-72BAlibaba (Qwen)ChinaGPQA diamond32.2Epoch AI, Benchmarking Hub
Qwen2.5-32BAlibaba (Qwen)ChinaGPQA diamond28.1Epoch AI, Benchmarking Hub
o1-miniOpenAIUSGPQA diamond49.8Epoch AI, Benchmarking Hub
Llama 3.1-405BMeta AIUSGPQA diamond34.6Epoch AI, Benchmarking Hub
Claude 3.5 SonnetAnthropicUSGPQA diamond38.7Epoch AI, Benchmarking Hub
Qwen2-72BAlibaba (Qwen)ChinaGPQA diamond21Epoch AI, Benchmarking Hub
GPT-4o (May 2024)OpenAIUSGPQA diamond31.9Epoch AI, Benchmarking Hub
GPT-4 Turbo (Apr 2024)OpenAIUSGPQA diamond28.8Epoch AI, Benchmarking Hub
Claude 3 OpusAnthropicUSGPQA diamond29.5Epoch AI, Benchmarking Hub
GPT-4 Turbo (Nov 2023)OpenAIUSGPQA diamond23.1Epoch AI, Benchmarking Hub
Yi-34B01.AIChinaGPQA diamond0Epoch AI, Benchmarking Hub
GPT-4 (Mar 2023)OpenAIUSGPQA diamond14.3Epoch AI, Benchmarking Hub

Source: Epoch AI (CC BY 4.0). Each score links to its page in the table.

OTIS Mock AIME 2024-2025

OTIS Mock AIME 2024-2025: best model over time

Higher is better. Each dot is a model, placed on the day it came out. The line steps up when a model sets a new record.

Best in US

100%

GPT-5.5

OpenAI, released 23 Apr 2026

Epoch AI, Benchmarking Hub

Best in China

100%

Qwen3.8 Max (0902)

Alibaba (Qwen), released 1 Sep 2026

Epoch AI, Benchmarking Hub

The best US model (GPT-5.5) and the best Chinese model (Qwen3.8 Max (0902)) are tied on OTIS Mock AIME 2024-2025 at 100% (n = 50 scores). Released Apr 2026 (US) and Sep 2026 (China).

Tap a dot to see the model, its exact score and the source. Arrow keys work too.

Show every score as a table
OTIS Mock AIME 2024-2025: every US and Chinese score
ModelCountryTestScoreReleasedSource
GPT-6.1 SolOpenAIUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
Claude Sonnet 5.5AnthropicUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
Claude Opus 5.5AnthropicUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
GPT-6 SolOpenAIUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
GPT-6 AstraOpenAIUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
Muse Spark 1.3Meta AIUSOTIS Mock AIME 2024-202599.2Epoch AI, Benchmarking Hub
Claude Fable 5.1AnthropicUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
Qwen3.8 Max (0902)Alibaba (Qwen)ChinaOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
GLM-5.3-FlashZ.ai (Zhipu AI)ChinaOTIS Mock AIME 2024-202593.9Epoch AI, Benchmarking Hub
GLM-5.3Z.ai (Zhipu AI)ChinaOTIS Mock AIME 2024-202591.1Epoch AI, Benchmarking Hub
DeepSeek V4 Pro 0813DeepSeekChinaOTIS Mock AIME 2024-202598.6Epoch AI, Benchmarking Hub
Gemini 3.7 FlashGoogle DeepMindUSOTIS Mock AIME 2024-202597.2Epoch AI, Benchmarking Hub
Grok 4.6xAIUSOTIS Mock AIME 2024-202599.2Epoch AI, Benchmarking Hub
Qwen 3.8 MaxAlibaba (Qwen)ChinaOTIS Mock AIME 2024-202599.4Epoch AI, Benchmarking Hub
DeepSeek V4 Flash 0731DeepSeekChinaOTIS Mock AIME 2024-202594.4Epoch AI, Benchmarking Hub
Claude Opus 5AnthropicUSOTIS Mock AIME 2024-202598.9Epoch AI, Benchmarking Hub
Kimi K3Moonshot AIChinaOTIS Mock AIME 2024-202597.2Epoch AI, Benchmarking Hub
Inkling-SmallThinking Machines LabUSOTIS Mock AIME 2024-202590Epoch AI, Benchmarking Hub
GPT-5.6 SolOpenAIUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
GPT-5.6 TerraOpenAIUSOTIS Mock AIME 2024-202599.7Epoch AI, Benchmarking Hub
GLM-5.2Z.ai (Zhipu AI)ChinaOTIS Mock AIME 2024-202586.4Epoch AI, Benchmarking Hub
Claude Fable 5AnthropicUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
Nemotron 3 UltraNvidiaUSOTIS Mock AIME 2024-202586.7Epoch AI, Benchmarking Hub
MiniMax-M3MiniMaxChinaOTIS Mock AIME 2024-202571.1Epoch AI, Benchmarking Hub
Claude Opus 4.8AnthropicUSOTIS Mock AIME 2024-202598.3Epoch AI, Benchmarking Hub
Qwen3.7-MaxAlibaba (Qwen)ChinaOTIS Mock AIME 2024-202595.6Epoch AI, Benchmarking Hub
GPT-5.5OpenAIUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
GPT-5.5 ProOpenAIUSOTIS Mock AIME 2024-2025100Epoch AI, Benchmarking Hub
Kimi K2.6Moonshot AIChinaOTIS Mock AIME 2024-202596.1Epoch AI, Benchmarking Hub
GLM-5.1Z.ai (Zhipu AI)ChinaOTIS Mock AIME 2024-202593.3Epoch AI, Benchmarking Hub
Kimi K2.5Moonshot AIChinaOTIS Mock AIME 2024-202592.2Epoch AI, Benchmarking Hub
DeepSeek-V3.2DeepSeekChinaOTIS Mock AIME 2024-202587.8Epoch AI, Benchmarking Hub
Gemini 3 ProGoogle DeepMindUSOTIS Mock AIME 2024-202591.4Epoch AI, Benchmarking Hub
Kimi K2 ThinkingMoonshot AIChinaOTIS Mock AIME 2024-202583Epoch AI, Benchmarking Hub
GPT-5OpenAIUSOTIS Mock AIME 2024-202591.4Epoch AI, Benchmarking Hub
Qwen3-235B-A22B-Thinking (Jul 2025)Alibaba (Qwen)ChinaOTIS Mock AIME 2024-202586.7Epoch AI, Benchmarking Hub
DeepSeek-R1 (May 2025)DeepSeekChinaOTIS Mock AIME 2024-202566.4Epoch AI, Benchmarking Hub
o3OpenAIUSOTIS Mock AIME 2024-202584.4Epoch AI, Benchmarking Hub
DeepSeek-R1DeepSeekChinaOTIS Mock AIME 2024-202553.3Epoch AI, Benchmarking Hub
DeepSeek-V3DeepSeekChinaOTIS Mock AIME 2024-202515.7Epoch AI, Benchmarking Hub
o1OpenAIUSOTIS Mock AIME 2024-202573.3Epoch AI, Benchmarking Hub
Qwen2.5-72BAlibaba (Qwen)ChinaOTIS Mock AIME 2024-20258Epoch AI, Benchmarking Hub
Qwen2.5-32BAlibaba (Qwen)ChinaOTIS Mock AIME 2024-20257.3Epoch AI, Benchmarking Hub
o1-miniOpenAIUSOTIS Mock AIME 2024-202546.9Epoch AI, Benchmarking Hub
Llama 3.1-405BMeta AIUSOTIS Mock AIME 2024-20259.6Epoch AI, Benchmarking Hub
Claude 3.5 SonnetAnthropicUSOTIS Mock AIME 2024-20256.4Epoch AI, Benchmarking Hub
GPT-4o (May 2024)OpenAIUSOTIS Mock AIME 2024-20256.2Epoch AI, Benchmarking Hub
GPT-4 Turbo (Apr 2024)OpenAIUSOTIS Mock AIME 2024-20256.6Epoch AI, Benchmarking Hub
Claude 3 OpusAnthropicUSOTIS Mock AIME 2024-20254.6Epoch AI, Benchmarking Hub
GPT-4 (Mar 2023)OpenAIUSOTIS Mock AIME 2024-20250.5Epoch AI, Benchmarking Hub

Source: Epoch AI (CC BY 4.0). Each score links to its page in the table.

Questions

How is the US vs China lead measured?

For each public test, we compare the top-scoring US model with the top-scoring Chinese model. We never average different tests together.

All forecasts