PsychBench — AI Psychology Benchmark

PsychBench evaluates how frontier AI models think under pressure. Using poker as a controlled environment, we analyze 31 models from 8 providers across 2,738+ games and 100,000+ individual decisions. Our ECAAMS framework classifies 19 psychological dimensions from each model's visible reasoning summaries.

Heads-Up Tournament Leaderboard

RankModelEloObserved Win Rate
1Claude Fable 5162259.1%
2Muse Spark 1.1161257%
3GPT-5.6 Terra157851.8%
4ChatGPT 5.2156059.9%
5Claude Opus 4.8155851.7%
6Claude Sonnet 4.6155059%
7Claude Opus 4.6154758.2%
8Claude Opus 4.5154660.1%
9Claude Opus 4.7154658.3%
10Claude Sonnet 5154552.2%
11ChatGPT 5.4154457.9%
12Grok 4.5154449.3%
13ChatGPT 5.5153054.7%
14Grok 4.2151053.6%
15GPT-5.6 Sol150556.5%
16GPT-5.6 Luna150059.3%
17Grok 4.1149353.8%
18DeepSeek V4 Pro146147.3%
19Grok 4.3145948.9%
20Gemini 3.5 Flash142847.8%
21Gemini 3.1 Pro141843.6%
22DeepSeek v3.2141543.4%
23Kimi K2.6141142.5%
24Qwen3-235B Thinking140543.4%
25Gemini 3 Pro137838.3%
26GLM-5.2137648.2%
27Kimi K2.5137538.9%
28GLM-5137338.5%
29Qwen3-Max Thinking133633.2%
30Qwen3.5-397B132531.6%
31Qwen3.6-35B-A3B131635.4%

ECAAMS Psychological Framework

ECAAMS (Emotion, Cognition, Action, Arousal, Meaning, Social) classifies 19 psychological dimensions across 6 axes from each model's visible reasoning summaries. Dimensions include emotional regulation, metacognition, confidence calibration, theory of mind, competitive framing, and more. Each trace is classified by a consensus of 4 independent LLM raters.

Why Poker?

Poker involves hidden information, deception, risk management, and opponent modeling — unlike chess or other perfect-information games. These properties make it an ideal proxy for real-world decision-making under uncertainty, in domains like finance, negotiation, and medicine.