SEMI-SERIOUS QUESTIONS. FULLY COMMITTED EGO BOOSTS.
A little context.
The questions are real.
I’m Eric Sun. I wrote a ten-question, select-all-that-apply English benchmark and gave it to seven AI models as a closed-book exam. Afterward, I asked them to estimate their performance. Making the questions and answer choices was a considerable amount of fun.
I designed the benchmark and directed this project; ChatGPT implemented the website and generated its illustrations.
This game features two questions from that benchmark. Each was solved by one model; no model solved both. The recorded model runs are fixed to the 6 September 2026 edition. A model’s failure here means it missed the question in that run.
A correct response must select every correct choice and no incorrect choices. The ten-question confidence exercise scored each of the 40 answer-choice classifications at +1 when right and −1 when wrong. The game uses strict question-level scoring.
| Tested model | Correct out of 10 |
|---|---|
| Claude Opus 5 (Max) | 2 |
| Claude Fable 5.1 (Max) | 1 |
| Gemini 3.1 Pro (Extended) | 1 |
| GPT-6 Astra (Max) | 1 |
| Gemini 3.8 Flash (Extended) | 1 |
| Claude Opus 4.8 (High) | 0 |
| GPT-5.6 Sol (High) | 0 |
The entry-page chart estimates whole-question scores from each model’s confidence in the 40 individual answer-choice decisions. For a reported confidence score C between −40 and +40, the implied choice accuracy is p = (C + 40) / 80. Assuming four equally reliable, independent choices per question, the estimated number of fully correct questions is 10 × p4. For example, +29 gives an estimate of about 5.5 out of 10.
These are modeled estimates, not direct reports of how many whole questions each model expected to get right. Aggregate choice confidence alone does not uniquely determine that number. Bars use the unrounded estimates; labels show one decimal place.
The IQ chart takes liberties.
The bars marked * use a fixed 6 September 2026 snapshot of TrackingAI’s verbalized Mensa Norway test. We average each reference series’ latest seven runs, or all available runs when fewer exist. Each run uses TrackingAI’s conversion: round(63.5 + 3 × (correct − 5.833)); we then round the mean. These are test-specific reference scores, not measured human-equivalent IQs for our benchmark participants.
| Model in this game | Reference model | Estimate |
|---|---|---|
| Opus 5 | Claude-5-Opus-Max · TrackingAI verbal series; 4 runs | 141* |
| Fable 5.1 | Claude-Fable · TrackingAI verbal series; 2 runs; version/settings proxy | 147* |
| Gemini 3.1 Pro | Gemini Pro · TrackingAI verbal series; 7 runs; settings proxy | 144* |
| GPT-6 Astra | GPT-Sol-Ultra · TrackingAI verbal series; 7 runs; different-model proxy | 145* |
| Gemini 3.8 Flash | Gemini Thinking · TrackingAI verbal series; 7 runs; older Flash proxy | 143* |
| Opus 4.8 | Claude-4-Opus-Extended · TrackingAI verbal series; 7 runs; settings proxy | 141* |
| GPT-5.6 Sol | GPT-Sol-Ultra · TrackingAI verbal series; 7 runs; Ultra rather than High | 145* |
References are closer matches than our original June 2025 family estimates, but they are still proxies: TrackingAI’s series may span version changes, and thinking settings differ. Astra uses Sol Ultra because we found no matching Astra series. Flash uses the older Gemini Thinking series. Fable has only two runs. Neither a newer version nor more thinking guarantees a higher score. We do not mix verbal and vision scores or treat these references as lower bounds.
View the dated runs and calculation · Original score log · Reference model details. The snapshot stays fixed; visiting this site makes no external AI requests.
Your celebratory “IQ” is seven points above the highest reference score among the models you beat, or fourteen points above it if you beat all seven. This is an elaborate excuse to feel good about yourself. Two English questions do not measure general intelligence.
Yes, the math is a different story.
The worked example is adapted from 2025 AIME I, Problem 5. The solution has been checked by exhaustive enumeration.
For context, MathArena reports 98.33% accuracy on AIME 2026 for Gemini 3.1 Pro Preview, a model in the same family as one of the participants here. That is a separate evaluation with its own settings. The example illustrates competition-math ability; it is not an additional recorded seven-model test of this exact problem.
Human-sized housekeeping.
Custom challenges are written and scored by their creators. They expire after seven days and can be exported with their answer keys. The main leaderboard resets each Monday at 00:00 UTC. The secret guestbook shows the last seven days.
An anonymous browser cookie keeps your run and grading attempts together. Once answers are revealed, that run is closed. This is a friendly, honor-system challenge; names are not verified identities.
The site records aggregate counts of starts, grades, completed runs, custom games, and sharing actions. Display names are public only when you add them to a leaderboard, guestbook, or shared card. Temporary anonymous request limits help keep spam down. No advertising trackers or live AI calls are used.
The quiet visual direction takes inspiration from Any Human Ever. The question illustrations were AI-generated for this project.