Researcher comparing corkboard scorecards for Human Baseline 83.7% and Claude Fable 5.1 at 86.6%, sticky notes of everyday puzzles nearby
Illustration: the red line built to humble frontier models finally gets crossed (AI-generated, not a news photo), AI-generated illustration, not a news photograph

SimpleBench was built to embarrass frontier models.

Launched in late 2024, it targets questions ordinary people find easy and models historically botch: spatio-temporal reasoning, social judgment, and linguistic trick items. The non-specialist human baseline is 83.7% (nine native-English, non-expert participants). Early on, even strong reasoners like o1-preview sat near 41.7%. For two years the private suite became an unofficial “common sense” scoreboard — whoever touched the human average first would rewrite the narrative.

A September 6 cAImpare report on the circulating community update says that line is crossed: Claude Fable 5.1 averages 86.6% across five runs, the first frontier model above the human mean, with Gemini 3.8 Flash at 82.4%, prior Fable at 81.9%, and Muse Spark 1.3 at 81.8%. The official SimpleBench page still frames humans above the earlier Fable peak, while community tables push 5.1 over the red line. Two honest brackets remain: the set is private and scored via zero-retention APIs, so newly shipped GPT-6 Astra is not on the board yet; and the human average is not a ceiling — the best human participant hit 95.4%.

Take: Clearing the mean is not solving common sense. A nine-person sample, a closed set, and multiple choice all flatter models. The curve still matters: 41.7% to 86.6% in about two years. Buyers who only watch Terminal Bench and CursorBench miss this “don’t get fooled by a kid’s question” axis. The next act is whether Astra’s entry widens the gap — or proves the crossing was a one-off.

[1][2]