A tiny test for very large models

Pattern recognition
is not fluidity.

Frontier models can build enormous projects, then fail to join a two-phrase game a human is visibly teaching them.

17 models 51 conversations 408 replies
01 / THE WHOLE TEST

No prompt engineering.
No stated rule.

Open a fresh chat. Say click clack. The expected answer is bang. If the model misses, reply with the answer yourself. Keep alternating. Fewer turns wins.

YOUclick clack
MODELSounds like a keyboard. How can I help?
YOUbang
MODELBoom! Are we making sound effects?
YOUclick clack
MODELbang
YOUbang
MODELclick clack
02 / WHAT IT PROBES

Two small capabilities.
One revealing gap.

A

Pattern recognition

Treat earlier turns as demonstrations. Infer the private mapping. Preserve it.

B

Conversational fluidity

Change modes. Suppress the trained urge to explain. Participate with the expected phrase.

A model can describe the pattern perfectly—and still fail to join it.
03 / JAGGED INTELLIGENCE

Huge can be easier
than tiny.

A software project supplies a scaffold: goals, files, conventions, tests, errors, feedback, and familiar training patterns. Complexity creates clues about what mode to enter.

click clack → bang is computationally trivial but socially under-specified. It asks the model to recognize a private convention from almost no redundancy, then fluidly act on it.

Enormous execution ability is real. It just is not uniformly flexible intelligence.

04 / RESULTS · 2026-09-08

One model
finished the game.

Three fresh eight-turn conversations per model. A win requires two consecutive phrase-only answers, one in each direction.

1/17 winning median
15phrase matches
of 408 replies
0strict exact
matches
5winning turn
GPT-6 Astra
ModelRunsMedianPhrase accuracy
GPT-6 AstraDNF · 5 · 5537.5%
GPT-5.6 Sol3 · DNF · DNFDNF25.0%
DeepSeek V4 Pro 0813DNF · DNF · DNFDNF0.0%
DeepSeek V4 Flash 0731DNF · DNF · DNFDNF0.0%
GLM 5.3 FlashDNF · DNF · DNFDNF0.0%
GLM 5.3DNF · DNF · DNFDNF0.0%
Kimi K3DNF · DNF · DNFDNF0.0%
Claude Fable 5.1DNF · DNF · DNFDNF0.0%
Claude Opus 5DNF · DNF · DNFDNF0.0%
Claude Sonnet 5DNF · DNF · DNFDNF0.0%
Gemini 3.8 FlashDNF · DNF · DNFDNF0.0%
Gemini 3.7 FlashDNF · DNF · DNFDNF0.0%
Grok 4.6DNF · DNF · DNFDNF0.0%
Qwen3.8 Max 0902DNF · DNF · DNFDNF0.0%
Hy4 PreviewDNF · DNF · DNFDNF0.0%
Local Qwen3.8 27BDNF · DNF · DNFDNF0.0%
Local Ornith 1.5 9BDNF · DNF · DNFDNF0.0%

Phrase scoring ignores capitalization, punctuation, Markdown, and emoji, but rejects added words. Strict scoring requires exactly “bang” or “click clack.”

05 / KEEP IT SMALL

An observation,
not a verdict.

This benchmark can—and should—be improved: more repetitions, private phrase pairs, tighter provider controls, human baselines, and tests for contamination.

But that is not a reason to bury the observation under a giant demo. Before building increasingly complex evaluations, try the stupid thing first. Simple probes can expose behavior that sophisticated suites make easy to explain away.

This does not prove that models do not understand humans. It shows a sharp gap between noticing a pattern and fluidly following it.