Kaiming He's Team Finds Top AI Models Were Taking the AGI Test Blind — and Claude Scores 100 With Its Sight Restored
A paper posted over the Golden Week holiday by Kaiming He's group at MIT argues the world's best models have been sitting the hardest AGI exam effectively blindfolded, fed a 64×64 grid of numbers instead of the game screen. Restore the image and add a searchable visual memory, and Claude went from 40 to a perfect 100.

The most-discussed AI paper of China's Golden Week did not come out of a lab in Shenzhen or a boardroom in San Francisco — it appeared on arXiv over the holiday from the MIT group of Kaiming He, the computer scientist behind ResNet and one of the most-cited researchers in the field, and by October 4 it had taken the number-one slot on 36kr's trending board, with AI engineers passing it around X. Its claim is embarrassingly simple: for the past half year, the world's leading models have been taking the hardest AGI test blindfolded — and when you take the blindfold off, the scores transform.
The test is ARC-AGI-3, the benchmark François Chollet, creator of the Keras library, released in March: twenty-odd puzzle games with no instructions, where the only way to learn the rules is to try things and watch what happens. But the official harness feeds a model the game as a 64×64 table of numbers — roughly 4,000 tokens per frame — with no picture at all. The VISTA paper's authors compare it to playing Super Mario while someone reads coordinates aloud. On that diet, Claude Opus 5 managed a score of just 40.
VISTA makes two changes and only two. First, it hands the model its eyes: the number grid becomes a proper 512×512 image of the screen. Second, it gives the model a lossless visual memory — an "album" in which every frame is archived at full fidelity, to be pulled up, compared side by side, zoomed into, and read for exact colours, with no cost in game moves. Nothing else changes. With sight alone, GPT-5.6 Sol jumped from 13.33 to 47.32 while using less than half the tokens. With the album on top, Claude went from 40 to a perfect 100 across all 25 public games, and cleared games it had never seen using 57.4 percent fewer actions than a first-time human player.

The paper's most quoted detail is almost anthropological. On one level, the model clicked an orange block, stopped, retrieved the four surrounding frames, laid them side by side, worked out what had changed, and only then moved on — 34 steps where a human newcomer needed 44. The contrast with this summer's leaderboard approaches is pointed: those systems, with names like Tycho and Schema, had models brute-force the games by writing code — one checkers-like game took 4,000 lines of Python to simulate. The VISTA model got by on three notes written in plain language. And the usual scaling levers did nothing: with VISTA running, enlarging the context window from 200,000 to 780,000 tokens dropped the score from 99 to 93.9, and blowing the images up sixteen-fold cut it to 88.3 — more input, worse play.
For an outside reader, the finding lands less as a capability story than as a measurement one. The industry's frontier labs and its benchmark designers are arguing, in effect, about what an exam should measure: whether a model that aces a game it can only hear described has demonstrated reasoning, or merely learned to navigate a degraded signal that no human player would accept. The eighteen unseen game environments that circulated with the Chinese posts look like children's puzzles; they remain unsolved by systems that can pass bar exams. Whether ARC's official leaderboard adopts image-based input may now decide which number was ever the real score — 40, or 100.