10 Comments
User's avatar
Louis Hemming's avatar

The economic re-definition of AGI makes a kind of sad sense next to your crystallized vs fluid framing. Crystallized is where the revenue is, so that's the half that gets to be called intelligence. A model that can train other models but cant figure out parking a game piece is apparently close enough for the press release.

Alberto Romero's avatar

Indeed. But my hunch is that crystallized intelligence gets you only so far. They're compensating for the lack of fluid smarts with extra tokens and the enterprise sector is catching up on that. They will start to ask whether those tokens actually bring revenue. I think Ant and OAI know this and are working to find better solutions. Another question is whether they will find them.

Understanding Intelligence's avatar

The goal of ARC-AGI is not to prove that a model has human-level intelligence, but to refute that it is the case. The philosophy of the test is Popper's falsification approach. Should GPT reach 100%, it will still be irrelevant. These benchmarks are saturated always in the same way: manually creating large datasets with games similar in spirit to those in ARC-AGI 3, so that the model will be prepared and optimized for the test. But intelligence consists precisely in facing completely novel situations: saturating the benchmark by accumulating experience defeats the purpose.

Keith Minler's avatar

Fascinating, instructive and well worth the read. I am wondering if is it possible for humans to try ARC-AGI-3 type tests?

Alberto Romero's avatar

Thank you, Keith. Sure, try here: https://arcprize.org/arc-agi/3 (note that this is the public set, which means the games are easier than those GPT-5.6 got a 7.8% on. It won't seem easy at first, but keep trying)

Marcus Seldon's avatar

I wonder a lot about hill climbing and overfitting to benchmarks though. The ARC tests are supposed to measure fluid intelligence, yet qualitatively even cutting edge LLM fluid intelligence feels a lot worse in real world scenarios than, say, those ARC 2 results would suggest.

Alberto Romero's avatar

Certainly. ARC-AGI-2 is still a much easier challenge than life

James Maconochie's avatar

Alberto, the reframe is the service here. Everyone will laugh at 7.8% or cheer the 20x jump; you did the harder thing and asked why this benchmark, of all of them, keeps doing this. And I think your own Ontario example contains the answer.

The driver handles the shape in the road, not because he carries a better world model in the abstract. He handles it because every novel situation he has ever faced cost him something, the near-miss, the skid, the fear that rewired his attention. His capability profile was smoothed as a consequence. That is what fluid intelligence is: what a mind looks like after errors have landed on it, over and over, and it has to carry them forward.

Which is why the model's profile is jagged, research mathematics on one side, 1-in-13 on children's games on the other. Nothing has ever landed on it. And it is why I'd bet the ARC series keeps resetting to near-zero: each version can eventually be trained through, but novelty under the stakes is precisely what training cannot pre-supply. You cannot memorize your way into orientation.

You played the demo yourself. So here is my question: when you finally got the hang of a game, what did your own wrong moves do to you that GPT-5.6's wrong moves do not do to it?

Georgi Kamov's avatar

The exact skill that breaks Sol, composing a coherent plan as the chain of inference gets deep, is the literal mechanism collective decision-making runs on. A room full of stakeholders is nothing but a long chain of inference with competing inputs that has to resolve into one plan. So the industry just spent a hundred million dollars of compute discovering that this is the hard part, at the scale of one mind. Nobody has asked what it costs at the scale of five. More here https://georgikamov.substack.com/p/the-adolescence-of-humanity