Calling the progress of Large Language Models (LLMs) in recent years “rapid” is a brutal understatement. In the span of two years, models that couldn’t complete our Physics homework have ascended to produce a solution to a Millennium Prize Problem.

Some attribute this progress to scale—the amount of data, compute, and model capacity behind these systems increasing by orders of magnitude. Fundamentally, we agree. But scale doesn’t need to be solved. The true problem, the inevitable one, is building the ideal architecture: one that lays the foundation not only for scaling to occur, but for that scaling to have a meaningful effect on how humanity solves its problems.

If you want a more in-depth technical explanation of this part, check elsewhere. For the purposes of this short essay, and for the interests of the intended audience, I’ll write mainly in layman’s terms. Large language models have seen an indescribable amount of progress in the span of a few years. But why them? Why hasn’t scaling produced the same effect in robotics, world models, and other neural networks?

In language, the combination of an architecture built to scale and a relatively simple training objective—next-token prediction—has achieved astounding results. We want to know where we can go from here. Can these principles transfer laterally across domains? Can we train robots the way we train LLMs? How will this system interact with the messy, ambiguous, real world? We have problems, for sure.

And these problems are inevitable. As soon as Google opened the door to the Transformer, and OpenAI opened the door to the modern LLM, these problems arose as a function of human knowledge being incomplete. But anything, anything obeying the laws of physics is possible. And therein lies the solution to such problems.

Of course, our larger mission is more than just “solving problems.” Leaving it there makes us theoretically indistinguishable from any group of people with an idea. You can think of this instead as our codified technical philosophy: what motivates Matthew and me to push our limits and explore a domain we know nothing about, with experience in a domain we know, quite honestly, little about, in an attempt to break a boundary we found.

What we ultimately want is for the progress of generative AI, and the service that it brings humanity, to be embodied in the real world. We want to see the same democratization of information and ability in the physical world that exists in the virtual. We want many things that I can describe and make sound good and altruistic, but most of all we want to reach the reality where we show it can happen, not tell you it will.

— Aryan