I Expect Rapid Progress — But Not Towards General Superintelligence
The next few years of AI progress will come from engineering acceleration—not from models becoming dramatically different in nature.
The core claim is that researchers are seeing massive acceleration in the infra and engineering capabilities of models, and that this will make models superhuman distributed GPU engineers within a few years. That will make experimentation and tinkering with model formulations far easier, but it will not change what the models fundamentally are. The author frames this as a shift back toward research being the bottleneck, after a deep-learning era in which researchers were judged by their ability to implement and scale ideas in complex infrastructure.
Much of the near-term progress comes from scaling inference-time compute with current tools rather than step-changes in model capability. Training metrics like tokens per second per GPU and inference metrics like FLOPs per token or cost per answer are highly verifiable and optimizable. The author expects AI agents to help optimize this process end-to-end in a few years, bringing inference capabilities close to the underlying maximum compute available on accelerators like GPUs. Companies have already captured 10-30% savings on cost-to-serve after announcing a model at a given price point. The effective cost of model intelligence is expected to decline near-exponentially, potentially faster than recent trends. A prediction is that pretraining research—at least in architecture and data selection for the current class of models—will be automated in 2-3 years.
This efficiency capture should take only a few years. Longer-term co-design of accelerators and models will add extra orders of magnitude on top of the flexible GPU platform. The author also flags RL environment quality as industrial-scale low-hanging fruit: multiple RL data companies have crossed $100M or $1B in revenue, yet the average output is remarkably low-quality, with countless researchers describing much of what they buy as "frankly crap." Leading labs still see clear ROI from buying the data, and the cruddy aspects are clearly fixable.
A Jevons paradox is expected for agentic models: as efficiency improves, demand will only increase, with the industry bottlenecked on orienting and delivering agents. Meta's Muse agent is cited as an early indicator, and more Muse-like experiences for different audiences and use-cases are expected. The value will come from understanding how agents work rather than pushing frontier performance.