Local Models, Frontier Models, and the Demand for Intelligence
I’ve been thinking about how local AI models and frontier models might coexist as both improve. We are building an extraordinary amount of centralized AI infrastructure at the same time that laptops, workstations, phones, and other edge devices are becoming much more capable. Those trends can fit together in several ways. I can think of three broad scenarios, and they lead to different conclusions about how much centralized compute we will need.
The first scenario is that local models never become good enough for most useful work. Small models improve, but they remain too limited for serious applications. Consumers and businesses see little reason to buy machines with substantial local AI capability, so most meaningful inference continues to happen in large data centers.
I think this is the least likely outcome. I can already do a surprising amount of useful work with a local model on a laptop and it is a couple of years old. Models, hardware, and software are all improving quickly, and a great deal of work simply does not require the smartest model available. Local models do not have to match frontier models to be extremely useful. They only have to become good enough for a growing class of tasks.
The second scenario is that we build far more centralized AI capacity than we need. Frontier models could stop improving rapidly enough to keep creating new demand while local models absorb more and more everyday work. The data-center buildout could then get ahead of actual demand and create something that looks a lot like the dark-fiber era after the telecom boom. We would end up with tremendous infrastructure built for demand that arrives much more slowly than expected, and prices would fall sharply until new uses emerged and demand finally caught up.
I think this is possible. Technology markets have overbuilt before, especially when many participants build in parallel. AI infrastructure could turn out to be the same.
The third scenario is that intelligence creates so much demand that we consume essentially all the compute we can produce. This seems most likely to me and aligns with Jevons’ paradox.
Local models get dramatically better. Frontier models get dramatically better. Hardware gets faster and cheaper. Inference prices fall. None of that reduces total demand because we are also learning how to get far more value from AI.
That last part may matter as much as the technical improvements themselves. We are still early in learning what to do with these systems. Companies are redesigning workflows around them. Developers are learning when to use models, when to use deterministic software, when to use tools, and when to combine all three. Individuals are discovering places where AI can save them time or let them do things they could not reasonably do before. Every improvement in our ability to use intelligence creates another reason to consume it.
Cheap local intelligence may even increase demand for frontier intelligence rather than reduce it. A local system that helps with hundreds of ordinary tasks can also recognize the handful of cases where a much stronger model has real value. More AI in more places creates more opportunities to go after the really difficult work. Local models can remove enormous amounts of routine load from centralized systems while simultaneously expanding the total market for intelligence.
The resulting architecture may look less like a choice between local models and giant data centers but more like a hierarchy of compute. Ordinary software handles the things ordinary software does well. Small local models handle routine AI work. Larger local models tackle harder problems. Systems burst to frontier models when the additional capability justifies the cost. Massive data centers concentrate on the problems that genuinely require massive amounts of compute.
Frontier inference could become dramatically cheaper in that world and still consume all the data-center capacity we can build. Prices might fall by 10x or even 100x while total demand rises rapidly. Cheaper intelligence would make existing uses less expensive, but it would also make entirely new uses economical.
Bandwidth followed a path like this. The cost of moving data fell enormously and networks became vastly more efficient, yet total bandwidth demand exploded because we kept inventing things to do with cheap bandwidth. Streaming video, cloud computing, video conferencing, and countless other applications that did not make economic sense at earlier prices.
I suspect intelligence may work the same way. Better models lower the cost of solving existing problems, while better tools and better understanding teach us how many more problems we can solve. The supply curve moves, but the demand curve moves too.
My guess is that the future is not local AI or frontier AI. It is both. Local machines will do an enormous amount of work, frontier systems will do things that local machines cannot, and increasingly sophisticated software will decide where each piece of work belongs.
The biggest uncertainty may not be how much compute AI requires. It may be how many valuable uses for abundant intelligence we have yet to discover.
