The AI space is really volatile, and we keep seeing technological improvements under the hood. AI also just keeps getting more capable at every size of model. Most people pay the most attention to the huge models: the biggest and most capable. But there are also the smaller AI models that are feasible to run on consumer hardware: these also keep seeing significant improvements. They’re a lot less capable than the big ones, but they’re improving fast.
It seems to me like a lot of the research on AI is still finding its footing, and that we will keep seeing improvements for a while longer, even if those improvements are mostly in efficiency rather than capability.
Eventually, the technology will reach a plateau, where any real improvements to capability would require a fundamental shift, and there is some early-days work being done there. LLMs are the peak of general-purpose AI currently, though. I think we’re not that far off from reaching the limits of what this technology is capable of.
When we reach that point, then the biggest problem with running an LLM becomes more important: they’re just so darn big and expensive to run. Even the smaller models can push consumer hardware to the limit.
Eventually, the race for capability will turn into a race for efficiency, because more efficiency means higher profit margins. There’s already been work done on more specialized hardware, like Google’s tensor processing units. I think that some companies will eventually bake a specific model into silicon: a chip specifically designed to run *one* specific model, with silicon-level specialization. It’d be an application-specific integrated circuit (ASIC) for one particular AI model. This could mean companies putting a specific LLM into future consumer hardware, for a hefty price of course. This would mean LLM inference at a much lower energy cost, but a huge amount of flexibility and user control would be lost in the process.
This isn’t a value judgement in favor of or against LLMs or AI in general. There is good being done by AI, but there’s also a lot of harm. But it exists, and we can’t un-exist this technology, even once the bubble pops. It’ll keep existing.
Evidence: Dedicated ASICs are things we’ve seen before. Crypto mining wound up using ASICs when GPUs just weren’t cutting it anymore, and I don’t think AI is any different. This would be possible today, but it doesn’t make any sense with how fast the landscape is changing. By the time a model-specific ASIC could be designed, manufactured, and integrated, there’d be some newer model that’s more capable. That’s why I think we won’t see anything like this until capability plateaus.
Date: The date is really hard to guess, as AI research is still very active. We keep finding better ways to design/train models for capability and efficiency, and none of this dedicated silicon makes any sense until that research slows down a lot. I’d say within the next 10 years, though.