The limit lies outside the model.
Power grids, cooling and efficiency help determine how far AI can actually scale.
When people talk about progress in AI, we almost automatically look at models: more parameters, larger context windows, better benchmarks. To me, that view is incomplete. Every model runs on very concrete infrastructure. It needs chips, electricity, grid connections and a cooling system capable of reliably removing the heat it produces. Depending on location and design, water can also play a role.
Compute is not an abstract resource
A new data centre cannot be scaled through software alone. Grid connections, substations, transformers and permits take time. Even when enough electricity is being generated, it must be available in the right place at the right time. Power supply therefore becomes part of the AI architecture, just like storage, networking and accelerators.
This does not mean that every larger model is necessarily the wrong direction. It means that progress should not be measured by model quality alone. What also matters is how much useful performance a system delivers per watt, per request and across its full lifecycle.
Efficiency is a product decision
I am therefore increasingly interested in the whole chain: smaller specialised models, quantisation, better utilisation, caching and the question of whether a request needs a large model at all. A data centre's location and cooling design are not details outside the product. They help determine how resilient and scalable it can be.
AI progress is not determined by algorithms alone. The next limit may be a transformer, a cooling plant or a missing grid connection. When building AI systems, we should optimise not only the intelligence inside the model but also account for the physical reality around it.