The most interesting question about the AI investment boom is not whether it is a bubble. It is whether the physical architecture on which the boom is being built will remain economically relevant for as long as the assets being financed.
Hundreds of billions of dollars are flowing into accelerators, power generation, transmission, cooling, buildings, chip fabrication, and grid connections. The International Energy Agency expects global data-center electricity consumption to more than double by 2030, to roughly 945 TWh, with AI the largest driver. In the United States, data centers could account for almost half of electricity-demand growth through the end of the decade.
These forecasts may be reasonable. But they are not forecasts of demand alone. Embedded inside them is a second forecast that is discussed much less often: that the present relationship between useful computation, hardware, and energy will persist long enough to justify the infrastructure being built around it.
That is a much more fragile assumption.
The hidden constant in the forecast
Today’s capex cycle implicitly relies on several things remaining sufficiently stable: large quantities of conventional accelerator compute, high energy use per useful result, centralized data centers, costly movement of data between processors and memory, and continuing demand for ever more hardware.
Once those assumptions are translated into steel and concrete, they become long-lived commitments. A grid connection, substation, cooling plant, or data-center shell is not purchased like a software subscription. Its economics depend on years—often decades—of utilization.
Yet the extraordinary resource cost of the current paradigm creates an equally extraordinary selection pressure against it.
If a company can deliver the same relevant inference outcome with one tenth of the energy and hardware, that is not a marginal feature. At current scale it changes total cost of ownership, cooling, power procurement, floor space, deployment speed, and capital intensity. A hundredfold improvement would not merely make existing facilities cheaper to operate. It could make parts of today’s infrastructure pipeline unnecessary.
The key phrase is relevant outcome. Users do not buy floating-point operations. They buy resolved support cases, completed designs, useful diagnoses, better decisions, and functioning software. Any method that reduces the computation required for those outcomes attacks the denominator on which the physical build-out rests. The same distinction between owning capacity and designing for the actual task also appears in decisions about large versus deliberately managed model context .
There is no need to wait for science fiction
The popular conversation tends to imagine only two states: today’s GPU-centered architecture, or a remote post-von-Neumann revolution. In reality, there is a ladder of possible improvements:
- Better models and algorithms can require fewer operations per task.
- Quantization, sparsity, and distillation can reduce memory and compute requirements.
- Better accelerators can lower joules per operation.
- Compute-in-memory and analog architectures can reduce the cost of moving data.
- Neuromorphic and event-driven systems can produce radically different load profiles.
- Photonic, spintronic, or other unconventional systems may eventually move parts of AI beyond the present architecture altogether.
These are not equally mature, and none should be treated as an inevitable winner. But they are not imaginary either. IBM researchers have demonstrated an analog in-memory chip for speech recognition and transcription that performs matrix-vector operations directly on memory tiles, reporting sustained performance of up to 12.4 TOPS per watt. Recent neuromorphic research has reported circuit-level energy-efficiency improvements of two orders of magnitude over a conventional comparison platform, while also stressing how far the field still is from broadly deployable, brain-like computing.
That mixture—real prototypes, difficult engineering, and uncertain commercialization—may explain why alternative architectures receive so little attention in macro discussions. Infrastructure analysts can count planned megawatts. They cannot easily assign a date or probability to a discontinuous efficiency gain. The measurable pipeline therefore dominates the narrative, while the technologies capable of invalidating it are relegated to the footnotes.
The productivity premise is not yet settled
There is another uncertainty underneath the capex story. Even within the current LLM paradigm, the aggregate productivity payoff remains unclear.
The evidence is not simply negative. A large customer-support study found that AI assistance increased issues resolved per hour by 13.8%, with the largest gains among less experienced workers. But a field experiment involving 7,137 knowledge workers across 66 firms found two fewer hours per week spent on email among active users without detecting broader changes in the quantity or composition of their tasks. A randomized study of experienced open-source developers using early-2025 tools found that they took 19% longer on eligible tasks—a result the authors explicitly framed as a snapshot of one setting, not a universal verdict.
This is exactly what an honest reading should produce: not “LLMs do not raise productivity,” but “the size, distribution, and durability of the productivity gain are not yet known.” A localized time saving does not automatically become firm-level output, economy-wide productivity, or a return sufficient to support every layer of the infrastructure stack.
The investment thesis therefore contains two coupled bets: that demand for AI services will expand dramatically, and that today’s capital-intensive method of satisfying that demand will remain competitive.
The capital attractor
Here the boom creates its own adversary.
When the United States and China direct vast amounts of industrial capacity toward similar compute infrastructure, they do more than intensify an AI race. Together they place an enormous global prize on discovering a cheaper computational paradigm.
Every expensive accelerator cluster enlarges the market for an alternative. Every delayed grid connection raises the value of efficiency. Every cooling constraint, transformer shortage, and financing bill attracts researchers, start-ups, and capital toward ways of doing more with less.
This creates the central paradox:
The larger the data-center boom becomes, the stronger the incentive to destroy the technological basis of the data-center boom.
That does not mean efficiency will automatically reduce total energy use. Lower costs can stimulate more consumption—the familiar rebound effect—and the IEA’s base case already incorporates continuing efficiency improvements while still projecting strong electricity growth. Nor does it mean new architectures will replace conventional accelerators everywhere. General-purpose ecosystems, manufacturing scale, and software compatibility are formidable advantages.
It means something narrower and more consequential for investors and planners: linear forecasts deserve a technology-risk discount.
“AI data centers will require X gigawatts in 2035” is not merely an extrapolation of demand. It is an extrapolation of a particular level of technological inefficiency across ten years. History offers few reasons to treat that inefficiency as a constant—especially when its removal has become one of the most valuable engineering prizes on Earth.
The ultimate winner of the AI capex boom may therefore be neither the owner of the most data centers nor the supplier of the most accelerators. It may be the technology that makes a meaningful share of both unnecessary.
