The Other Half of Compute
xAI stood up its first 100,000 GPUs in Memphis in 122 days. It doubled that in another 92. By early 2026 the site, Colossus, held around 555,000 of them, building toward two gigawatts of power, for a reported 18 billion dollars. 1
Two sophisticated people can look at that number and reach opposite conclusions.
Jensen Huang’s view is that the only real risk is underspending. He puts the buildout at a trillion dollars and counting, and argues the company that holds back capacity loses the decade. 2 Dario Amodei and Ray Dalio sit on the other side. Amodei has said it can be rational not to buy unlimited compute, because the revenue to justify it may arrive on a timeline that bankrupts whoever guessed wrong. Dalio keeps making a narrower point: a technology can succeed completely and still ruin the people who financed it. 3
Same buildout. Same dollar figure. One camp calls it the obvious move of the decade and the other calls it the setup for a wipeout. They are not disagreeing about the facts. They are reading the same number and the number is the problem.
What 18 billion dollars buys
Every token a model produces runs down a physical path. Electricity has to be generated, moved across a grid, and stepped down through transformers to a voltage a data centre can use. Chips have to be fabricated at advanced nodes, which in practice means TSMC and a single supplier of the lithography machines that make the process possible. The chips have to be wired together with optical interconnect, assembled into racks, and kept cold. None of those layers move at the same speed, and the slowest one always sets the schedule.
For four years the slowest layer kept changing. In 2022 the constraint was GPUs themselves. In 2023 it was the high-bandwidth memory stacked next to them. In 2024 it was the advanced packaging that bonds the two together. By 2025 it was photonics, the lasers and transceivers that move data between racks. By 2026 it had reached power and the grid, where a new high-voltage connection can take longer to approve than the cluster takes to build. Bringing a large new source of power onto that grid now takes a median of more than four years. 4
Each layer is real, each one becomes scarce in turn, and the scarcity moves to the next layer as the one before it gets solved.
Call it the capacity stack. It decides one thing: how much raw compute can physically exist. It tells you what you can run. It says nothing about how much useful work comes out the other end.
The binding constraint has moved through the stack for four years straight. Chips, memory, packaging, photonics, power. Each one stayed invisible until the one before it was solved.
The number that never makes the capex debate
Now look at a different figure.
In March 2023, running a million tokens through GPT-4 cost about 30 dollars. By the middle of 2024, the same class of capability through GPT-4o cost 2.50 dollars. By 2025 a GPT-4-grade model was available at roughly 10 cents per million tokens. 5 For the rougher GPT-3.5 tier the price fell from 20 dollars per million tokens to about 7 cents in two years, a drop of more than 250 times. Epoch AI, which tracks this carefully, finds inference prices falling somewhere between 10 and 50 times a year depending on the task. 6
Almost none of that came from adding watts. The capacity stack was straining the entire time. The cost of intelligence fell by two orders of magnitude anyway. These are list prices, so some of the fall is competition between providers, but most of it is a second stack that lives inside the software layer and does work the hardware never sees.
That second stack has its own layers. At the bottom is the attention kernel. The 2022 FlashAttention paper showed that a transformer was bound by memory traffic, the data shuttling between the fast and slow memory on the chip, and that rewriting the kernel to respect that traffic multiplied throughput without changing a single transistor. 7 Above it sits serving. Key-value caching, which means storing a conversation’s intermediate state instead of recomputing it on every new token, turned long contexts from a quadratic expense into something a business could afford to offer. Above that sits the model itself. Mixture-of-experts routing, the design behind Switch Transformers, broke the link between a model’s total size and the compute each token triggers, so a model can hold a trillion parameters and fire only a fraction of them per word. 8
Even the hardware gains are mostly architectural rather than brute force. NVIDIA’s GB200 NVL72 rack delivers up to 30 times the inference throughput of the same number of previous-generation H100 chips, at around 25 times less energy for the same work. 9 The watts per chip went up. The useful work per watt went up far more.
Each of these is a multiplier on the same physical base. Stack them and you get the hundredfold collapse in the cost of intelligence that the buildout debate never mentions.
The cost of GPT-4-class intelligence fell roughly 99 percent in two years. Almost none of that came from adding power.
Compute is a product
Raw physical capacity, multiplied by how much useful work each unit of that capacity buys. The capacity stack sets the first term. The efficiency stack sets the second. They run on different clocks, they are built by different people, and the one that is currently scarcer sets the ceiling on what you can do.
Once you read compute that way, the contradictions in the capex fight resolve.
Go back to the 18 billion dollars. Jensen Huang is right that physical capacity is scarce today. A grid connection does take longer than a training run, and the firm that waits loses ground it cannot buy back at any price. Amodei is also right that the return on that capacity is uncertain. Both of them are arguing about the first term and treating the second as a constant.
It is not a constant. It is improving 10 to 50 times a year. That cuts in two directions at once. A capex bill that looks insane against today’s efficiency can look cheap against next year’s, because the same site serves far more useful work for the same power. And capacity bought to serve a workload that the efficiency stack is about to make trivially cheap is capacity that strands. The danger in the buildout is owning the wrong term : paying for raw capacity after the binding constraint has moved to the multiplier, or perfecting the multiplier when you cannot get the megawatts to run it on.
Three years ago the next sentence would have sounded like a category error.
A 2-gigawatt site with a mediocre serving stack loses to a smaller site with a better one.
Where the constraint goes after silicon
The migration does not stop at the efficiency stack either. It keeps walking.
Once serving is efficient and the power is online, the slowest layer becomes the one furthest from the metal: whether an organisation can absorb what the stack has made cheap. Jensen Huang’s own example is the sharpest version of it. A 500,000-dollar engineer who consumes only 5,000 dollars of tokens a year shows the failure mode. 10 The tokens are nearly free, and the company still cannot route its own work to the capacity it already owns.
This is the layer Amodei and Satya Nadella keep returning to from opposite ends of the argument. The technical stack gets good faster than institutions reorganise around it. The final constraint on compute is organisational. It is how quickly people change what they do.
The tokens are nearly free. The bottleneck is the company.
A test you can run this week
Take any AI bet you hold, whether it is a position, a product, or a career, and do three things.
답글 남기기