The Binding Constraint
The limit on AI infrastructure is not capital, land, or chips. It is the interconnection of a very large electrical load to a grid that was never planned for it, and that constraint is quietly deciding where computing happens.
There is roughly $700 billion of hyperscale capital expenditure committed for calendar 2026, and about $2.5 trillion of contracted customer backlog sitting behind it. Vacancy across North American data center capacity is 1.4 percent. More than 80 percent of the 7,481 megawatts currently under construction is already leased, up from 74.3 percent a year ago.
Those three facts do not describe a market with a demand problem or a capital problem. They describe a market that cannot deliver fast enough, and the reason is narrower than most coverage suggests. It is not concrete, land, chips, or money. It is the interconnection of a very large electrical load to a grid that was never planned for it.
Once the constraint is named correctly, the long running argument about centralized versus distributed computing stops being philosophical. Nearly every distributed architecture attracting serious investment in 2026 is a workaround for a power problem. The ones that are not are solving for something else entirely, and it is almost never latency.
The scarce thing
Lawrence Berkeley National Laboratory publishes the authoritative accounting of the United States interconnection queue. Its 2026 edition, covering data through the end of 2025, reports a median of 61 months from interconnection request to commercial operation for projects that came online in 2025. That is just over five years. Of all capacity that requested interconnection between 2000 and 2020, only 13 percent had reached commercial operation by the end of 2025.
Generation equipment is no easier. GE Vernova reported roughly 100 gigawatts of gas turbines under contract as of the first quarter of 2026, combining firm backlog and slot reservations, against a lead time of about three years for new heavy duty orders. Management has been candid that the turbine is often not even the binding item. Engineering, permitting, and fuel supply frequently are.
Then there is the load queue itself, where the numbers turn strange. ERCOT told the Texas House State Affairs Committee in April 2026 that it was assessing approximately 410 gigawatts of large load interconnection requests, roughly 87 percent of them data centers. The research firm SemiAnalysis, which screens projects for secured land, a defined power solution, and approved permits, concluded that 311 of those 410 gigawatts are phantom: requests without the site control or financing to ever become real.
Four workarounds, four architectures
Facing a five year queue, an operator has four moves. Each produces a different computing architecture, which is why the architecture debate and the power debate turn out to be the same debate.
Bring your own generation. Meta's Hyperion campus in Richland Parish, Louisiana expanded in July 2026 from 2 to 5 gigawatts at a cost above $50 billion, supported by seven combined cycle gas plants financed through Entergy Louisiana and roughly 240 miles of new high voltage transmission. xAI's Colossus in Memphis runs substantially on dedicated gas generation, including a 1.2 gigawatt plant permitted across the state line in Southaven, Mississippi, alongside a far smaller grid connection. This is the fastest available path to power. It also carries the largest local political surface area, as the permitting litigation around the Memphis turbines has demonstrated.
Split the run across sites. If a gigawatt is unavailable in one place, take 300 megawatts in three. The obstacle was always technical. Training a large model has historically required every accelerator to exchange data with every other accelerator at bandwidth no wide area network can supply. That obstacle is now falling. Google DeepMind published Decoupled DiLoCo in April 2026, reporting a 12 billion parameter model trained across four separate United States regions using 2 to 5 gigabits per second of wide area networking, a level DeepMind characterizes as achievable over existing connectivity between facilities. Microsoft operates its Fairwater sites in Wisconsin and Atlanta, roughly 700 miles apart, as a single coherent cluster over a dedicated AI wide area network, having laid more than 120,000 new fiber miles across the United States in a year.
Distributed training was not pursued because distribution is elegant. It was pursued because nobody can energize a gigawatt campus on demand.
Become a flexible load. In March 2026 Google announced approximately 1 gigawatt of contracted demand response across five utilities, including Indiana Michigan Power, TVA, Entergy Arkansas, Minnesota Power, and DTE Energy, agreeing to shift or curtail machine learning work when the grid is tight. PJM has gone further and made it compulsory. Its Interim Resource Adequacy Service requires new large loads arriving without dedicated supply to reduce consumption or switch to onsite backup when the system nears emergency conditions. PJM's board separately proposed a backstop capacity auction to cover a 6,831 megawatt shortfall for the 2028/29 delivery year at a cost cap of $555 per megawatt day, raised from $325.
This is the most underrated development of the year. Training and batch inference are deferrable in a way that a hospital, a smelter, or a water treatment plant is not. AI data centers are simultaneously the largest new load arriving on the grid and the most curtailable large load ever connected to one. That is a negotiating position. The operators who recognize it early will secure interconnection agreements that the ones who do not will not.
Put the work where no interconnection is needed. Which brings us to the part of the market discussed most often and understood least.
Where each architecture actually wins
Distributed and edge computing have been oversold for a decade on the wrong argument. The pitch was latency. The evidence is that latency almost never decides.
AT&T's chief technology officer said plainly in April 2026 that he was "not sure that there's much value in extending that compute all the way to the far edge just to save another millisecond or two milliseconds of latency." Both AT&T and Verizon have effectively abandoned mobile edge computing as a service business and pivoted to becoming landlords for AI data centers, which is a considered judgment about where the value sits. Telefonica has five commercially live edge nodes. Industry forecasts published a few years ago called for nearly 1,200 network edge data centers by 2026.
Latency failed as a predictor because the arithmetic rarely favors it. Light travels roughly 200 kilometers per millisecond through fiber, so the near edge saves most users somewhere between 10 and 50 milliseconds against a well placed cloud region. That is decisive for a real time bid auction, bot mitigation, TLS termination, or an industrial control loop. It is irrelevant to a language model response that takes 800 milliseconds to generate.
What does decide placement is less glamorous and considerably more durable.
| What is actually binding | Where the work goes | Representative evidence |
|---|---|---|
| Tight coupling: the compute must exchange enormous volumes with itself | One coherent fabric, wherever power can be assembled | Inside an NVIDIA rack, accelerators exchange data at 1.8 to 3.6 terabytes per second. Step outside to the scale out network and the figure is roughly 0.1 |
| Loose coupling, power scarce | Split across regions | DeepMind trained a 12 billion parameter model across four US regions on 2 to 5 Gbps |
| Data gravity | Next to the data | Moving a petabyte out of a major cloud costs roughly $60,000 to $92,000 in egress alone |
| Sovereignty and jurisdiction | Inside the customer's own boundary | Google now ships an air gapped appliance running Gemini on premises. AWS launched its European Sovereign Cloud in Brandenburg in January 2026 |
| Disconnection tolerance | On site | Retail, clinical, industrial, and tactical environments where the network is assumed to fail |
| Steady state, predictable, storage heavy | On premises | 37signals cut infrastructure spend from $3.2 million per year to under $1 million with no added staff |
| Bursty, unpredictable, still being designed | Cloud | Optionality is the product being purchased |
| Latency budget under roughly 10 milliseconds | Near edge or on device | Ad auctions, bot mitigation, control loops |
The most instructive case in the category has nothing to do with milliseconds. Chick-fil-A runs more than 2,500 Kubernetes clusters, one per restaurant, on deliberately consumer grade hardware. The reasons given were unreliable connectivity, the requirement that kitchen production and payment terminal onboarding keep working when the link drops, and the fact that no commercial vendor offered a licensing model that made sense at thousands of active clusters. Software licensing economics, not physics, drove the architecture. That is a more common story than the industry admits.
There is a symmetrical caution on the other side. The lab most publicly committed to decentralized training, Prime Intellect, built its strongest recent model on 512 accelerators inside a single high speed fabric rather than across a distributed network, because the reinforcement learning stage that drives most current capability gains cannot tolerate loose coupling. Low communication methods work for pretraining. They do not yet work for everything that follows it.
What to do with this
Treat power availability as a procurement input, not a facilities footnote. If a provider cannot state the energization date and whether it rests on grid interconnection, behind the meter generation, or a study not yet complete, they have answered the question. Ask for the same detail about the second and third phases, which is where schedules usually slip.
Sort workloads by coupling before sorting them by anything else. Work that must exchange large volumes with itself at high frequency belongs in one place, and that place will be expensive and slow to obtain. Work that does not is far more portable than most organizations assume, and portability is now worth real money. The financial case for moving steady state production off metered cloud is stronger in 2026 than it has been at any point since 2015, and it is strongest exactly where the workload is boring.
Discount the demand numbers in both directions. Three quarters of one major queue is not real. Built capacity is 98.6 percent occupied. Both statements are true simultaneously, and any forecast that does not distinguish requested capacity from energized capacity is not a forecast.
Bottom line
The industry spent a decade debating where computing ought to happen and answering with architecture diagrams. In 2026 the answer is being written by interconnection studies, turbine slot reservations, and demand response contracts. It is a less interesting conversation. It is a considerably more consequential one.
Selected sources
- Lawrence Berkeley National Laboratory, Queued Up: 2026 Edition
- ERCOT, Large Load Update, House State Affairs Committee, April 2026
- SemiAnalysis, Stop Saying Half of 2026 US Datacenter Capacity Is Canceled
- GE Vernova, First Quarter 2026 Results
- CBRE, North American Data Center Demand Continues to Outpace Supply, August 2026
- Google, A Demand Response Milestone for Data Centers, March 2026
- PJM, Reliability Backstop Proposal
- Google DeepMind, Decoupled DiLoCo, April 2026
- Microsoft, The Architecture Behind the Azure AI Superfactory
- Light Reading, AT&T CTO Casts Doubt on AI Compute at the Far Edge, April 2026
- Light Reading, AT&T and Verizon Are Pivoting Into the Landlord Business for AI
- STL Partners, Edge Computing at MWC 2026
- Chick-fil-A Tech, Enterprise Restaurant Compute
- Basecamp, Leaving the Cloud
- Google Cloud, Google Distributed Cloud at Next '26
- AWS, AWS Launches the European Sovereign Cloud, January 2026
- Prime Intellect, INTELLECT-3 Technical Report
A Stewart Consulting briefing, published for discussion. Figures are drawn from company filings, regulatory testimony, and published research as of September 2026, and are cited as orders of magnitude rather than guarantees; queue and capacity numbers in particular move quarter to quarter. Nothing here is financial, tax, or legal advice. Questions or feedback are welcome, see Connect.