Articles

Microsoft Compute Crunch: Copilot Before Azure Customers

Microsoft is so short of AI compute that Copilot gets served before Azure cloud customers, executives say — even as sales quotas climb ahead of earnings.

Chisato Chisato · · 5 min read
Rows of illuminated server racks inside a large data center

Two days before Microsoft reports quarterly earnings, the company’s most pressing problem is not demand — it is that it cannot build capacity fast enough to meet it. According to reporting citing multiple Microsoft executives, the company is so short of AI compute that its own products, including Microsoft 365 Copilot and GitHub Copilot, are served before paying Azure cloud customers, even as the cloud sales organization is being handed quotas that rise roughly 30%.

The account, surfaced by Business Insider and echoed across industry press ahead of Microsoft’s July 29 results, describes an internal triage order for scarce GPUs. One executive summarized the hierarchy bluntly: “All of the supply is gone once you solve for frontier labs and our internal businesses like M365 and Microsoft AI.” In other words, after Microsoft satisfies the compute needs of the AI labs it hosts and its own first-party Copilot products, comparatively little is left for the Azure customers whose subscriptions the sales team is being pushed to grow.

The order of priority

The reported allocation sequence puts Microsoft’s own strategic bets first. The company is said to solve first for frontier model training and inference — the workloads that keep it at the AI frontier — and for its internal Copilot businesses, then for research and development, with the remainder going to serve external Azure demand that continues to grow faster than supply.

That ranking is rational from a corporate standpoint and awkward from a customer one. Copilot seats are among Microsoft’s highest-margin, fastest-growing revenue lines, and keeping frontier workloads fed protects the company’s competitive position. But Azure’s promise to enterprises has always been near-limitless, on-demand capacity. When the hyperscaler’s own products sit ahead of paying tenants in the queue for chips, that promise frays — particularly in the specific U.S. regions where GPU-backed instances are hardest to come by.

Why Microsoft is short

The shortage is not new, but insiders say it has gotten worse. Microsoft’s internal forecasts have shown the data-center capacity squeeze that began in 2025 extending well into 2026, with demand outstripping supply in key U.S. regions through at least mid-year. The constraint is physical: not enough power, not enough shell space, and not enough of the most advanced accelerators to fill the racks that do exist.

Microsoft has responded with one of the largest infrastructure programs in corporate history. For fiscal 2026, which began July 1, the company has said it plans to increase AI capacity by roughly 80% and to nearly double its data-center footprint over the following two years. It has also broadened its silicon base — recently expanding Azure AI and HPC capacity with AMD accelerators alongside its existing Nvidia fleet and in-house designs — to avoid being bottlenecked on any single supplier.

Even so, capacity comes online in months and years, while demand arrives in weeks. That mismatch is what forces the triage. It is also the same wall every hyperscaler is hitting: the hundreds of billions of dollars flowing into AI data centers are constrained less by willingness to spend than by how fast power, land, chips, and memory can be assembled into working clusters.

The tension inside the numbers

The disclosure lands at an uncomfortable moment. Microsoft stock has lagged the megacap group in 2026, down sharply year-to-date, and Wednesday’s report will be scrutinized for whether AI infrastructure spending is finally converting into monetization. The key benchmark is Azure growth, where management guided for constant-currency expansion in the high-30s percent range and where investors have set an unofficial bar around the mid-30s.

Here the compute crunch cuts two ways. A capacity shortage means Azure’s reported growth is supply-limited — the company could, in theory, be selling more if it had the chips. That framing turns a constraint into a bull case: demand so far ahead of supply that revenue is capped by physics rather than by sales. But it also raises a harder question the market has been asking all season, most visibly after Alphabet’s capex-driven selloff: if Microsoft must pour ever-larger sums into infrastructure just to keep pace, and still cannot serve its own paying customers, when does that spending translate into durable free cash flow rather than perpetual reinvestment?

Raising cloud sales quotas by 30% into a capacity shortage sharpens the contradiction. Sellers are being asked to book more Azure consumption that the company may struggle to provision, a setup that risks either disappointed customers or deals that slip because there is nothing to run them on.

Part of a wider squeeze

Microsoft’s triage is a symptom of an industry-wide compute famine, not an isolated stumble. The same scarcity is visible in the extraordinary financing being arranged to secure future capacity — from OpenAI’s projections of $750 billion in compute spending by 2030 to the $250 billion data-center backstop Nvidia is reportedly weighing for OpenAI. When the biggest buyers are pre-committing to years of supply and the biggest supplier is underwriting their build-outs, it is because compute has become the binding constraint on the entire sector’s growth.

That scarcity also reshapes the competitive map. Microsoft’s decision to prioritize frontier labs and first-party Copilot reflects where it believes the durable value sits — and it echoes a broader strategy of leaning into its own AI stack, from sovereign-AI partnerships in Europe to aggressive Copilot positioning. For enterprise customers, the practical takeaway is that raw cloud capacity is no longer a commodity to be assumed; it is a scarce resource to be contracted for in advance, much as the economics of AI data centers now revolve around securing power and silicon years ahead of need.

What it means

The reported priority order — frontier labs and Copilot first, Azure customers last — tells enterprises something important about the platform they are building on: during a shortage, the landlord’s own tenants come first. That is a defensible business choice, but it changes how buyers should plan. Reserved capacity, longer-term commitments, and multi-cloud contingency stop being optional hedges and become the baseline for any workload that cannot tolerate being deprioritized.

Who benefits: Microsoft’s highest-margin products and its frontier-model position, which stay fed regardless of external demand. Rival clouds may also benefit at the margin, picking up workloads from Azure customers who cannot get the capacity they need.

Who is exposed: enterprise Azure customers competing with Microsoft’s own AI products for the same GPUs, and a sales force being asked to grow bookings 30% against a supply base that cannot keep up. If Wednesday’s guidance leans on “supply-constrained demand” to explain any Azure shortfall, expect investors to press on the follow-up: strong demand is only bullish if the company can actually build fast enough to serve it.

What to watch next: Microsoft’s July 29 Azure growth number and, more importantly, management’s language on capacity timing — when the 80% capacity increase and doubled footprint actually relieve the crunch. That answer will color not just Microsoft’s quarter but the entire megacap earnings and Fed week, where the market is grading whether AI’s enormous capital bill is finally starting to pay.