Google has put caps on Meta’s access to its Gemini models because it does not have the compute to serve them, according to Financial Times reporting that cites three people familiar with the matter. Meta buys Gemini access through Google Cloud, and in March 2026 it learned it would not receive the full capacity allotment it had requested. Staff were told to use AI tokens more efficiently while workloads got trimmed to fit.
Read that again: one of the two or three biggest AI spenders on the planet is being rationed by a direct rival, and the rival’s reason is not strategy, it is stock. Both companies declined to comment on the report.
What Meta was actually buying
The arrangement is stranger than the headline. Meta uses Gemini for content moderation and safety work: flagging harmful posts and rooting out scams across Facebook and Instagram. Per the FT’s sources, Gemini outperformed Meta’s own Llama models at these tasks, which is why the company that built Llama was renting a competitor’s model for some of its most sensitive production workloads.
When the caps landed, Meta began shifting moderation work onto Muse Spark, a newer internal model developed under its Superintelligence Labs division, per coverage of the report. It has also told employees to treat inference tokens as a scarce resource, which at Meta’s scale they now literally are.
The numbers
Google is not short of money, it is short of machines. The company is spending more than $180 billion on capital expenditure in 2026 and still cannot serve everyone: Google Cloud booked roughly $20 billion in first quarter revenue, its order backlog nearly doubled year over year, and CEO Sundar Pichai has already acknowledged that compute shortages held back stronger results. Google has even agreed to pay SpaceX roughly $920 million a month for bridge capacity from about 110,000 Nvidia GPUs, per the same coverage.
Meta, for its part, has guided to $115 to $135 billion in 2026 capex and is pouring money into its own giant data center sites. In May it cut about 8,000 jobs while reassigning 7,000 people into AI roles. None of that buys capacity fast enough to matter this quarter, which is the whole story: money is abundant, delivered compute is not.
The signal
Compute scarcity is now shaping alliances between rivals, and it decides which projects live. When Meta rents intelligence from Google, Google’s allocation desk effectively becomes a planning input for Meta’s product roadmap. That is an extraordinary amount of leverage to hand a competitor, and it is the same dependency trap we flagged in our look at free compute deals: whoever supplies the model can reprice it, ration it, or cut it off, and the customer’s fallback is whatever internal model it can promote under pressure.
It also confirms the supply-side thesis for 2026: the binding constraint on the AI buildout is not demand or capital, it is memory, power, and delivered accelerators. If a $180 billion capex budget cannot keep Google’s largest cloud customers whole, the shortage is structural, not a scheduling hiccup.
The caveats matter here. The entire account rests on three unnamed sources, Google and Meta both declined to comment, and no actual cap figures have been disclosed. “Rationing” could describe routine cloud capacity allocation under a heavy backlog, dramatized by whoever leaked it. Meta also has an obvious interest in the framing: a story about Google failing to supply capacity conveniently explains a migration to its in-house Muse Spark model and hands Meta leverage in any pricing renegotiation. The claim that Gemini outperformed Llama comes from the same unnamed sources, not from a published benchmark.
What to watch
Three tells over the next two quarters. First, whether Muse Spark actually absorbs Meta’s moderation load, and whether moderation quality on Facebook and Instagram visibly wobbles during the handoff. Second, Google Cloud’s backlog commentary at second quarter earnings: if the backlog keeps compounding, more customers than Meta are being rationed and are simply quieter about it. Third, whether Meta pulls forward its own buildout timelines; a company that just learned a rival controls part of its safety systems has every incentive to overbuild.
The AI economy keeps producing these arrangements: rivals selling each other the inputs both of them need, right up until the inputs run short. Compute is the new allocation politics, and Google just showed everyone who holds the clipboard.
AI and business tech news, verified by a physician who reads the filings. One email a week, no noise.
