Week of 2026-07-06 to 2026-07-12 · Five frontier models shipped in one stretch and half of US enterprise tokens are now Chinese — while the labs borrowed and leased tens of billions for power and silicon. The token raced to the floor; the capital went vertical.

Two things happened this week, and they are the same thing.

The first: capability got abundant and cheap. On July 8 and 9, three more frontier-class models reached general availability — OpenAI’s GPT-5.6 in three tiers (Sol, Terra, Luna), xAI’s Grok 4.5, and Meta’s Muse Spark 1.1, with Tencent’s Hy3 landing alongside. Add the models already live — Fable 5, Sonnet 5, GPT-5.5 — and for the first time a developer could open five frontier-class models on the same afternoon. Meanwhile CNBC reported, using OpenRouter data, that Chinese open-weight models hit a weekly peak of 46% of US enterprise token usage — up from 11% averaged over the prior year, and above 30% every week since February. The single largest vendor on the platform is DeepSeek, not OpenAI.

The second: capital got expensive and concentrated. In the same seven days, Amazon raised at least $25 billion in bonds for data centers (on the way to ~$200B of 2026 capex), Anthropic signed a 20-year, ~$19 billion lease for 401MW of power at a former Kentucky aluminum smelter, and Meta confirmed it will start producing its own inference chip, “Iris,” in September on the way to 14 gigawatts by 2027.

Put the two next to each other and the week has a shape. The product raced to the floor. The money fled to everything around it. The token is now a commodity input you multi-source and haggle over. The margin has left it. So the capital is pouring into the four things a commodity can’t erase overnight: power, silicon, the surface where work happens, and the data only a used product produces.

The commoditization is not a forecast anymore. It’s in the bill.

We have written the mechanism of this for a month — the price cut as a weapon, the open-weight floor, the export ban that couldn’t touch the capability. This week the mechanism became a number on an invoice.

The number is 46%, and the reason is boring: the Chinese models are 60–90% cheaper and close enough. DeepSeek V4 Flash lists at $0.14 per million input tokens against GPT-5.5’s $5.00. GLM-5.2, shipped last month under an MIT license, saw the fastest adoption of any model Vercel tracked in 2026 — daily tokens up ~27×, customers up ~80× in its first week. Airbnb’s Brian Chesky says the company runs Alibaba’s Qwen for customer service because “it’s fast and cheap” and it cut resolution time from three hours to six seconds. The clean version of the margin-collapse thesis making the rounds this week is right even if the forecast is wrong: commodity inference has already arrived, and increasingly it is a commodity you can run yourself.

The multi-lab launch day and the 46% number are the same story told twice. When five labs can ship frontier-class capability in one week and half your tokens already come from models you can download, the model has stopped being the differentiator. Simon Willison, testing GPT-5.6, put the practitioner reality plainly: OpenAI claims a new high on Agents’ Last Exam, but “so far it hasn’t struck me as better than Fable at the kind of complex coding tasks I’ve been using,” and “price-per-million tokens doesn’t tell us much now that the number of reasoning tokens can differ.” That is what a commodity market sounds like from the inside: the spec sheets diverge, the felt difference narrows, and the buyer starts optimizing on price.

So the money went where a commodity can’t reach.

If the token won’t hold margin, capital chases what will. Four directions, all visible this week.

Power. Anthropic’s $19B, 20-year lease is not a model bet — it’s a bet that electricity and cooling will be the scarce input long after any given model is obsolete. It follows the same logic as its $40B TPU commitment and Meta’s 14GW. This is the deep end of the circular-financing story we dived last week: the risk in the AI trade sits not at the solvent hyperscaler end but in the levered middle that finances the buildout. A 20-year lease on a smelter site is Anthropic putting its balance sheet exactly where the moat is.

Silicon. Meta’s Iris chip and DeepSeek’s own reported inference chip are the labs-go-vertical thread advancing on both sides of the Pacific in one week. Nvidia keeps ~70 cents of every inference dollar; owning the chip is how you claw it back. When the token sells at the floor, the only way to keep any margin is to own more of the stack under it.

The surface. ChatGPT Work launched July 9, an agent that produces finished documents across your apps and files, positioned squarely against Microsoft 365 Copilot. Anthropic answered with Claude Cowork for mobile. Both labs are climbing up-stack into the knowledge-worker application because the model layer no longer defends itself. This is the channel war one rung higher than the terminal: not the harness, the workspace.

The data. Grok 4.5 shipped trained partly on Cursor IDE session data after SpaceX’s $60B buy of Anysphere — the accept-button moat we dived a week ago, now shipping in a product. It landed fourth on the Artificial Analysis Intelligence Index, behind Fable 5, GPT-5.5, and Opus 4.8, with no system card. Human preference-on-correctness is the one input that hasn’t commoditized. Everyone is now buying the surface where the accept happens.

Anthropic pulled the flagship — then blinked.

The sharpest local evidence that the token has become a place to defend margin, not win share, is what Anthropic did to its own best model. Starting this week, Fable 5 leaves the Pro/Max/Team subscription and bills through pay-as-you-go usage credits at $10 / $50 per million tokens — double Opus 4.8, and the highest per-token price Anthropic has ever listed for a shipped model. The frontier stays premium while the floor collapses under it. That is the whole pricing thesis in one product page.

There is a wrinkle worth naming, because we bet against exactly this shape and lost that bet in W27. Facing backlash, Anthropic extended included Fable access to July 12 and said the metering is temporary, “until capacity allows.” In W23 we predicted GitHub would blink on Copilot pricing under pressure; it didn’t — it tightened. So does Anthropic’s blink vindicate the instinct? Not really. GitHub held a structural reprice and refused to walk it back. Anthropic extended a promo by four days and kept the meter. The direction is identical on both: the flagship gets metered, the subsidy dies. A capacity-gated promo extension is the meter finding its level, not a reversal of it.

The one lab that didn’t ship is the tell.

Google is the only major lab without a publicly available frontier model on the busiest launch day of the year. Gemini 3.5 Pro slipped to July 17, a full architectural rebuild, after early testers flagged token-efficiency and coding gaps — and after four senior researchers, including MoE co-author Noam Shazeer, left for OpenAI and Anthropic, helping wipe ~$225B off Alphabet in June. In a week that proved capability is abundant, the company with arguably the deepest compute and research bench couldn’t get a competitive model out the door on time. Abundance is not the same as ease. The frontier is crowded and hard, which is why the labs that can ship are the ones spending on the stack underneath, not just the weights on top.

So here is the read a working engineer should carry out of this week. Model choice is now the least interesting decision in your stack. Five frontier models, a Chinese commodity tier at half the price, and an open-weight floor you can self-host means the leverage moved to everything around the model call — your eval discipline, your routing, your context scaffolding, and, increasingly, which company owns the power and the surface you’re renting. Pick your model like you pick a cloud region: cheaply, reversibly, and with a fallback that’s already eval’d.

And there is one commodity so cheap and so widely adopted that most teams are already running it without having decided to — Chinese open-weight models, quietly, through American routers and clouds, while Washington tries to figure out whether it can stop them. That’s the deep dive.

Also this week

One thing to watch

The token went to the floor this week; the capex went up. Prediction (70% confident): across the next two earnings cycles (through Q1 2027), at least two of {Amazon, Meta, Microsoft, Alphabet} raise their 2026–2027 AI capex or infrastructure guidance even as frontier API list prices fall or hold — the divergence between a commoditizing token and a compounding infrastructure bill widens rather than closes. If capex guidance flattens or falls alongside prices, the commoditization is eating the whole stack and this call is wrong.

Deep dive commissioned: the 46% number — how Chinese open-weight models came to serve nearly half of US enterprise tokens, why the export-control playbook can’t reach them, and what a product engineer should actually do about the models they’re already running.