Newsletter/Models

The Frontier Premium, Part 2: Open Just Got Big

The Frontier Premium, Part 2: Open Just Got Big

Kimi K3 lifted the floor. Eight days later, Opus 5 reset the frontier above it. Both are demand for compute.


Five weeks ago, in the GLM-5.2 piece, I wrote that the clearest disconfirmer of the frontier labs' rolling monopoly would be the open-model lag shrinking: Chinese releases arriving weeks behind the US frontier instead of months.

On July 16, Moonshot AI released Kimi K3 through its own products and API. On July 27, the weights and technical report went public, making it the largest open-weights model ever released: 2.8 trillion total parameters and 104 billion activated parameters per token. On GDPval-AA v2, a benchmark built from real tasks across 44 occupations, Moonshot reports a score of 1,687. Artificial Analysis independently scores it 1,668, which lands it in the same place: third behind Claude Fable 5 Max and GPT-5.6 Sol Max, above the Opus 4.8 tier.

Read that again. A downloadable model outscores what was Anthropic's flagship at the start of this year. Eight days after its launch, Anthropic replaced that flagship tier at the same price. That sequence, not either event alone, is what this piece is about.

GDPval-AA v2 at K3's July 16 launch: Fable 5 Max, GPT-5.6 Sol Max, Kimi K3, Opus 4.8

Launch-period reported results; benchmark versions and reported scores can change.

The lag has not collapsed to zero. Fable 5 Max still leads, and the newest closed tier remains out of reach. But the prior frontier tier, where API price matters most, now has an open-weights competitor that anyone can download, subject to material commercial terms for some large model-as-a-service operators.

GLM-5.2 shrank the frontier premium on price. K3 challenges it on scale, and in doing so breaks the assumption that has organised the whole open-versus-closed debate: that open models are the small, cheap, good-enough tier. The consequence that matters is not where intelligence lives. It is where the demand for computation goes.

The Scale Inversion

For most of this cycle, the mental model was stable. Closed labs built the giants; open labs built efficient followers. DeepSeek's breakthrough was doing more with less. GLM-5.2's pitch was frontier-adjacent capability at one-tenth the price. Open meant lean.

K3 abandons that positioning entirely. 2.8 trillion parameters, 896 experts, 104 billion activated parameters per token, a million-token context, reasoning switched on by default. Moonshot did not build a cheaper follower. It built the largest model ever released as open weights. And the cadence is the sharper point: K2.6 in April, K2.7-Code in June, K3 in July. Three releases in four months, each positioned closer to the frontier, from a lab working under hardware constraints Part 1 assumed would keep the lag wide.

For buyers, the price challenge is immediate. K3's API lists $3 per million uncached input tokens and $15 per million output, against $5 and $25 for the Opus tier it was reported to beat. A 67% premium is hard to defend against a cheaper model of similar measured capability, and buyers only need a credible reason to run the test.

Then Anthropic changed what that premium buys. On July 24, three days before K3's weights went public, Claude Opus 5 arrived at the same $5 and $25, claiming a new state of the art on this benchmark. The tier K3 was measured against no longer exists at that capability level, and the price did not move.

Both things are true at once, and the tension is the story. An open-weights model reset expectations for what the prior frontier tier should cost, and the closed lab answered within eight days by resetting what that tier is, without moving price. The lead held. The floor still rose.

One caveat applies to the closed side too, and it is the same one this piece applies to K3. A benchmark peak is not a production verdict. In CodeRabbit's code-review evaluation, Opus 5 used roughly 60,000 input and 9,500 output tokens per call, well above GPT-5.6 on those tasks. That is one workload, not a general cost estimate, but it illustrates the point: unchanged per-token pricing does not mean unchanged cost per finished task if the model spends more tokens to get there. The number on the price page is not the number on the invoice.

It is also a model built for demanding workloads. Always-on reasoning can generate thinking tokens on every query; the million-token window invites repository-scale workloads that products will be built to fill. Part 1's cost-per-successful-task equation moves in opposing directions here: fewer dollars per token, more tokens per task, and substantial memory state to support long, concurrent sessions.

That last clause is where the story turns physical.

Open Weights, Industrial Demands

The GLM piece noted that a 1.51TB model is open but not operationally free. K3 answers the same question at a much larger parameter count, and now with published weights rather than estimates.

The release settles the arithmetic that was an estimate a week ago. K3 ships natively in MXFP4 four-bit precision. At 2.8 trillion parameters, that implies roughly 1.4 TB of raw weights before runtime overhead; a 16-bit representation would be roughly 5.6 TB. A community self-hosting overview reaches the same order of magnitude.

Serving K3
Native format MXFP4 (4-bit)
Raw weight footprint, MXFP4 ~1.4 TB before runtime overhead
Raw weight footprint, 16-bit ~5.6 TB before runtime overhead
Deployment shape Multi-accelerator; topology depends on throughput and context targets

Open weights are not the same as an unrestricted commercial licence. Moonshot permits copying, modifying, and deploying K3, but a business operating a model-as-a-service offering whose group revenue exceeds $20 million over 12 months must reach a separate agreement before commercial use. Commercial products and services using K3 above 100 million monthly active users or $20 million in monthly revenue must also display "Kimi K3" prominently. Internal use, and use through Moonshot's official products or certified inference partners, are exempt. That matters because the licence is least frictionless for some of the large API and hosting businesses most able to operate K3 at scale.

Those figures are before any context loads. vLLM's K3 preview shows non-disaggregated serving working, while its disaggregated, expert-parallel path remains in final validation. K3 can be self-hosted; the open question is what a replica costs at production concurrency. The million-token context adds serving state on top of the resident weights, which is what Kimi Delta Attention exists to contain, and the real per-session number will firm up as independent operators publish it.

In a mixture-of-experts model, sparsity primarily lowers per-token compute. A high-throughput serving system still has a strong incentive to keep the full expert set close to compute rather than repeatedly moving inactive experts from slower storage, even though lower-cost offload configurations will be possible. The important point is not a precise per-replica figure. It is that production-quality K3 serving is likely to require a multi-accelerator, memory-heavy system, with additional capacity for concurrency and long context.

Follow that through. Open weights democratize access to the model; they do not make production-grade serving free. The cost ranges from a large hourly cloud bill to capital-intensive owned infrastructure, depending on configuration, utilisation, and redundancy.

So a serious open release can work like a demand broadcast, if adoption follows. Organisations that were never going to train a frontier model can justify renting or buying the infrastructure to run one, and the beneficiaries would be the neoclouds, sovereign programs, and enterprises that own or rent serious compute, plus the suppliers of the capacity those systems consume. If adoption follows, the model layer commoditises while the serving layer absorbs more value. That is the same migration Part 1 described, away from raw intelligence toward what surrounds it, except the beneficiary here sits underneath every model provider at once.

Who Collects When the Model Layer Commoditises

The week K3 launched, TSMC and ASML both reported record or guidance-raising quarters, with capacity commitments running into 2028. Days later Alphabet raised its 2026 capex guidance to $195 to 205 billion from $180 to 190 billion, with Google Cloud revenue up 82% and cloud backlog at $514 billion. None of this is evidence that K3 created demand, since semiconductor orders and hyperscaler budgets are set quarters or years ahead. The companion piece works through what those prints mean for the cycle. The point here is narrower, and it is about who is positioned when models stop being scarce.

TSMC is the purest expression. Whether the winning model is American or Chinese, closed or open-weights, it is likely to create demand for leading-edge accelerators, a large share of which are fabricated by TSMC. TSMC does not need a view on the model race. It invoices a large share of the participants.

That is the structural asymmetry of an open-weights world. A frontier lab's revenue depends on holding a lead. A foundry's does not.

Extending that logic down the stack gives the public-market map.

Layer Names Read Why
Foundry and litho TSMC, ASML Cleanest Model-agnostic, as above
Equipment AMAT, Lam, KLA, Tokyo Electron Positive Logic and DRAM capacity gets funded either way
Memory SK Hynix, Samsung, Micron Positive fundamentally, crowded Serving state scales with size and concurrency; Nvidia locking supply
Cluster plumbing Broadcom, Arista, Vertiv Positive Inference footprints multiply, and they need power and networking
Neoclouds Merchant capacity Mixed New tenant class, but Meta may enter as a competitor
Accelerators Nvidia, AMD Mixed-positive Inference is where competition is hardest
Closed-model economics Microsoft, API-moat software Pressured A credible open alternative at 60% of the price
Hedged Google Neutral to positive Own models, own silicon, cloud backlog $514B; capex commitment is real

Three rows deserve more than a line. Memory is the most direct fundamental beneficiary if the K3 pattern spreads, and the least clean way to express it: the demand per deployment stays unproven until independent implementations arrive, and it is already the most crowded trade in the complex.

It also just received a conspicuous vote of confidence. On July 24, Nvidia and SK Group expanded their partnership across AI factories and next-generation memory, announcing letters of intent under an initiative reported above $500 billion. Those arrangements include a long-term Nvidia-SK Hynix memory collaboration and a two-gigawatt SK Telecom facility planned around Vera Rubin accelerators and HBM4 from 2027. Read it for what it is: the largest accelerator vendor and a major memory supplier planning years ahead, a statement about expected scarcity. Read it with caution too, because a multi-year, multi-entity letter of intent is not capex in a quarter, and reciprocal supplier commitments can make each side's demand look firmer than an arm's-length order book would. Announced and delivered can differ by a lot.

Neoclouds gain a tenant class and a competitor in the same month. A July report says Meta is exploring reselling surplus compute, which would make Meta a potential new merchant supplier of the capacity that neoclouds sell.

And the pressured row is subtler than it looks. The labs are private, so the exposure surfaces through Microsoft's OpenAI economics and through software whose moat is a frontier API call. But momentum-priced open-lab proxies cut both ways: K3 is a reminder that no open lab holds the crown for long either.

The policy fight underlines the stakes. On July 24, Nvidia, Microsoft, Meta, Palantir and roughly twenty other companies, joined by IBM, Hugging Face, Mistral, Mozilla, Perplexity, Andreessen Horowitz and Y Combinator, warned Washington against "premature restrictions" on open-weight models. Anthropic and Google did not sign, and OpenAI reportedly added its name later. The roster does not map to a clean incentive split, since Google publishes open-weight models and sells cloud, and the letter argues on innovation grounds rather than revenue. What it does establish is that a broad coalition, from chipmakers to venture investors, sees enough at stake in open weights to lobby together over it.

One caution runs under all of it. Everything above is an argument about fundamentals, and July kept demonstrating that fundamentals are not what is moving these names. A strengthening demand story does not, by itself, resolve a positioning problem. That is the companion piece's subject, and its conclusion applies here unchanged.

What Could Break This

The obvious risk is that K3 does not survive contact with production. Launch benchmarks select for strengths; long-running agents expose reliability failures that demos miss. Artificial Analysis corroborating the ranking is a stronger signal than a lab's own scorecard, but a benchmark position is still not the same as displacing Opus-class models inside real workflows. If K3 needs more retries and supervision, its cost advantage narrows the way it might have for GLM-5.2.

The release could still disappoint. Public weights do not establish that they reproduce the API model's results, nor do they settle the licence, reliability, or hosting-cost questions. Until independent parties serve and benchmark them, the hosting economics above are informed estimates.

The scale-inversion argument also has a soft spot: one release is not a trend. If K3's size buys little over models a third as large, the industry may conclude that 2.8 trillion parameters was a flag-planting exercise, and the efficient-follower model reasserts itself. Open-weights demand for memory would then grow with adoption, not with model size.

And geopolitics is now the live risk, not the hypothetical one. Washington is reported to be weighing restrictions on Chinese models, including K3. Against that policy backdrop, the open letter arrived. If procurement bans or compliance rules make serving Chinese frontier weights difficult, the demand broadcast reaches a smaller audience and reaches it later. The effect would be uneven: in restricted markets, removing the cheapest credible alternative hands some pricing power back to closed models until a permitted open model fills the gap; elsewhere, serving demand follows whatever weights remain available, since the hardware does not care about the model's passport. Restriction redistributes the premium by jurisdiction more than it restores it globally.

What We Are Watching

Cost per finished task, on both sides. With weights out and vLLM's working preview available, the open question is no longer whether K3 can be self-hosted but what a replica costs at concurrency, thinking tokens included. The same test applies to the closed tier: CodeRabbit clocked Opus 5 well above GPT-5.6 on tokens per call, and a benchmark lead bought with more tokens is a softer lead than the price page implies. The Jevons mechanism only kicks in if cheaper per-token cost expands total usage faster than efficiency saves it.

The closed-lab cadence. Opus 5 answered the first round by resetting the tier through capability rather than price. Watch whether that holds, whether pricing eventually follows, and how fast the top keeps moving. The frontier premium does not disappear; it retreats upward. The question is whether it can outrun the weights.

The lag itself. Part 1 said a sustained move from four-to-six months toward weeks would break the rolling-monopoly thesis. July cut both ways: K3 landed above the tier Anthropic was selling that week, and Anthropic replaced that tier eight days later. The next two Chinese releases, measured against Opus 5 rather than its predecessor, will say whether the lag is compressing or whether July was a high-water mark.


The frontier still belongs to the closed labs, and Opus 5 was the proof, delivered eight days after K3 and at the same price as the tier it replaced. Nothing here changes who owns the newest, hardest capability.

What is changing is the floor. The weights are public, which makes the open tier enormous, cheap, and roughly one generation back, and that forces the closed labs to keep moving. Opus 5 is what keeping moving looks like. That is the loop worth watching: open weights lift the floor, the frontier resets above it, and the previous frontier becomes something anyone can download.

Every turn of that loop could mean more inference to serve, and more hardware to serve it on. The labs are competing over the model layer. The bill lands underneath.

GLM-5.2 showed that being behind the frontier matters less than it used to.

K3 will test what that costs to serve.


This is part of an ongoing series on AI infrastructure economics. Part 1 covered GLM-5.2 and the shrinking frontier premium; earlier pieces covered the physical bottleneck sequence, the memory demands of agentic work, and why price per token no longer reflects the cost of completing a task.

This is analysis, not investment advice. Model specs, prices, and benchmark scores move quickly; recheck the figures against the primary sources before relying on them.


Sources

Get the next analysis

AI infrastructure, model economics, and agentic software, delivered through Substack.

Subscribe