The standard coverage of Chinese artificial intelligence laboratories asks whether they have caught up. It is the wrong question, and it has been the wrong question for at least two years, because it measures the variable that matters least to the people making purchasing decisions.
Enterprises do not buy benchmark positions. They buy a task completed reliably at a cost that clears an internal threshold. The competitive development represented by Moonshot AI and its peers is not that they have matched United States laboratories on the hardest reasoning tasks, which on published evidence they have not. It is that they have made capability sufficient for the majority of commercial workloads available at a price and through a channel that United States providers cannot match without abandoning the revenue model their infrastructure spending assumes.
This piece is written to remain useful after the next model release, so it concerns cost structure and distribution rather than benchmark tables.
The Cost Curve Is the Competitive Surface
There are two costs in artificial intelligence and public discussion conflates them constantly. Training cost is one time, enormous, and concentrated in a handful of organisations that can finance it. Inference cost is recurring, per request, and paid by every user of the resulting model for as long as they use it.
Training cost determines who can build a frontier model. Inference cost determines who can afford to deploy one, and it is therefore the only one of the two that shapes the market for delivered capability. A laboratory that spends less to train but the same to serve has achieved nothing commercially. A laboratory that serves comparable output at a fraction of the price per token has changed the market, regardless of how it performs on the hardest available evaluation.
The gap in posted pricing between leading open weight models from Chinese laboratories and premium tiers from leading United States providers has run to roughly an order of magnitude on comparable context lengths. Readers should verify current figures directly against posted provider pricing rather than trust any secondary summary including this one, because the numbers move monthly. The direction has not moved. It has gone one way for three years.
Most Workloads Are Not Frontier Workloads
The premium tier is priced as though every request requires the best available model. Very few do.
Consider what artificial intelligence is actually deployed to do inside a large organisation. It classifies incoming documents. It extracts structured fields from unstructured text. It summarises meetings and long threads. It routes tickets. It drafts first versions of routine correspondence. It answers questions against an internal knowledge base. It translates. It writes and reviews boilerplate code.
Each of those has a quality threshold rather than a quality maximum. Once a model clears the threshold, additional capability produces no additional business value, and the buyer is simply paying more for output they cannot distinguish. A model one generation behind the frontier clears that threshold on the majority of these tasks, and it does so at a small fraction of the cost.
This is the mechanism by which a laboratory that is genuinely behind on capability can nonetheless take the majority of the volume. It is the same mechanism that has operated in every technology market with a wide quality distribution and a low quality threshold for most use cases, and there is no reason to expect artificial intelligence to be the exception.
The Compute Constraint Produced an Efficiency Advantage
The most consequential and least intended outcome of the export control regime is worth stating plainly, because it inverts the expectation the policy was built on.
Controls on advanced accelerators genuinely restrict the scale of training runs achievable inside China. That constraint is real and it is binding. Its effect, however, is not to stop research. It is to redirect it. A laboratory that cannot solve a problem by adding compute must solve it by using compute better, which pushes effort toward architectural efficiency, mixture of experts routing that activates only a fraction of parameters per request, aggressive quantisation, and distillation of large model behaviour into small model footprints.
Every one of those techniques reduces the cost of serving a model. And every one of them is published in open literature, presented at open conferences, and reproducible by anyone within weeks. The asymmetry is stark: compute cannot cross the border, but the knowledge of how to need less of it crosses instantly and in both directions.
The result is that a compute constrained competitor has been systematically pushed toward exactly the capability that erodes the pricing power of a compute abundant incumbent. That is not a prediction. It is a description of the last three years of published research output.
Set what you believe decides the outcome. The ranking reorders against your weights. Country scores are fixed structural assessments and do not change.
Scores are LUMINAIRE structural assessments drawn from public compute, capital formation, grid capacity, research output and policy records. They describe position rather than capability of any single model, and they are reviewed rather than regenerated.
Open Weights Are a Distribution Event, Not a Discount
The pricing gap receives most of the attention. The licensing decision matters more, and it matters for a reason that has nothing to do with cost.
When a model's weights are published, a customer can download it and run it on their own hardware or in their own cloud tenancy. Three things follow immediately. The vendor relationship ends, along with the recurring revenue, the usage telemetry and the upgrade path. The two objections that slow enterprise adoption most, which are data residency and vendor dependency, disappear, because no data leaves the customer's environment and no external party can change terms. And the cost structure converts from a variable per token charge into a fixed infrastructure cost the customer already understands how to budget.
For a regulated institution, a defence contractor, a government department or a healthcare provider, that last point is frequently decisive on its own, independent of price. Many such buyers were never going to send sensitive material to an external inference endpoint at any cost. Open weights do not compete with the premium tier for those buyers. They serve a segment the premium tier could not access.
This is why open weights are not recoverable through price competition. A provider losing on price can reduce its price. A provider losing to a self hosted deployment has lost the customer, and there is no price at which they return, because the objection was never about money.
A provider that loses on price can cut its price. A provider that loses to a self hosted model has lost the customer entirely.
The Question the Capital Expenditure Assumes
This is where the cost curve connects to the financing story running alongside it.
The scale of data centre investment currently underway is underwritten by an implicit assumption: that inference revenue will grow faster than inference cost falls. That assumption is not obviously wrong. Volume growth has been extraordinary and shows no sign of slowing, and cheaper inference expands the set of viable applications, which raises volume further. It is entirely possible for prices to collapse and total revenue to grow.
But it is an assumption, and it is the one the market is actually pricing when it prices the artificial intelligence build out. The forced liquidation of a forty five billion dollar artificial intelligence fund at the end of July 2026, examined in the companion piece to this one, was not caused by a change in that assumption. It was caused by leverage. What it demonstrated is how quickly the assets priced on that assumption reprice when anyone reconsiders the slope of the curve.
Readers assessing which countries are positioned to influence that curve should read the accompanying country by country reference, and can use the scoreboard above to weight the inputs themselves.
What to Watch
Four series will tell the story better than any model release. The first is posted price per million tokens at comparable context length and quality tier, tracked over time rather than at a point. The second is the share of new enterprise deployments choosing self hosted open weight models over hosted premium endpoints. The third is whether premium providers introduce cheaper tiers that concede the general competence market rather than defending it, which would be the clearest possible admission of where the pricing power actually sits. The fourth is whether efficiency research output from compute constrained laboratories continues to outpace that from compute abundant ones.
All four are observable without a forecast, and all four are more informative than the next leaderboard.
