Open Models Will Absorb Most of the AI Spend
Here is my bet: open models (open weights and open source alike) will end up absorbing most of the money and compute the world spends on AI. The proprietary frontier models get the headlines and the IPO valuations, but developers and AI teams see something different up close. Open models are improving fast, the gap to the proprietary leaders keeps narrowing, and the economics point one way. OpenAI and Anthropic are already starting to warn us about open models, and I expect those warnings to get louder. Here is why I think the spend follows.
Follow the money
Frontier pricing is running on a treadmill. The best closed models hold a clear lead for two to six months before someone matches them or distills the capability and gives it away. That is a brutal foundation for a business, and it is why OpenAI and Anthropic are pushing into applications like coding, security, healthcare, and legal, where margin sits closer to the customer. Watch the IPO math too. Both reportedly want to go public above $800B, valuations that assume pricing power the market may stop believing in. I would not bet my roadmap on today’s API prices holding.
Getting real value from this? Consider becoming a paid supporter 🙏
Cheaper per task is not a smaller bill. Cost per finished task is dropping fast, helped by setups where an expensive model plans and a cheap one executes. But Jevons paradox is real: when a resource gets more efficient and cheaper to use, total consumption often rises rather than falls. Cheaper tokens unlock so much more usage that your overall bill can climb even as each task gets cheaper. Budget for cheaper units and heavier use, not for savings that never show up on the invoice.
The premium model is becoming a planner, not a workhorse. Route the hard reasoning and verification to a frontier model, and push the bulk execution to a cheaper one. It is good engineering and good economics, and it quietly breaks the assumption the big labs were built on, that every step runs through a premium API. If you are not architecting this way yet, your competitors already are.
Watch the gap, carefully
Chinese open weights have reached parity in specific, high-stakes areas. In cybersecurity bug-finding, Zhipu’s GLM 5.2 has been measured matching or beating Anthropic’s Opus 4.8. That matters, because finding security bugs is one of the highest-stakes things you can ask a model to do, and open weights help attackers as much as defenders. Still, being good at one thing is not the same as being good at everything, so test these models on your own work instead of trusting news headlines.

Coding agents are still a clear US stronghold, and that gap is not closing quickly. US models ride a flywheel where serious users generate failure data, which trains better models, which pulls in more serious users. Chinese labs are dragged by chip shortages, over-reliance on distillation and benchmaxxing, and pressure to ship consumer products instead of doing open-ended research. That can change, but today the gap holds, so I would not rush to swap out your coding stack.
The “China caught up” story is genuinely contested. Some observers still put the best Chinese labs a year behind, but most estimates land in the three to twelve month range, and it shifts with every release. It matters less than it sounds. For most routine work, my own experience is that open-weight models match the proprietary leaders at a much lower cost, and you can often post-train them to come out ahead. So test on your own tasks, not on anyone’s leaderboard, mine included.
Rethink your stack
Open weights are about control, not just price. You can run an open model on your own hardware, tune it on your own data, and use it without a third party who can change the rules on you. The Fable shutdown made this concrete. When Washington blocked foreign nationals from using the model, Anthropic cut access for everyone, and European buyers started asking whether US models are safe to build on. That is a data moat, a privacy posture, and a hedge against having the rug pulled, not just a cheaper bill. The catch: attackers can run open models just as freely, and many EU and US buyers still will not touch Chinese-origin ones.
A tested fallback is leverage even if you never use it. Just having a working open-weight alternative caps your exposure to a provider’s pricing, rate limits, behavior changes, and policy calls. But the leverage is only real if the fallback is real, running in production with real evals, not a line in a deck that says “we could switch.” A backup you have never tested is not a hedge, it is a hope.
Inference efficiency is the next cost lever. The newest gains are not smarter models, they are faster serving. One Chinese system reports 57 to 85 percent faster generation and several times the throughput under heavy load, with almost no overhead. That attacks the exact cost structure the big labs depend on, so watch how models are served, not only how they score. The cheapest token often wins on plumbing, not intelligence.

Do this now
Run your models as a portfolio, not a religion. Use closed models where they clearly win, open weights where they are good enough or cheaper or more private, and keep the mix swappable. The right question is not “are open models better,” it is “where are they good enough to change my cost and control position.” A portfolio of models still needs governance, evals, and clear ownership, or it turns into expensive complexity.
Build the switching layer before you need it. Routers, OpenAI-compatible endpoints, eval harnesses, memory, logging, and clean abstractions let you change models without rewriting your product or going blind on quality. And treat your prompts, code, and agent traces as strategic assets: decide deliberately what flows to a closed API versus stays local. Those traces are training signals, not exhaust, and you only get to give them away once.
Optimize for resilient, not open or closed. If you remember one thing, make it this. The useful frame is controlled versus resilient. Closed AI gets dangerous only when it becomes infrastructure you cannot replace, because then the vendor’s terms, access rules, and whatever government pressure lands on them become part of your operating environment. Architect so that no single lab, or government, is ever a single point of failure. That is the whole game.

Where I think this all lands
First, the US posture toward China looks self-defeating. Blocking foreign nationals from a US model while clearing advanced chips for export does not slow China down, it hands it an opening and pushes its labs onto a homegrown stack, exactly what the earlier chip bans did. Cutting research ties and treating foreign talent as a threat drives away the people American labs depend on. Export controls have not stalled Chinese progress, and the damage lands on US companies and their customers instead. I would rather see cooperation than a wall, and the durable answer to strong Chinese open models was never to ban them, it is to build better ones.
Second, supply. Open weights are a business decision that can change on a dime, and I used to worry the supply could dry up. I am more optimistic now. Between consortiums, government and enterprise-backed efforts, and new labs stepping in, I expect a steady stream of capable open models, not a drought. Thinking Machines’ Inkling makes the point: not the best model out there, but an open-weights base built for post-training and customization. You do not have to wait for the perfect open model to start cutting your dependence on a closed one.
Your prompt logs and agent traces are strategic training assets, not empty exhaust.
That is the real challenge for OpenAI and Anthropic. In an earlier piece I argued the hybrid stack was coming for their pricing power, and this is the sharper version of that argument. Enterprises are now convinced of three things: they do not want vendor lock-in, they can offload a growing share of compute to open and smaller models they tune themselves, and they would rather not let their own data train someone else’s next model. Put it together and the days of a two-company oligopoly commanding most of the world’s inference are drawing to a close. Which may be why both are racing to go public, building a war chest before that pricing power erodes and pushing into domain-specific markets where an application is easier to defend than a token price. Smart move, and also a tell.

China’s AI Export Controls at a Glance

