Three New Models, One Signal About Where AI Spending Goes Next

Three frontier level models landed within weeks of each other this fall, GLM 5.2 from Zhipu, Kimi K3 from Moonshot, and Gemini 3.6 Flash from Google, and I wanted to capture early developer reaction so I can monitor how feelings about these models change over time. The individual verdicts differ, but together they hint at where enterprise AI spending is actually headed.

Subscribe to our weekly newsletter

GLM 5.2 Makes Good Enough Hard to Ignore

Initial press coverage of GLM 5.2 was that it’s the strongest case yet that open, cheap models are closing in on the frontier, not matching it outright, but close enough that the price gap becomes the actual story. One developer built an entire weekend project, working agent included, for about 20 dollars, the kind of task that runs past 100 dollars on premium models for similar scope. Independent testing landed it just behind the current frontier leader on agentic coding, a genuinely strong result for an open weight model. A security research firm even found it out hunting Claude Code for vulnerabilities at roughly seventeen cents per bug found, though that comparison ran through a custom pipeline built specifically for this model, so I’d treat it as promising rather than proof one model is simply smarter than the other. Because Zhipu uses familiar API styles, pointing your existing tools at GLM 5.2 can be close to a one line change, at least on paper.

The complications show up once you get past the demo. Running it yourself takes something like 80,000 to 150,000 dollars in server grade hardware, so cheap mostly means cheap to rent, not cheap to own. Multiple developers reported it getting stuck in loops or losing the thread on harder problems, with one task that a rival model finished in under five minutes taking GLM 5.2 close to an hour, and once you add up the extra debugging time, the token savings can evaporate fast. There’s also a real geopolitical overhang, with export control moves in both directions still unsettled. None of this proves the model is bad. It does mean the loudest claim going around, that cheap open models are about to trigger a broad price collapse across the industry, might prove premature. Cheap competition has coexisted with healthy margins before, and the bigger question is whether businesses that need support contracts and someone accountable will actually switch over a price difference this size.

Reading the Room on Kimi K3

My take after going through early reactions from developers is that Kimi K3 matters less because it beats the top proprietary models outright and more because it gets close enough to make the difference feel almost academic for a lot of real work. Independent benchmarks put it just behind the two reigning frontier models and ahead of everything else, and enough developers backed that up in their own coding sessions that it’s hard to write off as marketing. The price gap is what will end up changing buying decisions, somewhere between a third and a fifth of what the top US labs charge for comparable capability. A handful of demos, including a full simulated desktop built from one prompt and an autonomous chip design run, added real buzz, though I’d treat any flashy showcase as a preview of what’s possible rather than what you get on an ordinary day. Hanging over all of it is a genuinely unresolved argument about whether K3 was partly trained on another company’s model output, something nobody has settled with hard proof either way.

A key complaint circles back to one habit, K3 thinks too much. It burns through reasoning tokens on questions that don’t need them, so the advertised price per token and the real price per finished task can tell very different stories, and more than one tester found a subscription quota gone in a single sitting. Self hosting the open weights sounds appealing until you price out the hardware, which puts it out of reach for anyone but large companies and specialized hosts, and the usage terms shift depending on whether you’re using the app, the API, or the eventual open release, its own kind of tax on anyone doing procurement. None of this settles who has the best model. What it does suggest is that the gap between frontier and good enough has shrunk to the point where the real competition is moving to whoever builds the better product around the model, not whoever owns the smartest one.

Gemini 3.6 Flash: The Cost Play, Not the Capability Leap

Unlike the last two models I’ve covered here, GLM 5.2 and Kimi K3, this one comes with no open weights to download, no option to self host, and no community fine tuning ecosystem. Gemini 3.6 Flash is API only, and that shapes how I read the reaction to it. Nobody is arguing this is a smarter model than what came before, and Google isn’t really claiming that either. What developers actually praise is speed and price for narrow, repetitive work, things like e-commerce listing classification, batch translation, and backend agents that run all day without much supervision. That’s a more honest signal than any benchmark chart, since most production AI spend goes toward exactly that kind of repetitive task rather than open ended reasoning. Google also seems to still hold a real edge in multimodal work, one person described automating a tedious ten to fifteen minute image cleanup task for a jewelry business, and others pointed to strong results on non English text and general image analysis. None of this reads as a frontier leap. It reads as Google doubling down on being the cheap, fast option for the huge middle layer of AI work that never needed the smartest model in the first place.

The complaints tell a more interesting story than the compliments do. The loudest one isn’t really about the model at all, it’s about Google’s own product plumbing, a confusing tier lineup, enterprise billing that lags behind personal accounts, and a coding tool people got pushed out of after a subscription tier quietly disappeared. Pricing has also crept upward across versions, with some developers citing a jump of more than six times on output pricing for the cheapest Flash tier since last year, enough to push them to move workloads to a competitor entirely. And the question hanging over everything is the flagship Pro model that keeps slipping, with outside reporting suggesting an earlier version got pulled back for underperforming the competition. Put it together and this looks less like a race for the smartest model and more like Google tuning a workhorse for its own economics, the sheer volume of queries it already runs through search every day, rather than for outside developers comparing full platforms. That might be a perfectly sound business decision. It is a different bet than the one OpenAI and Anthropic are making, and worth knowing before you build a workflow around it.

Open Weights Will Absorb the Volume Layer

Step back from the model by model scorecard and a clearer signal about spending shows up. GLM 5.2 and Kimi K3 look set to absorb the volume layer of enterprise AI, the routine coding, extraction, and tool calls that make up most of what companies actually run, well before either one challenges the frontier on the hardest problems. Developers are already running these models alongside Claude, Codex, and DeepSeek in the same workflow, using the cheap model for routine implementation and saving the pricier one for planning, review, or anything that needs real judgment. That isn’t a hypothetical architecture, it’s what people described building within days of each release.

None of that means the bill actually shrinks, or that the closed labs are in trouble tomorrow. The savings tend to evaporate once you count the extra debugging and hardware costs I walked through above, and cheaper tasks mostly just mean more tasks get automated rather than a smaller invoice. The longer term pressure is a different story, and it’s the same one I flagged recently: frontier pricing runs on a treadmill, today’s lead gets matched or distilled within months, and every release like this narrows the window a premium lab can charge full price for something genuinely exclusive. What’s more likely than an outright collapse is that the money moves rather than disappears, out of paying frontier rates for every routine step and into routing, evaluation, serving, governance, and the domain specific applications built on top of these models. GLM 5.2 and Kimi K3 are the first releases since then to make that case with real numbers instead of an abstract argument. That shift, not any single model’s benchmark score, is what will actually determine how AI spend gets allocated over the next year.


Related Content

Discover more from Gradient Flow

Subscribe now to keep reading and get access to the full archive.

Continue reading