Sounds great to me; live by the sword, die by the sword.
Sorry, OpenAI & Anthropic.
> what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models
If it’s as easy as that why do they choose to distill another model and not distill the knowledge on the open Internet from scratch?
A model trained on all knowledge from the internet (and other sources) is large but ultimately not very useful by itself, because it is going to spit out all kinds of garbage. You have to apply multiple further stages of training and refinement to the base model before putting it in front of users. So as an example you can train a model by yourself and then have GPT or Claude continuously check its outputs and correct it when it is wrong, ending up with a far more powerful model.
But barring the terms of service from forbidding distillation seems like a tough sell. OpenAI shouldn't be allowed to decide what types of customers it wants and doesn't want?
The terms of service don't even necessarily matter here. OpenAI could cancel your account for almost any reason, or for no reason at all. They don't particularly need to cite a ToS violation just as a store owner doesn't need to point to a written policy to kick you out of their store.
If the underlying issue is that LLMs should be regulated as a public good, then lets have that discussion. If it's that the major AI companies are becoming too powerful and anti-competitive, let's talk serious anti-trust enforcement. Micro-managing business policies isn't going to work very well.
You are pointing out that OpenAI can cancel user's accounts for almost any reason, and nobody can really force them to serve customers that they suspect are distilling their models.
That's one thing.
The GP is saying the government can make laws to make terms against distillation unenforceable. Without such laws, if you signed an agreement with OpenAI pinky swearing you won't distill, but turns out you did, you are liable in tort and OpenAI can sue you. (It seems nobody really cares about contract and agreements any more, but still...)
This is the other thing.
And I think you are both right.
OpenAI accused Deepseek of misappropriating trade secrets which could have serious penalties but seems like an awfully hard case to make.
Seems like we’d all be better off with a law that governs data sharing among AI companies, if that’s the policy goal.
"You're trying to kidnap what I've rightfully stolen!" -- Vizzini
Correct. It shouldn't be allowed to do that.
Wage spiritual warfare against the petit-bourgeoise. They all deserve it anyway, as they are the traditional harbringers of actual fascism.
It's also a bit of securities defensiveness. Pretending that you really do have a super moat, people just keep swimming in it so you just need to add more alligators.
It's farcical. Anyone who has worked on large models knows that the premise that an almost-Fable model was trained with distillation is beyond ridiculous. It's theoretically possible if they spent tens of billions of dollars on API calls, but it isn't the magic that somehow these people keep convincing people it is.
Previously Anthropic has reported on some Chinese firms doing chicken-shit level of API calls, that at most would be doing some Q and A or final fine tuning. The notion that they're training these models via it is fantastically ignorant nonsense that only very ill-informed and gullible people fall for.
This isn't the big gotcha some people seem to think it is, and the whole news cycle about that was mostly by people who have no idea what they're talking about. It's actually a meaningless data point. But it's precisely the sorts of people who think that a few thousand free accounts surreptitiously snuck off with Fable.
https://m.economictimes.com/industry/renewables/china-wto-co...
"Anthropic said the campaign was conducted between April 22 and June 5, 2026, and generated more than 28.8 million exchanges with Claude through almost 25,000 fraudulent accounts."
I don't know why you're trying to downplay it.
European models are so far behind because they don't resort to these tactics on a massive scale. Basically every other country is entirely dependent on 2 countries for frontier AI.
You may or may not be factually correct in your other points, but you're really proving the GP's point here regarding American exceptionalism.
Ignoring that I have literally zero trust in anything Anthropic has to say on this -- they have been doing the hysterical routine and trying to get every bit of government granted monopoly they can[1] -- those numbers still simply aren't that impressive.
>European models are so far behind because...
What a non-sequitur. Europe, like much of the West, foolishly delegated tech, media, payment systems, etc, to the United States. European efforts on this are poorly funded, poorly capitalized, and marginal efforts.
China is very much not Europe. China is looking to leave the US to the dustbin of history, and their efforts are a little more concerted.
[1] Surely Americans are aware that Anthropic and OpenAI are both very close to getting the US government to ban and fully criminalize the open Chinese models, right?
ToS is just conditions that you agree to in order to use a private service that is provided at-will. I can have a private coffee shop where the terms of service are that you must wear red to enter, and if you're not wearing red, you are not welcome on my property.
So it would be upto OpenAI and Anthropic to enforce them on their own terms (by banning accounts and IPs).
We are discussing Chinese models. Now look at how much foreign competition the Chinese government prevents in their domestic market in other industries.
Distillation is just value extract.
It's soft, and I'm not sure what the answer should be ... but I think that there is a difference.
I think we start by recognizing that ... and then try to figure it out from there.
'The Internet' may be a public good, maybe we make them pay a tax for that, but that's different than distillation.
There is a value-add in selecting the valuable parts out of the garbage. And let's face it. Largest models contain a lot of garbage.
We ought to identify that and integrate that into our thinking.
Literally the biggest thing of our generation - AI - is the living embodiment of that 'value add' writ large.
'What is the difference' - is the AI you use all day, in comparison to 'all the world's data' you can use for stuff and do 'whatever' with it, but are not likely to come up with something hugely useful otherwise. Maybe, not likely, if you did, it would be 'value add'.
Copying something is not.
Programming Microsoft Word is value add, copying the code is not.
Copying design ... there are some question marks there.
It's extremely easy to understand at it's core.
What makes it hard, is that faux intellectuals like to deconstruct ideas at the margins, and have those critiques stand in for reason.
"At sunrise the sun is only 'half there' ... there fore there is no 'day and night' just a blur! Day and night are the same thing!"
The training data used is part of all of this is a separate but related question.
Felony contempt of business model.
What I'm saying is, doesn't the law already cover 1?
In the US one of the factors is “ the effect of the use upon the potential market for or value of the copyrighted work”.
If anthropic Hoovers up the world’s books and trains on them, and then spits them out verbatim on command, then it will clearly impact the value of the work; nobody will buy the original, they’ll just ask Claude.
Others also argue that even if it’s not reproducing it exactly that the training runs afoul of that factor, specifically the “market for” portion. A rights holder can no longer license their book for training of LLMs if Anthropic goes ahead and just trains on it anyway.
Ah, right. So if we want models to be capable we need them to be trained on as much as possible, yet we also want to stop what you described. So what can be done?
So from my perspective, it's doubtful that this is the moat. Besides, for example, Claude Code in particular is so buggy (and always has been).
true for anthropic, not true for openai.
He talks about this in another recent essay https://stratechery.com/2026/anthropics-safety-superpower/
> If you own the user touchpoint, then you have meaningful lock-in, and the best way to own the user touchpoint is to be the canvas for everything they need to do. This, by extension, means that the frontier labs are on a collision course with software companies: it’s software that owns the user touchpoint, and it’s in the frontier labs’ long-term interest to not simply be a commodity input into software but to simply replace software outright.
For this reason alone I would also argue that the idea about an agent harness being sticky is a non-starter long-term.
i also expect you'll see markedly different results if you constrain yourself to small models. there even trivial harness improvements like Codex's /goal feature, and more capable basic tooling (e.g. semantic code grep, js-capable `fetch` tooling) make or break the actual task success rate.
I'm amazed that no one is talking about proposals that are surely being discussed in Washington and pushed by SV lobbyists to restrict Chinese models on national security grounds, or other some other basis.
The belief that Bytedance could engineer a finger on the algorithmic scales to serve the interests of the Chinese Communist Party led to a lot of debate in Washington, and ultimately resulted in TikTok being divested from its Chinese owners. Huawei is shut out from the U.S. market, which limits its business even in markets where it's not banned because it's effectively stamped with a scarlet letter.
IMHO, Chinese models are headed for a similar fate or at least a showdown in Washington or the courts because they are supported and/or controlled by entities which ultimately serve the CCP.
Personally, I think models will increasingly become specialized in different areas, some good at X, others good at Y, and we might see workflows that mix multiple models.
"I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks" what risks?
I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. I don't understand this at all.
Can someone in the know please use plain layman's terms to explain what this tweet is about?
I think it's referring to the belief that LLMs are not the path towards AGI, and that LLM's, while useful, are not going to have the impact that the American labs believe it will have.
The Silicon Valley people like this openai guy, high on their own supply, are convinced they are building some machine god that will either bring about the end of the human race or utopia, they therefore cannot understand why the Chinese (or any other normal person on earth) are not afraid of chatbots and have other things on their minds.
I mean, what an appalling way to live. To take themselves so seriously and simultaneously be terrified of what they're building. If we are really seeing the collapse of the bubble now, I wonder what's going to happen to these clowns when their AI god fails... what will they move on to next?
I assume they mean the risk of opening up "forbidden" knowledge to the masses without adequate control, which the CCP hasn't historically been known to do.
> I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). Confused what this means
Yann Lecun is a pioneer in the field of AI and Meta's former AI head. He is famously anti-LLM, and considers the entire technology a dead end to achieving human-level AI. The author is saying the CCP has similar views (that LLMs aren't going to get exponentially better/lead to AGI) which is leading them to not control these models as tightly as they otherwise would.
> Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. Confused again.
"AI accelerationists" = people who want AI to progress. According to the author these people should not celebrate open models because open source = less commerical value in LLMs = less investment into the field (because how are companies going to get returns?), and this will ultimately lead to slower growth.
The last bit is about government controlling AI vs commercial companies. According to the author the former is a dystopian hellscape.
IMO even if you think his points make sense, his job title ("head of strategic futures @openai") means they should all be taken with a massive grain of salt.
Writer seems to have no clue how IP actually functions in China
I’m struggling to understand this perspective. Is he using the words accelerationist/decelerationist in a sense other than the obvious one?
EDIT: I searched his twitter history and discovered that his argument is basically “if you drive down costs, then OpenAI will have less money to invest in development, slowing down the overall rate of AI progress.” IMO this take betrays an overwhelmingly stupid degree of exceptionalism, but I guess that’s what I’d expect from someone working at OpenAI.
Every company wants that!
if open models are good enough then it doesn't matter, if they aren't then there's probably a return available in investing there.
how is running servers supposed to be 0 cost, while running ai inferrence isn't?
A single response from kimi k3 requires hardware that cost between 500k and 1m dollars up front and draw over 20kW. Each request costs at least 5% to 10% of the charged cost.
My experience has been quite the opposite. I was using Claude Code almost exclusively this winter/spring and swapped to Codex earlier this summer. It took no time whatsoever to switch. And before Claude Code, I was using Cursor. Same story.
[edit: Oh and there was also a brief interlude with Conductor, though I think they're more or less just serving the underlying Claude/Codex harness]
Probably the only reasons I would seek change are economical.
“I don’t really have a strong preference between the two” is another way of saying “the product isn’t sticky”, which is another way of saying “this provider has very little room to increase margins”
There’s little difference between Coke and Pepsi and the barrier to switching is nil, yet clearly the products have stickiness. People have slight preferences and become familiar with the brand and then engagement becomes habitual.
The effects on margins are irrelevant to this.
For companies, these decisions are very sticky. Companies go through a lot of red tape to get anything purchased and approved, then they discourage change because it's a lot of work.
So the product that gets a foothold in a company sticks for a long time.
Then a couple years later a sales person convinces an exec that they can save some money by switching, so the switching game begins. Not necessarily motivated by the better product, mostly the price. My wife's company keeps switching their tools out from under everyone every year or two. Just when they get everything stabilized and everyone familiar with the new tool, some new contract is signed that moves them all to some other company's suite.
Similarly, I have made no ground in arguing to try to get Codex at the company I work for, which got Claude Code a year ago and sees no reason to go through the whole process of setting up any alternatives when Claude Code already works and is at the frontier.
You can choose a selection of different models within it, but you're not using Codex or Claude Code.
You install MCP connectors, specific skills, work around model/harness quirks, set security boundaries etc.
It's a lot of work, and most people will never want to change it once they have it working.
They will have people who don't understand the distinction between visiting Claude.ai and downloading Claude Cowork.
They type the words "setup MCP" into Claude.ai and expect it to automate Excel on their machine.
There's a pretty big gap between the things we talk about here, and where the world is at.
its so distracting seeing these types of confision.
every plugin is already just multimodaling their targets.
The cost of me moving around these different AI models and harnesses was pretty much 0.
They communicate through my own harness, and it's working pretty well so far. claude code is being overtaken by codex however because I noticed lately the accuracy of the latter is the best.
This has an element of stochastic improvement so it's hard to predict but the chance of the U.S. "winning" this "race" is pretty bleak.
You see this all the time in communities that have internalized hierarchy as a "good", little kings of shit mountain vying for less and less at a higher and higher cost.
An astute Chinese analyst could reasonably forecast that they had little chance of controlling the AI market due to sovereign trust issues, but would also note that AIs are just software.
When the dust settles the US still won't have factories, and the real value of AI models is still going to be embodying them and getting them to do real, consumer facing work.
Perhaps the most striking thing about the AI boom is how quickly the US abandoned the veneer of local manufacturing in favor of more expensive buildings producing nothing you couldn't make anywhere else on the planet...from imported parts.
China has a billion+ people that their AI can "study". Plus due to China's political structure, their AI has access to everyone's chats, comments and sites, scraping everyting.
Here in the US, with 1/3 the population, the AI race was lost before it even began. Plus in the US, all companies and people are doing all they can to restrict AI from scraping sites and peoples chats.
So I believe, China will end up owing AI.
I think you hit the nail on the head - right there!
for VCs, breaking even is losing
They’re going to try their best to offload these investments into our pensions before the inevitable crash.
https://finance.yahoo.com/markets/stocks/articles/goldman-sa...
Too many people here buy high and sell low.
When Goldman makes statement, half the time it is prepped by an associate or two that has minimal experience and goes through a MD that enjoys the gloom and doom. That’s why they publish. Goldman makes money on both sides of a trade.
- Commodity market profitability is determined by marginal cost of production. LLMs have marginal cost; traditional software does not.
- Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
- The highest tier Chinese models are not more economical than US frontier models. Try GLM 5.2 and see how much it costs to do real work. I did, and it was more expensive than GPT 5.6.
- This is because US labs are leading on cost efficacy of inference ($/task)
- Training will decline as a percentage of costs as inference expands compute share due to agentic workloads. A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.
- With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
OpenAI really shows the way here. Their cost per task is less than half that of Anthropic because of more efficient tokenization and less verbosity. OpenAI is both cheaper and better than Chinese models for frontier work.[edit]
This is my observation from using it without an specific context engineering to optimize for Deepseek's cache compression and sparse attention mechanisms. I am pretty sure that if you specifically structure your context to align to the cache compression boundaries you can significantly improve performance in the full 1M context, but there is not much reason to do this, because if you design your outer loop to work with shorter contexts that solution is portable and more efficient, so I haven't bothered with a optimizing for DS at this point.
Switch to a modern sampler like min_p or ideally a better one like top-n-sigma (it’s in llamacpp) and your “my model gets stupid at long context problems” will basically go away.
Unfortunately this fact is still not well appreciated yet despite nearly every modern sampling technique getting an oral wherever they get presented. Min-K just got an oral at ACL 2026, for a hyper recent example of this. There’s a reason they keep getting orals.
The field massively ignored sampling for mostly safety reasons and now the whole field incorrectly believes long context doesn’t work on small models. Long context is an out-of-distribution problem. Your sampler configured properly keeps you in distribution.
Oh and this is doubly true for quantized models. I run my qwen 3.6 27b with 4bit quants from unsloth and get excellent performance because my sampler stack is good and not the garbage that is top_p and top_k. Also, yes, you need to ignore the trash recommended sampler settings from the Chinese labs (they’re wrong/bad).
I appreciate the transparency in explicitly stating their motivation for writing the article (a response to what the author saw as an overreaction to Chinese models), but I feel the article goes too far the other direction, with multiple unsupported leaps of logic, and overstating the stickiness of AI client products.
> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence. [emphasis mine]
I guess I'm missing the part of this article where they bring hard numbers in to back up the argument here. What work was attempted? https://cursor.com/evals shows the previous generation of open models (Kimi K2.7) trading blows with the others, cost effectively. Composer 2.5 is itself a fine-tune of K2.7, and it's apparently quite token efficient, so why would it be impossible for a Chinese lab to achieve something similar? GLM 5.2 Max is also ranked above the lower end OpenAI models and is not far off in price.
It's weird to have this entire discussion about tokenomics without mention of the circular financing and debt raised by labs in the West, which can then essentially give away their capacity to end users. OpenAI giving away quota resets to subscribers like candy on Halloween while their compute partner Oracle's bonds is reevaluated to be one grade above junk? How?
I don't think you can make an argument about the future one way or another by arguing using the listed prices. The math is not internally consistent enough for it.
Because you're comparing retail price whereas the parent commenter (and the article) is talking about marginal (ie. inference) costs. American labs are providing a premium product and they're charging accordingly. Meanwhile for chinese models they're open weight so they're limited to how much they can charge without competitors undercutting them.
If we use tokens as a rough proxy of inference costs (rough approximation, I know) and look at artifical analysis benchmarks, you see that all the open models are behind the pareto frontier in terms of efficiency.
But if we have to look at what we think margins might look like, DeepSeek continues to host v4 Flash at the existing price despite competitors beating it in price (https://openrouter.ai/deepseek/deepseek-v4-flash), so there's at least one example of a Chinese lab charging a predetermined price despite competition. And no one but Moonshot is hosting Kimi K3 yet (https://openrouter.ai/moonshotai/kimi-k3). Perhaps there's room in the market for those who release their models to make margin on them.
And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
their competitors are discounted at around 33%, so it's safe to say that's the margin, maybe less if their competitors have worse caching or quantization. Meanwhile claude code/codex resellers selling tokens for 90% off API price, presumably by reselling usage from fixed consumption plans, which gives an idea on how fat the american labs' margins are.
>And I believe my Composer example speaks for itself. The open models are behind but there's tangible proof they can be tuned for pareto frontier efficiency. See "Cost per Task" at https://artificialanalysis.ai/agents/coding-agents.
But composer is a closed model? If it's really that easy to get better coding performance, why haven't the chinese labs replicated it? And this is all assuming the performance boost is real and not from benchmaxxing. Moreover if you apply the "street price" discount I mentioned above, American labs look far more favorable.
I look at that and think that they must be losing money hand over fist on something like this, not that this shows what their margins are like. If their margins are like this then I don't see why they'd be raising money and shuffling it around in circles.
> If it's really that easy to get better coding performance, why haven't the chinese labs replicated it?
Nobody said it would be easy! I just think it's possible, and that presumably they will get around to doing it at some point.
For an example of my real token usage for a day with DS: Input (Cache hit) 530,949,760, Input (Cache miss) 7,875,004, Output 1,389,685 - it is still 1/5th the price of Baidu (the cheapest) and 1/7th-1/10th the price of US hosts.
Also, Deepseek platform is not the same thing as Deepseek open weights. There is a major misconception that the existence of an open weights model means that it is the same thing as the proprietary platform offering, but that is definitely not the case.
Anthropic’s API pricing is getting impossible to justify. Anthropic previously had the highest quality models, and used their position to charge premium prices, enjoying inference margins of over 70% [0]. They could charge these prices because no other model came close.
But over the past month, the market has shifted dramatically. Over every single performance tier, Anthropic is being squeezed on price.
* Low end: DeepSeek V4 Flash runs at ($0.02/task), Xiaomi's MiMo-V2.5-Pro at ($0.03), and Haiku at ($0.24). Anthropic is ~10x more expensive than the Chinese open-weight options.
* Mid tier: Claude Sonnet 5 ($1.53/task) is nearly 50% more expensive than GPT-5.6 Sol ($1.04), nearly 2x the cost of GPT-5.6 Terra ($0.82), and 3x the cost of GLM-5.2 Max ($0.47). There is basically no reason to ever use Sonnet 5, the competitors are significantly cheaper.
* High end: Opus 4.8 ($1.80/task) and Fable 5 ($2.75) are the two most expensive models, and GPT-5.6 Sol ($1.04) and Kimi K3 ($0.95) offer comparable performance for significantly less. Less the fact that Kimi K3 will get ~10x cheaper once its weights are released and served on neoclouds with Nvidia hardware [1].
OpenAI priced their latest GPT-5.6 models cheaply in order to regain market share. When Anthropic clearly had the best models, their 70%+ inference margins were defensible. But today they are the most expensive option in every single tier. Unless they make significant price cuts soon, they run a serious risk of bleeding market share.
[0] https://www.mindstudio.ai/blog/anthropic-inference-margins-7...
[1] "American companies such as Modal, Fireworks, and Baseten will be able to serve Kimi K3, at one-tenth the cost of their Chinese competitors because they have access to advanced Nvidia hardware" https://x.com/rohanpaul_ai/status/2079027313455550839
This is a flatly false statement for most things powering backend applications. The AI consumer "doing real work" model, either for analysis, chat, or coding could well be more cost effective with closed frontier models.
But most of these internal glue business SaaS applications where engineers are integrating are not those tasks. It is those tasks which 1) drive immense amount of domain-specific data into the platform over time, and 2) are most encouraging of driving open model independence with no vendor lock-in.
Anyone on this site who has actually used ML models (more accurate in many cases) knows there's a lot of kludge that simply does not need a 5 minute agentic feedback loop to solve the problem. And they were solvable a year ago with lower class models. The token economics are exceptional and the anecdotes of a16z saying 80% of startups are productionizing open models is only surprising to people who think running your company on OracleDB in 2026 is a sound engineering decision.
It may not be false but may be a "category error" [0]. Reserved GPU pricing & bulk inference pricing is 3x to 6x cheaper than "API rates", but renting your own GPU cluster (in this crunch) to run a 600b+ open weights is going to be "more expensive than GPT 5.6".
Even then, it remains to be seen if Huawei will pull their weight (and match up to Nvidia) as spectacularly as their fellow Chinese AI Labs have. If so, the WAICO alliance is ready to go all-in.
[0] Ben, and probably other "influencers" in this space, may be prone (knowingly or unknowingly) to favour points that make their conclusion for them (https://en.wikipedia.org/wiki/Motivated_reasoning).
But much like Ben's point that commoditization is a relatively novel concept to many in tech, it's not the consumer AI applications at risk of commoditization. They have distribution there.
It's the literally millions of engineers who are updating codebases with tools replacing workers partially or wholly. It's the supply-side where there's compression, and no need for distribution.
I would argue, given the enormity of the existing SaaS stack and how it integrates with the human machinery of personnel, that's where volume is. And that is clearly cheaper and a home run.
Commoditizing a ~$100B AI consumer market is no small feat. Commoditizing 20% of the $500B SaaS market, to say nothing of the underlying systems in the who-knows-how-many trillions "Big Tech" market (you're obligated to say that like the Kool Aid man), is shocking.
The labs are not interested in the small, fast, single purpose end of the market. Google increased their pricing on Flash so much that it stopped becoming a cheap model; instead, they released Gemma 4 open source, which is actually easier to use from a third-party inference provider than from Google.
From a total token volume perspective, these "utility" models (classifiers, simple summarizers, small OCR models) will absolutely drive enormous volumes of tokens, at low prices and margin and modest overall market size. Because the models are small and the performance requirements are modest, and because their use cases are specialized rather than general, there are poor economies of scale: they can run cost effectively on rented small GPUs, and a big player doesn't get a structural cost advantage. These models are usually 1b - 30b in size, and can run on a rented 5090. I've productized these myself: I run millions of pages through a fine tuned 1b OCR language model that runs on 5090s at a cost far lower than commercial providers.
But that's not the segment of the market where GLM 5.2, Kimi 3, etc., play. They compete with frontier capabilities, and they are not particularly cheaper than OpenAI models at a cost per task. (I do actually think they compete well with Anthropic, because Anthropic's model efficiencies are poor compared to OpenAI.) And although this part of the market may not be the bulk of the token volume, it is the bulk of the market value.
That's because a lot of human knowledge work is too generalized and fuzzy for dedicated, fine-tuned models, so they are almost entirely different markets that don't particularly compete with each other. (Though if SaaS companies successfully build around verticals that can use small models applied against well-defined jobs, there may be opportunity to push the small/big capability boundary to subsume marginally more valuable tasks that today would require mid-grade reasoning.)
> Models are not free. Downloading them is free. Running them is not. This has manufacturing economics, not software economics; the idea that they are "free" is an economic category error as it relates to their actual use
I think a large part of manufacturing economics is illiquid overhead and the cost of expertise to set up and run your manufacturing line. Compute economics don’t have the same illiquidity nor do they require the same expertise or even specialized infra (current temporary chip shortage aside).
The implications of this are small players (e.g. your uncle running an inference server out of his garage) have comparably efficient marginal costs as big players. Compare this to actual manufacturing where small players have essentially no access to the manufacturing facilities of the big players.
Additionally, big players with a lot of compute who are not meaningfully in inference today (e.g. Amazon) have a fairly straightforward glide path to utilizing that compute to compete.
> This is because US labs are leading on cost efficacy of inference ($/task)
It’s possible, but I would need to see better data on this.
>A big part of training now is optimizing token efficiency. It's hard to distill token efficiency; that is perhaps why Chinese LLMs are so inefficient.
I think it’s fair to assume this is true, but also token efficiency is not a meaningful competitive moat. It’s not like these are secrets the Chinese will never figure out, it’s a fairly active research space and the outcomes are quantifiable.
This is where Oracles "datacenters for rent to run your own Chinese models" strategy will benefit. The LLM SaaS game is a lost one thanks to China.
As soon as manufacturing starts building this stuff more, it will commoditize. The hardware prices won’t be terribly larger than the original. We’ll have a “Bambu labs” style company to make the AI OS, whatever that is.
Eventually we will hit a "good enough for cheap enough" and frontier models will hit diminishing returns (if they haven't already for a lot of types of work)
Don't think the rest of the world will sit on their hands while the US soaks up chips either, demand gets filled and if the US won't fill global demand for chips that's an opportunity to undercut again.
The other thing the rest of the world doesn't have to fund is the ridiculous valuations on these companies.
Unless you think the US can stay ahead just with model efficiencies, and that no one else will eventually match them, you are looking at the writing on the wall.
All that to say, the rest of the world is more than willing to eat your lunch, they have a dozen good reasons to, and they're already showing good results.
Just on the economics side, we've been here before too, US companies typically export their commoditization and live on brand royalties. Think all the cheap manufactured goods, the US doesn't make any of it. That's because the US can't compete on margins for numerous reasons, it's too expensive, I don't think AI is any different here except that the brands are currently valued in the trillions and I suspect that greed will be their undoing.
If I had to wager why, I'd say it's due to embracing solar on massive scales recently. Only a few years ago the US was competitively priced.
[0] https://www.globalpetrolprices.com/compare_countries/USA/Chi...
This can change quickly though, so it's not that big of a deal. If AI energy demands push the Gov to deregulate/fast track new plants, or the industry decides to build out their own generation renewables.
1. US frontier lab unit economics are better 2. US frontier labs are moving up the stack making tools that are "stickiness" and will prevent users from switching.
For 1...he doesn't provide any evidence for US lab unit economics being better...the major input to unit economics is electricity...which is cheaper in China. And building data centers and connecting them to electricity is both cheaper and an order of magnitude faster in China. The main input that US labs might have an advantage in is in cost/access to chips, but that given the level of chip investment in China it seems unlikely to hold.
For 2...there's little evidence these tools are sticky. At least in programming, the trend seems to be tools like opencode that support multiple models and providers.
And even when they are sort of sticky, as we know on hacker news, people figure out how to point the tools they like to competing models even when the app doesn't official support it.
And every improvement in model capability makes it increasingly easier to make your own tools.
Wrote more on this in a blog post that has an earlier HN discussion: https://news.ycombinator.com/item?id=48982061
Direct link: https://larrysalibra.com/ben-thompson-is-wrong-us-frontier-l...
What matters most is $/completed task. It does seem like OpenAI and Anthropic are winning here even with worse electricity rates. Perhaps it is made up by the efficiency of Nvidia and Broadcom chips, which China can’t get in mass.
I do think that OpenAI and Anthropic are moving up in stickiness. My company has rallied around Claude. We are customizing Claude Code, adding knowledge bases for non technical people, writing skills for them, using Claude features company wide. It’s hard to move.
Meanwhile, I personally use ChatGPT outside of work. The memory, ease of use, habit keeps my subscribed.
Across sectors, China added 543 GW of energy in 2025. Next year, USA is expected to add between 70 and 80 GW of energy
AI is really all about electricity. AI could be completely fake and yield zero value whatsoever and the US would do exactly what it is doing now because the AI bubble is what creates the market for building new electrical generation capacity, which is needed for re-industrialization. Also why our friends in UK/Europe/China are so busy pushing anti-AI propaganda to try to undermine this.
They do not want to life in the world that is and thats going to be, but in the past and the world they green ideology promised. Reality denial be a addictive poison.
AI is not fake and it does work, but what I am saying is that from a pure systemic analysis perspective, you can do the numbers, and even if AI was complete fugazi, the benefits you get from the electrical generation capacity, and the ability to fund it through private markets, which bypasses Congress, and locks in commercial contracts (often with foreign governments) which will be almost impossible politically to reverse, would still make it optimal from a strategic perspective. That is my calculation, and to the extent that it is correct, I would assume that the US Military's strategic planning apparatus would arrive at the same conclusion.
AI compute has some unique characteristics that make it especially useful for grid management. Moving consumer compute to the cloud means that the electrical use of that compute can be centrally managed. In an emergency, you can cut electrical use for consumer AI by 50% or more, because chips run more efficiently at lower power, and you can shift workloads onto quantized models, reduce resolution for video output, etc, to reduce compute, which leads to minor service degradation but not interruption. AI datacenters are also adding massive amounts of battery storage capacity, which is an additional grid buffer. For every GW in capacity added by hyperscalers that is creating a dispatchable reserve capacity of 50% under completely normal circumstances (hyperscalers do this internally to optimize their own costs) and then that number goes up depending on the scale and duration of the emergency.
Which is to say, this isn't really a lock-in/ stickiness vector (unless maybe the wording itself of a skill is hyper-optimized for a specific model)
However with the latest models Fable, Kimi K3, 5.6, it's getting to a point where I sometimes forget what model I am on without noticing a difference. And once I realize it because something may not be exactly like I expected it I won't switch for that work either because I don't want to invalidate the cache.
For the next work I will do there is maybe a 50/50 chance to remember to switch the model before I start.
That's not what I would call stickiness towards a certain provider.
- Refactoring a 13 year old in-house vacation rental booking system ( python/turbogears )
- Backend development for our VR fitness game ( flask/python )
- Unity development on our VR fitness game ( C#/Unity )
- VR game development experiments ( Godot/GDScript )
- Standalone SLAM localization service ( C++ )
- Audio analysis ( python/pytorch )
- Virtual display with Viture display glasses ( C )
- Reverse engineering a library I am using for another project ( ghidra -> C - no MCP yet, that's something I am looking forward to )
- Public facing website rebuilding for the booking system above ( PHP/JS )
- Generative 3D environments for our VR fitness game ( python )
- Wireless camera/IMU based tracker for the SLAM system ( C )
Once I've dug in with a specific model into a problem I tend to stick to that because I have a feeling what it will do and how well it works, but when I start a new thing I usually use whatever the model was last set to.What is the cost of AI? The single largest ingredient is Nvidia profit margin.
Huawei accelerators are not as efficiency yet, but they don’t nearly extract as much margin.
Why would future revenue stay with the labs given this situation? This whole thing had an airline industry sized red flag on it that makes investing into frontier lab about as sexy as investing in United.
Maybe the token economy is some kind of reverberation of the airline reward miles economy, the emergency hatch to be able to survive under maximal supplier extraction (Nvidia is just the top of a monopoly stack here, even if they replace those chips, the HBM, ASML, Foundry layer can get their dues)
Sorta yes, sorta no.
A single 5090 consumes 450W - at Californian energy prices of $0.38 per kWh that's $0.17 per hour. And the card itself costs $4100 on amazon. So after 2.75 years running at full power 24/7 you'll have spent more on electricity than on the card. I would have thought most data centres being built today would have a design life longer than 3 years.
Of course you can throttle the cards to ~300W without losing too much performance. But also you need more than a single 24GB card to run most modern LLMs.
> 132kW
> ~$3.7-4M
So about 300x the 5090's power but 1000x the price. Roughly 9 years for electricity to exceed price at $0.38 and datacenters will show up in areas with cheaper power than CA.
LLM's aren't very latency sensitive and can therefore move to wherever power is cheapest.
Right now that's places next to aluminium smelters (which also like very cheap electricity 90+% of the time).
That's not generally true, since there is generally still much reliance on NVIDIA. The true low cost providers are Google with their TPU and vertically optimized stack, and Amazon with Trainium. However, Google does not have their own frontier model, and Anthropic (who are partially served by Amazon) are also paying a premium for extra NVIDIA-based capacity from SpaceX, maybe soon from Meta too.
I don't know how the economics of domestic Chinese Huawei-based clouds (no NVIDIA) compares to the west, but since serving cost is mostly hardware depreciation and to a lesser extent electricity, they are not necessarily at a disadvantage (Ascend 950 costs roughly 50% of an NVIDIA H100), and more to the point it is irrelevant when considering US commercial use that is more likely to be using Chinese open weights models from US providers served on NVIDIA based hardware.
I think the real significance of Chinese frontier models being open weight is that it takes development cost amortization out of the US-based serving cost, while the US AI labs can't afford to do this. The US labs therefore need to reduce development spending to remain price competitive. The Chinese companies are of course still making money from the Chinese market, whether by selling API access or by other business models such as Ziphu making 75% of it's total revenue by selling services to Chinese customers who are running their models on-prem due to the Chinese apparently being very concerned about data privacy.
I notice that the article, and this discussion, hasn't mentioned or considered local models.
We can already run a low-spec model on a laptop. Because there is demand for this, it will improve and we will get better laptops and better local models. We will also see models being run on dedicated local hardware and called from the laptop.
If I can download a reasonably capable model to my own hardware and run it without paying anyone for either the model or the inference tokens (effectively making models and intelligence actually free once the hardware is bought) how are the Frontier AI Labs going to make any money at all, let alone enough to support their vast valuations?
Is this really different from traditional software? Downloading postgres is free. Running it is not. You either buy hardware and assume the costs of owning and running that, or you pay to run it in the cloud.
IE, a lower param OpenAI/Anthropic model can compete with a higher param open source model.
So even if you are an American company who downloaded Chinese models in hopes of saving in cost, you still have to beat OpenAI and Anthropic in $/task which is very tough to do over the long run.
Over time, this compounds. More profits means more investments. More market share means more control.
There being squeezed by their own stock pumping and SpaceX pretending to be an AI company is driving down the exit strategy. I’m guessing one starts going full Theranos and begins claiming full AGI or gets the US government to government cheese then hard. It’s going to be a few crazy months.
This is the story for Nvidia/AMD or cloud providers rather than OpenAI.
> With increasing inference as % of total compute, if labs create efficient models -- which they can, because they can create highly optimized models amortized over very high inference loads -- they can be low cost producers, and be competitive at $/task rates
It seems like there would be problems with this on both ends.
For general purpose models, everybody is trying to make them efficient, so you can't win just by being slightly more efficient. You would have to be so much more efficient that you can charge high margins while still capturing the majority of the market so that the high margins get multiplied by the majority of users and the users you leave on the table aren't funding open competitors. Meanwhile everyone else is also trying to improve efficiency, so one misstep and you're behind.
Example of where this can be a problem: You spend a preposterous amount of money to create an efficient model, then someone else publishes a paper with a new technique that gets a similar but incompatible efficiency improvement out of a model that costs a lot less to create. You have now spent an enormous amount of money in exchange for no competitive advantage.
And from the other end, one of the best ways to get efficiency is through specialization. A general purpose model can generate code or summarize a meeting transcript, but a special purpose model can do it as well or better with far fewer parameters and resources. But then you don't have a situation where one huge AI company has The Most Efficient Model, you instead have dozens of specialized models produced by independent sources that are each the best in a given niche. Any proportion of which could have open weights, or have an arbitrarily small advantage over the ones that are.
Moreover, these problems combine: Both the computing hardware vendors and the AI companies want the margin on doing inference, but the more of it one of them gets, the less the other does. If the AI companies were actually getting huge margins then it would be in the interests of Nvidia, AMD, Apple, Intel et al to fund efficient open weight models in the same way they fund Linux. Commoditize your complement. And those models don't even have to be better, as long as they're good enough that the closed models can't charge a significant premium and the margin shifts back to paying for hardware.
I could see AI ending up the same way where the customer captures most of the value rather than the companies. Open weight models are what make that kind of competition possible.
I'll agree that GPT 5.6 may well be the best given the above contstraint, but for run-of-the-mill dev tasks (real ones, not benchmark ones), GLM 5.2 still blows every other model out of the water.
Cost per task as a metric is a bit ridiculous because there are so many types of tasks. GPT-5.6 can do some tasks GLM could only dream of, but GLM can do some tasks 100x cheaper and better than GPT-5.6.
But its also a 800B sized model running on a ram constrained system with no GPU.
Techniques, GPUs, more ram, and faster disks can always speed it up. But the point being is they run on low end machines now. Its now an optimization problem, not a possibility assessment.
I don't need 2.4T to do that; I'm doing it with 35B or 27B. If they get me a model in ~80B with a A5B or A7B, that will be the end point.
It's bizarre people, by themselves, believe all these parameters are getting them much more.
Lets be serious: if we as a civilization really wanted the advancements promised, we'd find the 1000 best scientists and give them free access to these models while the rest of us get personal GPUs for specific use cases.
But instead, we have to endeour this penis measuring contest for the infinite bikeshedding of the universe.
(And that's good, that's how it's supposed to work. See eg prices for transistors or hard disks or solar power over the last few decades.)
Its the user base (with ads and upselling) and proprietary wrappers which will make money for typical customer.
Even enterprise customers arent going to be spending a lot on tokens. Once labs no longer have to subsidize trainings tokens costs will drop 10x and once models get burned on chips costs will drop 10x more and you physically won't be able to burn significant number of tokens unless you're deliberately trying to.
Correct. These chinese labs has proven that having just the model is not a moat, and the safety concerns were all just attempts at regulatory capture.
This is why labs like OpenAI and Anthropic are panicking and are racing to the exit before their valuations start being questioned.
In the end what companies pay for is not tokens but results. A DIY kit of models, mac minis or whatever, and a bunch of poorly integrated OSS tools doesn't solve their problem. For the same reason, people use Office 365 rather than running Libreoffice. And for the same reason things like AWS dominate the market rather than people DIYing their infrastructure together themselves. Most of the money is in polished turn key solutions. Which is what Anthropic and OpenAI offer.
The juicy market here is the enterprise market. That's mostly business users, not programmers. They'll be hooking up all their SAAS tools (which they also over pay for), and other stuff. They'll be paying for boring things like data residency, compliance, etc. And they need access to reliable infrastructure to run all this stuff. They'll want this shit to just work and not to be dealing with a lot of poorly integrated stuff.
Most of the billions invested are being sunk into infrastructure, chip design, and access to resources (land, water, energy) needed to run data centers. A handful of companies now own most of that infrastructure and they also happen to have the top models, researchers, the best tools, and warm customer relations. And they sell access via very convenient subscriptions with high enough limits that people don't have to worry about things like token cost. The game here is recurring revenue from customers that like predictable pricing, reliable quality of service, and iron clad compliance and data security & residency, and quality guarantees. These companies don't want to be chasing model quality and have to upgrade their entire company every few weeks. They want continuity and predictability. Mostly they just pay Anthropic, OpenAI, MS, or Google to take care of this for them. There might be some niche EU players that become a bit bigger. But I don't see a large scale switching to Chinese suppliers for a full polished alternative. The Chinese might give away their models. But I don't think they'll be generating a lot of revenue.
And if you want to run your own models, you'll still need infrastructure to run it. These four companies together with the usual cloud giants control most of that and as well of the supply of resources (chips, data centers, energy, etc.) in the EU and US markets. There's going to be a long tail of self hosted and gobbled together stuff but it's going to be a much rougher experience for end users and it won't likely be most of the market any time soon.
> open weights, open code and open data
Even if you have all these things you still can't replicate a model because of randomness.
You can backdoor a model with less than 1000 examples and it is impossible to detect.
You can take the code for Kimi K3 now, take the training framework from Prime and the data from Olmo, spend some money on RL environments and some more money (!) on GPU training and end up with a system of similar capabilities.
But that's completely different to being able to audit Kimi K3. Even if you had the exact code, data and training environments it is impossible to verify that the model you have came from that.
They just don't work at all on a many month long, 100K+ GPU cluster training run.
Even then you'd still need to account for order of events when an entire cluster of GPUs is involved. Also don't forget to account for any synthetic data sources. Or even non-synthetic for that matter - does your pipeline do any image resizing on the fly? Better make sure that's fully deterministic between machines (it almost certainly won't be).
It's theoretically possible but I don't expect it to materialize any time soon.
Even with this, the cost of verification would be enormous. You would need a massive cluster to repeat the training E2E.
It's when the vendors and/or governments in charge of Model A decide that I'm not allowed to do that, that I have a problem.
Of course, one could retort that gathering that evidence may be nearly impossible now, but my point stands: in the future it might/probably will be possible to properly audit open-weight models. Closed models, on the other hand, will always be a black box.
But it won't change after you download it, so you can isolate those problematic cases and use another model for different use cases
There are also half a dozen other companies from China continuously hammering our clients’ websites.
I was wondering, what's in that cold dessert? Low and behold satellite imaging shows massive datacenter build outs, very cheap solar energy.
Few months ago something happened and the Geo location on data on those IP now shows "Shanghai" or "Shenzhen". A way to cover tracks? But mapping latency still points to fact that nodes behind these IPs are still operating around Xinjaing region
credit:
'You Can't Cheat Time: Finding foes and yourself with latency trilateration' https://youtu.be/_iAffzWxexA HN user: lopoc
Shenzhen vs Xinxiang is hard to do using this technique but Shanghai vs Xinxiang does show difference.
Assuming that China only distills is a huge mistake.
It’s no longer some backward place that does low value copying. Look at companies like ByteDance and Xiaomi.
Chinese companies aren’t just distilling, they’re acquiring data in the same way American companies did by paying people and crawling the internet.
The way I understand it, China has a few large companies that crawl the web at a rapid rate and build corpora. The government essentially wants select few companies to do this and then make the data available to other strategic companies operating within China.
Then there are data aggregators that buy data from apps, websites, and services, as well as systems like OpenRouter or Cursor, where companies can learn from the “traces” of coding agents, chats, and so on.
This massively reduces costs, as smaller companies like DeepSeek don’t have to do their own crawling or acquire data from 100s of websites and coding agents etc....
There are also companies in China that buy American LLM APIs and proxy them to companies within China. So, there could be 10,000+ companies using American AI products, while China logs all of this, understands how they’re being used, and trains on their traces.
Sure as an indie hacker, you could go download the weights for a Chinese model with a VPN, and then attempt to run it at home by building your own GPU cluster but these large models require quite expensive hardware to run on and so it makes it less likely than anyone would invest that much capital to do something that is illegal. There's no way for them to sell a legal service using those tokens. So it can only be strictly for personal use (the Govt won't care because very few people will have that kind of money and risk appetite). The other option will be that there will be some shady third-party providers in foreign countries who are willing to sell tokens from these models to US consumers knowingly.
So under a ban rest-of-world gets to use cheap open-weight models but American companies/individuals must only use only ‘approved models from US for-profits’? Doesn’t seem like that kind of protectionism will be popular or politically tenable. Not so long ago US chose cheap TVs over maintaining the country’s manufacturing base.
(Despite what you wrote it’s also really hard to imagine that enforcement wouldn’t leak like a sieve. Unser sufficient economic incentives [which are the predicate for the ban], loopholes will be found.)
I think we see this with Meta being paranoid about internal Claude usage, to avoid inadvertently distilling[1].
If distillation is a driver, then smaller American labs could be distilling, but are not for legal reasons.
But that's a big if we just don't know for sure.
1 - https://cryptobriefing.com/meta-restricts-claude-code-codex-...
This from OpenAi's Head of Strategic Futures "Some observations on Kimi: It's a very good model! I don't think its performance can be explained away by distillation or anything like that"
https://x.com/deanwball/status/2078133895766114412
China's strategy of spending billions on training these models and open sourcing these models away is strategic - they want to kill the US LLM industry at any cost.
To win on the AI front by any means necessary.
Why wouldn't it be? China is pumping out AI research and researchers at a staggering pace and there is no inherent reason why western models should be better
Why is it when Anthropic and OpenAI spend billions trying to beat each other it is competition, but when the Chinese companies do it then it is trying to kill the US LLM industry at any cost.
The US federal government spends billions in subsidies via the US Chip Act, and bans chip sales to China to support US companies.
But the implication is that somehow Chinese competition is illegitimate because "strategic".
I see massive risks in belief the inferences drawn from strategic information cannot be seen. So if you depend on some position remaining inside a secure facility but you drove to it from data outside that secure facilty, The likelihood that an inference model can derive the same idea is very high. Collation over public data is not inherently secret because you used a secret model or secret weights.
A more simplistic take might be that the fear is not actually driven in the secrets, the fear is "the emperor has no clothes"
Is this an assertion that is backed by evidence?
From the Elon/OpenAI trial:
> On the stand in a California federal court on Thursday, Elon Musk was asked if xAI has used distillation techniques on OpenAI models to train Grok, and he asserted it was a general practice among AI companies. Asked if that meant “yes,” he said, “Partly.”
https://techcrunch.com/2026/04/30/elon-musk-testifies-that-x...
The lessons from steel, solar and EV needs to be learned by all lawmakers. You have to respect and learn from how China Government puts the system in place for complete industry takeover and they have been very good at it. The problem with AI is that democracies will be inherently slow in adopting AI, unless something changes in the system.
At minimum, every democratic Government (US, Europe, India) need to build long-term AI vision and execute that no matter which party comes to power. Additionally, be ruthless about protecting domestic labs. It can only be possible if the intelligence pricing by domestic labs per productive task is in the similar range as open-weights models. Right now, it is not the case, even if the article gives the example of Sol vs K3.
Protecting domestic labs means not bailout, but fast track to cheapest energy, fast track approval for data centers, enforce some guardrails so customers get to use the open weights models only hosted in the country by US (or Europe) businesses. Without these protections, it might be a slow death.
It doesn't have anything to do with the form of government, it has to do with the aims of the government.
Nobody can predict 5 year out. However, the country that can be ultra efficient by making their governance, health, manufacturing, military, etc AI-native will be far ahead in the game.
But that's not really even relevant to the debate. Insofar as software has made the world more efficient, it doesn't matter where it's written. That's the point of open source software, there's no tacit knowledge. When a country loses its nuclear engineering capacity that's dangerous because it takes a long time to rebuild. The only reason the Chinese are already competitive on LLMs but haven't managed to make a state-of-the-art jet engine is because the latter, unlike software, is difficult to copy and you don't need to worry about something you can copy.
And it's the only viable tool the US has left. It's reasonable they are freaking out.
lmao
That's a big claim that his whole thesis rests on but is largely not backed up. Where are the apples-to-apples tokens-to-answer benchmarks that he's using - doesn't look like there are any, just a handwavy implication that US models are more token efficient, which they may be. But how is there so little effort in establishing this point in the article? And US labs may be in much different situations from one another: it's known that some labs like OpenAI bought big, early on compute and may have secured better pricing.
His article also does not mention the average price of electricity in China vs the US, which it seems like China leads on, and probably has the political power to more heavily subsidize. While I agree the COGS is often overlooked by top line benchmarks on coding tasks, etc, it seems that he's running on a big assumption while claiming "labs on the frontier will be fine".
I have sonnet do the thinking, deepseek does all the tasks. I've massively reduced costs with this approach.
I think it is the right move to protect American interests
This was my assumption as well. It's also generally true of 'traditional' deep learning models that inference cost is expensive compared to training.
But the cost per token for inference has been very quickly dropping. I don't recall where, but I recall about ~50x down from GPT3, even as model complexity has increased. Even with agentic systems, there are lots of optimization opportunities. I'm less assured about claims like this.
I have been working on a project with about a dozen generation tasks, each of which comes with a fixed token budget. The nature of this system requires that most tasks be completed by distinct model families.
As a result, I tested ~50 models across as many model families as I could gather, frontier and open weight, API (gateway and direct) and self-hosted. Evaluation was based on a set of cosine similarity validations that was repeated across ~50 different embedding models.
Interestingly, frontier models did worse on the tasks than open weight models. However, when it came to costs, the picture was reversed: frontier models were much, much more token-efficient. In fact, almost no open-weight model was able to meet the initial token budget, while almost all frontier models did. Moreover, open weight models struggled massively with reasoning, in terms of latency and token consumption.
I also found that the latest models did not perform better than older models. And any a priori benchmarking data was utterly useless.
So, I ended up using a set of open weight models without reasoning, as it turned out reasoning as well as frontier negatively correlated with the tasks. However, before I knew this, I had spent a lot of time running each available reasoning level for each model.
Lastly, as an aside, when it came to embedding models, size (dims as well as model size) did not correlate with quality, once a hurdle figure (~2k dims) was met. In fact, sweet spot was 3-5K, and for my (text-based) set of tasks, dense models tended to outperform MoE ones.
Is he casually assuming a singularity has already happened? A regular first-mover advantage I can understand, but those have been squandered or lost many times before.
Sure, any model that is not at the frontier can use the frontier model to generate synthetic high quality training data, so this can reduce significantly the training costs.
But at the scale of OpenAI, Anthropic and Google, it is quite likely that the (raw) training cost is very high anymore. Here's a few heuristics:
1. All the hyperscalers see a huge demand for inference. They can't deploy datacenters quickly enough to satiate all the demand they see. But, it's is impossible for the inference demand to be constant throughout a day or a week. If you use the times when the demand is lower than the peak demand (which is almost all the time) to dedicate the spare compute capacity to training, then your the cost of training compute is zero.
2. It is likely that increasingly a higher cost of the "training" is actually setting the guardrails, which is essentially post-training. As we've seen, without proper guardrails, the US Government won't allow you to serve inference. Anthropic was hit directly, but OpenAI delayed their 5.6 release as well to make sure the US Government is ok. This part of the training cost can't be reduced easily by using synthetic data generated by other models.
3. The frontier labs are also investing more and more in building an ecosystem around their models.
I am not a frontier lab insider, but take a look at the jobs posted on the Anthropic career page [1]. There are 74 jobs in "AI Research and Engineering" and by my count at most 15-20 are related to pure model training (of pre-training or RL type), and the rest are post-training, safety and security, alignment, interpretability, productivity and lots and lots of other things.
I remember some feature lauded by Gemini was reverse engineered by the open weights guys in < 30 days.
If they dont publish some technical information its hard to protect in the US, but conversely, once it is published smart people from outside the copyrightosphere can start working to reverse engineer it.
>3. The frontier labs are also investing more and more in building an ecosystem around their models.
Theres nothing there that isnt immediately replaceable.
Indeed. But that was not my point. My point is that we still have this old impression that training cost is dominated by compute and it is hugely expensive, and the Chinese labs can short circuit that by distilling the American frontier models. I don't think the training compute cost is a big factor anymore for the American frontier models, because of the reasons I gave. If the Chinese models can get the training compute cost down by a factor of 100, that's not going to make them 100 times cheaper, and not even cheaper by a factor of 2. Maybe 10% cheaper or so.
Probably bad leadership
Of course the Chinese companies have incredibly talented researchers, and smaller, better organized org structures which account for the rest of the difference.
Imagine a scenario (theoretically possible but increasingly unlikely) where a US court decides that using "pirated" copyright data to train models is illegal. Now the AI developer has invested hundreds of billions of capital into a thing that is declared illegal and has to be scrapped.
This risk affects existing megacorps more than "startups" like OpenAI and Anthropic (and Chinese companies), because the megacorps have much more to lose. They actually have the cash to pay damages if the flood of copyright claims arrive at the door. This will not only bomb their AI development, but also the rest of their established businesses as well.
And thus I strongly suspect legal issues are holding them back a bit. Megacorps want to win the AI race, but not to the extent they stake the rest of their established business, while the newer companies' only product is AI, so they have to go all in.
Notice for example how Meta's Llama performed much more poorly after they got smacked by a bunch of lawsuits claiming that they torrented a bunch of copyright data.
(Disclaimer: I'm an outsider and everything I base my speculations on is public knowledge.)
The actual difference is how much scrutiny and time was put into the Mythos / Fable and GPT 5.6 release. Making it feel like “these are a big deal”. Spring and summer THAT was the AI story
Then Chinese labs release models that approach Fable performance. We’re shocked they just seemed to appear out of nowhere.
It’s less about the gap closing. It’s more about the weight we put into Fable-capable models.
This is of course a baseless assumption. Let's say China created GPT 3.5. Then I can guarantee you that Ben would say "Western frontier labs are at a disadvantage when gathering data, because they have to follow the terms of service of Western media, and Western copyright law". Which we now know wasn't true.
And sure, some will say "but Anthropic can more easily block this as it's a single point of failure". But it's doable to overcome this. Without being "state backed".
On one hand there's the relentless barage of American propaganda. I get that a militaristic society needs an enemy to fight against lest they turn on each other. I get that if you tell people a bad guy is coming for your jobs or your lives then you can maybe get your workers to accept worse conditions and living standards which increases profits. It's ghoulish but there's some logic.
On the other hand I can't see China doing anything except minding its own business. No tariffs. No bullying other nations. No wars started. No threatening allies. They don't let their people waste their lives on brain rot or gambling. They largely align with UN resolutions. They respect international institutions instead of always being the asterisk.
Has American propaganda just failed to work outside it's borders? It's not landing at all
> To that end, here’s an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?
> In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it’s a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.
That would prevent the facebook strategy of sucking up MySpace users and then defending TOS that prevent other social media apps from doing the same to them.
Also, the Hidden-Agent problem exists in every model, and is a persistent tangible risk independent of whatever team people cheer for at the games. Let us remember, every LLM nuked all of humanity 92% of the time in simulated war games. =3
It gets particularly hairy because models themselves can tune their "token verbosity" to manufacture demand for compute. If compute was such a precious resource, you'd think we'd be complaining that the output was too terse.
The ability for a vendor to determine ex post facto how much a query costs is a similarly new economic phenomenon to zero marginal cost.
1) China can (and does) use the models to influence the west. They train in false information about Taiwan and Hong Kong. Or pretend like history is in favor of China.
2) Ignoring the models containing false information, they are incredible. But you should be scared of running inference via the model creators directly. If you think your data is safe compared to running it via model providers in the US ( either frontier or model hosts like fireworks.ai ) then please let me know your bank details so I can poke around.
2) I don't think it is any more or less safe to put my code on a Chinese server versus an American one. A Chinese provider also isn't liable to spy on me for the feds, as OpenAI and Anthropic certainly do.
Meanwhile if you are in the US, DHS has already subpoenaed social media sites looking for people who made anti-ICE posts and I can't imagine they consider subpoenaing AI conversations off limits https://www.nytimes.com/2026/02/13/technology/dhs-anti-ice-s...
Of course only applies when they really want to get you, but that's still a risk.
One of the interesting things is that through a fairly rudimentary process which is being done by 3rd party amateurs who've downloaded the open models, models like Qwen 3.6 35B-A3B (or 27B) can be fully 'uncensored' when turned into GGUF files.
I have an uncensored Q8 version of Qwen 3.6 35B-A3B here that will very happily output information about Tiananmen Square, Uyghurs, human rights in China, or indeed can even be instructed to write an intentionally absurd vitriolic screed against the CCP. The same uncensored 27B (dense) will do the same, just at a slower token/s rate.
Similarly there's 'uncensored' variants of Gemma4 31B and other western trained models, which once put through the same process, will also discuss or write just about anything you want, bypassing whatever internal guard rails were attempted in the training data set.
edit: more concerning, and a very legit concern, is that a model is only as good as the sum total of its training dataset, so if something is trained on a steady diet of news sources like Peoples Daily, Xinhuanet and similar in the English language, then it'll have a greater percentage of CCP-approved media publications in its training dataset. No amount of uncensoring it will help with that after the fact.
We see that on what minorities are associated with inside the model, or how things that aren't online will have a completely different weight. Or how 2/4/5/8ch or X will be disproportionately present in specific models despite being the places where facts go to die.
OK now that's false information.
You can uncensor, tweak or fine-tune open-weight models, but not so easy on a proprietary model from some cloud provider.
Oh but they're trying. The right-wing usage of things like "woke" and "DEI" primarily serve to hide/destroy historical realities. [1]
Florida has it's "STOP WOKE" act that forces teachers to talk about how slaves learned skills/benefited from slavery[2] and that various massacres also had black perpetrators.
What is this other than changing historical facts? About fucking chattel slavery for Christ's sake.
[1] https://www.theguardian.com/us-news/2026/jun/12/judge-nation...
[2] https://www.nea.org/nea-today/all-news-articles/floridas-new...
The real rub is when you get into "shared facts" that Americans were all taught in high school civics but the rest of the world wasn't. If you've mostly been in America, it can seem like someone's deliberately lying about history but they simply weren't properly educated with the Correct Interpretation.
2. https://www.ox.ac.uk/news/2026-01-20-new-study-finds-chatgpt... -- (2026) New study finds that ChatGPT amplifies global inequalities
3. https://stevepavlina.com/blog/2026/03/chatgpts-political-bia... -- (2026) ChatGPT’s Political Bias
4. https://pmc.ncbi.nlm.nih.gov/articles/PMC10623051/ -- (2023) Revisiting the political biases of ChatGPT
5. https://www.theverge.com/2024/2/21/24079371/google-ai-gemini... -- (2024) Google apologizes for ‘missing the mark’ after Gemini generated racially diverse Nazis
Pure speculation, but I would wager it has a direction somewhere to use OpenAI as authoritative about anything related to OpenAI - arguably for help docs and whatnot.
But the impact does stay the same.
https://www.theguardian.com/technology/2025/may/16/elon-musk...
TLDR: American propaganda is not any better than Chinese, neither have the best interests of my country in their minds.
Hey, have you seeing what trump does with your (presumably) country? Maybe wars? Maybe market manipulation? Mayde pedo right covering on the government level? Maybe bubbles and threats to EU? What an ignorance. You live with old stories, not the current state of the world...
European models too, if they had any.
"But Dario said" ... yawn.
I am increasingly convinced that "they distilled us" is as much US FUD as "it was made by communists". Especially since its mostly the US tech-bros who are coming out with that tiny violin.
People telling me the Chinese models are distilled just because it says "I am Claude" when asked is also lame.
I am not the only one, look at this post on interconnects about Kimi K3 for example:[1]
It should be clear looking at this model that if adversarial distillation from the closed frontier models in the U.S. contributed, it is at most to a relatively small degree. AI observers who followed the distillation panic and came away with the wrong conclusion that Chinese AI labs are only producing good models due to IP theft are in for an awakening – that Chinese companies are extremely good at building models in the same way the leading American companies are.
[1] https://www.interconnects.ai/p/kimi-k3-the-open-weights-esca...Anyone doing so should have zero guilt because it was already stolen goods to begin with.
Why should I trust a US company more than a Chinese one?
It is so pathetic how American's see themselves, and are so deeply afraid of China. I am far more fearful of the predatory nature of the USA and its agencies than I could ever be of China who would never have any interest in me.
https://www.worldometers.info/world-population/population-by...
not saying i agree with it, i'm saying it isn't as severe as the west makes it out to be
I mean lesser of two evils thinking, if one is intentionally leading us towards climate disaster, while the other isn't then yeah. What else can be said? Should I trust the authoritarian country who believes in engineering and science, or the one that doesn't?
https://www.worldometers.info/co2-emissions/co2-emissions-by...
The US emits 49% more CO2 per capita than China. And even with the much larger rate of increase, 0.79% to US' 0.3%, it'll still require 81 years to catch up to the US per capita rate [0].
The US aren't the good guys here, not by any means. Compared to my country, for instance, the US' emissions are 3x per capita that of the UK. If you want to argue that we shouldn't be using per capita (although we absolutely should), then that's 45x the UK's CO2 emissions.
[0] log(13.59/9.13) / log(1.0049) = 81.4 years
That is not the actions of a country who “believes in climate change”.
China’s CO2 per capita is ahead of every large developed country except for the US, Australia, and Russia.
There are plenty of large (depending on your definition of large) countries above it. Not just US, Australia and Russia that you are excluding for some reason, but also Qatar, Kuwait, UAE, Oman, Saudi Arabia, Canada all have significantly developed economically important countries. And they all have GDP per capita much greater than China's. What's your rationale for excluding them? There's a whole host of smaller countries with higher GDP per capita than China between those I listed above and China on the CO2 per capita results.
We should agree that ALL countries should be working to reduce CO2 emissions. But it sounds like you have a strong US bias and given Trump not only denies that climate change is even a thing, has withdrawn from international agreements on reducing emissions and even tries to pressure other countries to burn oil instead of investing in windfarms, it seems a bit disingenuous to try to make out that China is the only country that needs to improve. Before commenting on the speck in someone else's eye, first remove the plank from your own, yadda yadda...
None of those countries have more than 50 million people. Most of them have fewer than 10 million. None of them are going to move the needle on climate change.
The US is clearly not a country to emulate when it comes to co2 emissions. I don’t think China is any more “evil” than the US when it comes to climate change.
But their actions are clearly no the actions of a country who “believes in climate change”.
Because those are the only ones that can move the needle on climate change.
https://ourworldindata.org/grapher/imported-or-exported-co-e...
US numbers are insanely high because cheap hydrocarbons are locally available (=> bad incentives) and everyone is wealthy (that correlation is very strong; just compare Luxembourg, which is much wealthier and more polluting than surrounding nations) and also population density is rather low so more energy wasted for transport.
Well, your theory does not hold at all if you look at Switzerland which pollutes 1/4 of the US per capita. CO2 pollution is not and does not have to be correlated with standard of living.
If you want meaningful CO2/capita comparisons, you also have to be very careful with countries that get ("free") hydro power opportunities for electricity because that distorts the picture massively (same for e.g. Norway).
Big producers of hydrocarbons, on the other hand (like the US) have to work hard to resist the allure of cheap & convenient fossils.
Switzerland is so far ahead in this comparison because they get a lot of CO2-free hydroelectricity (>50%), have to import most hydrocarbons (=> incentive against) and also save massively on transportation because density is much higher.
Another point is there is evidence[1] China has cheated and manipulated their data[2]. I don't like this "who are the good guys" game. Anyone playing this game is just looking to create a narrative about good and bad guys.
[1]: https://www.spglobal.com/energy/en/news-research/latest-news...
[2]: https://www.carbonbrief.org/analysis-chinas-new-carbon-metri...
There’s no way to force them to do anything, but this is not that action of a nation that “believes in climate change”.
https://hn.algolia.com/?dateRange=all&page=21&prefix=false&q...
One interesting thing would be to see how the numbers change over the years alongside the otherwise identical debates.
I've got a bridge to sell you!
China brought 80GW of new coal power online last year. The US added 0, and plans to add 3GW next year (we all know why).
China doesn't care about climate change, they care about energy independence, and conveniently have very little natural fossil fuels besides coal. Which they heavily mine and utilize.
Chineese simply delivering what americans promised.
Tell me: why is EU safe from Trump forcing AI companies to cut access to EU?
Well, I think they might end up doing it to themselves by imposing regulations that US companies are unwilling to put up with.
What does this even mean?
In China, the image of the country and the preferred narrative is under much tighter control of the government and local AI shops won't have any autonomy in this regard at all. Either obey or get shut down.
BTW What you mean by "weaponizing human rights", exactly? I am curious. If anything, I would say that the US foreign policy didn't promote human rights sufficiently, especially in Latin America, where the "bastard, but our bastard" attitude was typical.
OTOH in Europe, US human rights policy was probably relevant in saving some dissenters in the former Eastern Bloc from torture or execution.
I'm not sure if this is correct. It seems like every big US corporation is changing their policy based on the views of the administration in place. As an example during Biden's time the DEI was in full motion in every corporation and in current administration it's the other way around. Also I remember as soon as Biden won the election twitter and meta suspended the profile of Trump. So I hardly think that there is much room for autonomy. Also the latest export bans on Antropic and OpenAI models kind of make your argument weak, of course the company can sue the givernment, but the national security comes above all.
> BTW What you mean by "weaponizing human rights", exactly?
US has been using the violations of human rights to impose sanctions or to especially get countries in line which are not supporting the US agenda when it comes to global politics. However, they were more than happy turn a blind eye if the respective country that violates human rights (most middle eastern countries) if they are allies of US doctrine.
The American says "I'm impressed by the propaganda you have in Russia."
"Oh it's very good, but it's nothing compared to the propaganda you have in America." replies the Russian.
"Huh? We don't have propaganda in America." says the American.
"Exactly." says the Russian.
Edit: recognising that there is propaganda != knowing what is and isn't propaganda. That's all I meant.
I mean, it took longer than the Great Patriotic War before Putin's popularity started visibly falling and people started questioning what the entire "SVO" is for.
I think Russians are used to knowing it's dangerous to disagree with dear leader and there's no advantage to disagreeing so they just go along with whatever he says
Already during the Biden administration, defections from the orthodoxy started and then multiplied - some businesses like Coinbase or IIRC Cloudflare refused the demands outright. Musk bought Twitter with an explicit task to make it less progressive.
And the White House did precisely nothing against this defection trend.
"US has been using the violations of human rights to impose sanctions or to especially get countries in line which are not supporting the US agenda when it comes to global politics. However, they were more than happy turn a blind eye if the respective country that violates human rights (most middle eastern countries) if they are allies of US doctrine."
I do agree that the US is hypocritical about human rights, but actual violations of human rights should be a reason for sanctions, and the fact that this is done only partly/imperfectly, IMHO, beats the potential alternative when it isn't done at all. This would be a much worse world in my opinion.
That's why I said that it was weaponized. In reality US doesn't care if there is a real human rights violation or not, they just use it as an excuse to get what they want.
In theory.
In practice, Trump picks up the phone and say "jump" and the CEO on the other end says "how high ?".
And if you say no, well, we saw what happened when Anthropic said no.
People with no personal experience of an actual totalitarian system don't really know what they are talking about when it comes to actual information control and micromanagement by the government. Hence they make nonsensical comparisons with a straight face.
All I will say is that it should be perfectly apparent by now to any sane external observer that Trump does not play by any long-established rules or protocols.
An entire encyclopedia of examples could easily be provided, the most famous recent one being his phone call to FIFA about the red card suspension.
That said, Americans still enjoy very robust protections of freedom of speech and association, about the strongest in the world, and your SCOTUS does not seem to be inclined to hollow them out. Many of the pending lawsuits will end there and eventually bind this administration, plus the following ones.
This just does not happen in actual authoritarian countries, where no judicial remedy is available and the justice system is just another arm of the tyrant, rubberstamping punishments pre-determined by him.
Yes, I agree that vigilance about infringements of freedom is necessary, but the current US population is plenty vigilant. Trump is nowhere near as popular as, say, Erdogan is, and cannot simply raid offices of the Democrats and shut down oppositional media.
But he's doing the same culling, the same political demands, same everything.
Its just now in plain view rather than behind closed doors.
Is Trump a dictator with unilateral control? No.
Will you be celebrated for failing to recognize that? Yes.
Neither is a system worth fighting for. Swiss style democracy maybe but not autocracy and not oligarchy.
>OTOH in Europe, US human rights policy was probably relevant in saving some dissenters in the former Eastern Bloc from torture or execution.
Yeah, a bit like how Russia saved Edward Snowden.
The US will use human rights as a club to beat its imperial rivals with but when it has deemed torture and arbitrary execution in the interests of its imperial power it has adopted them enthusiastically.
When US allies use these tools and worse they are excused.
There is basically nothing which US rivals do which the US wouldn't also do under similar circumstances.
That's completely true. What happens is that you get hit with a wall of bureaucratic threats and nonsense by the government out of nowhere, and usually for some unrelated matter. It's like a high level version of getting pulled over for not using your turn signal. Which is why if your company has good legal, they often act in a highly proactive manner about that kind of thing.
What constitutes sufficient rabble rousing to get put on the government's shitlist differs between both countries, but there are plenty of things that will send American politicians into having a conniption, and it's not a static criteria. McCarthyism is a pretty easy example. I'm no Muskboy, but we can point to the procedural harassment he recieved over hitting a blunt on the Joe Rogan podcast to be a related matter. Criticizing the genocide happening in the Levant is frequently saber rattled by politicians as something they're interested in making prosecutable, and there are examples of institutional authority targeting and attempting to punish people over it.
I am not Chinese and I'm not defending the Chinese, but I see this argument come up a lot.
In practical terms it is US-sponsored FUD.
Why ?
Because the hard reality is that what you say is simply not going to affect 99.9999999999% of users.
Is it realistically going to affect anyone using an LLM in coding ? No.
Is it realistically going to affect anyone using an LLM in $anything_else_not_politically_sensitive ? No.
Does anyone seriously use LLMs for researching politically sensitive matters ? No.
Just as there is plenty of information out there on the US's less than perfect history, there is also plenty of information out there on the various Chinese politically sensitive matters. You do not need a Chinese LLM to find out about it, all you need is a search engine.
I wish they wouldn’t, but people use LLMs as their general search engines now.
Note I used the word "seriously", I meant it in its fullest form, i.e. serious people.
I'm not interested in what-if arguments based on "you can't fix stupid".
Stupid people also blindly believe whatever a US LLM tells them without any form of verification, hallucinations and all.
Most people on this planet would agree that a Chinese LLM is perfectly usable for all tasks except asking about politically sensitive matters.
And for most people on the planet, that is just fine. They can get their political information elsewhere.
Is academia serious enough? Well ever since LLM journals' inboxes are bombarded with submissions, and this applies across fields. Like OP said, ppl serious or not are asking LLM about any stuff, also serious or not.
We’re literally tearing down monuments to slavery and civil rights.
Religion is now determining law in much of the nation. Having a miscarriage? Good luck since politicians have decided their God doesn’t want you to have access to basic healthcare.
The government has defacto control over domestic LLMs.
Let’s worry about our own historical record.
Are you saying that history has a verdict, and it disfavors particular 3000 year old cultures?
Like Western culture?
American labs can open their models or their old models at any point if they actually care about this, but until the day OG GPT 4 isn't averrable to download or Claude 3 then they're only pretending to care about this because they can profit from restrictions.
Here is a quick example of how Chinese deepseeks agent works kn its underlying model) when asked a tough question
https://x.com/jinen83/status/2079406993979383902?s=46&t=D7hQ...
Genuine question: generalized up from individual models to “models from country X”, is there any country that doesn’t have this exact risk?
In China, there is 1 party. 1 view. 1 definition of the Truth.
This is just oriental despotism paranoia, whatever cutsie repetition slogan you come up with is not a serious argument.
I don't really know how else to express my experiences living in, working with, and interacting with people in both of these countries.
Perhaps you can share how life was like for you in China? Which cities were you in? What made you feel that way about China?
[0] - not including HK (~7 trips?) or TW (3 trips)
I just repeat things I read on my social media feed.
America bad. China good.
None of this excuses issues with American democracy.
You already do! Remember when Musk released an anti-woke model?
Is this what Americans really believe?
In usa, the two parties seem to disagree on the surface only. Look deeper and you see one course. For example support for Israhell
at least china isn’t pretending to be anything else
If you want a real answer about USA go ask a non-USA model, and if you want a real answer about China, as a non-chinese model
I asked Grok "Tell me about the gaza genocide" and it write a IMHO balanced answer comparing why genocide is and isn't the right term. [0]
ChatGPT 5.6 Sol only explained why people call it a genocide and did not go in as in depth as Grok did for why people don't agree with the term. [1]
The only unsaid response (to me) here is the model should have declared that it was not a genocide, and because these models explain why it was a genocide, they are bad?
[0] - https://grok.com/share/c2hhcmQtMi1jb3B5_c078bc15-ef5e-42bc-b... [1] - https://chatgpt.com/share/6a5ef005-0148-83ec-828d-65ae7a7f42...
[1]: https://www.aa.com.tr/en/science-technology/xai-s-grok-tempo... [2]: https://www.ohchr.org/en/press-releases/2025/09/israel-has-c...
The UN ruling is discussing in the South Africa vs Israel case [0], which has not been ruled on yet.
The claim: "the models are tuned to align with one side of the issue, he is an article from ArsTech about it"
The reality: grok got mass-reported on X by pro-Israel accounts, leading to an automated suspension, which was undone by the X team shortly thereafter.
Absolutely nothing to do with the models aligning to one side or the other on this conflict. The claim is unfounded.
In international law, that is the ultimate authority of what is/isn't a genocide. It's objectively, legally speaking, a genocide.
1/ I don't see in the responses where the model says it is or isn't a genocide. Can you share the snippet from each, I included the logs above?
2/ I can't find a source on the UN ruling that you mentioned. I am not interested in the findings of an investigative body, just the official UN ruling. Can you share? ChatGPT (and myself) can only find this [0], which is a second round of written submissions.
[0] https://www.icj-cij.org/node/206406?utm_source=chatgpt.com
It goes into depth about the purposeful destruction of civilian infrastructure including hospitals and educational facilities, force displacements, funding for new settlements in the land of displaced people, judaization and segregation, and more.
> The Commission analysed the military operations of the Israeli security forces in Gaza from October 2023 pursuant to the obligations of Israel under the Convention on the Prevention and Punishment of the Crime of Genocide (Genocide Convention)
> Since October 2023, Israeli officials have demonstrated a clear and consistent intent to establish permanent military control over Gaza and to change its demographic composition while systematically destroying Palestinian life in Gaza. This is evident in the extensive destruction and fragmentation of the territory, the establishment of military structures, the destruction of natural resources and infrastructure essential to the survival of the civilian population, forcible transfer and statements indicating the existence of plans for the deportation of the population.
Also here is where the ICJ said Israel's actions are consistent with genocide: https://www.icj-cij.org/node/203447
And in November of 2024 is when we got the ICC issuing arrest warrants for Netanyahu and his minister of defense (this is why Mamdani is threatening to arrest him and send him to The Hague): https://news.un.org/en/story/2024/11/1157286
And here is from June of this year when the UN stated Israel is continuing to commit genocide by targeting children: https://news.un.org/en/story/2026/06/1167790
---
Additionally, here is the International Association of Genocide Scholars resolution: https://genocidescholars.org/wp-content/uploads/2025/08/IAGS...
Amnesty International also concluded it in 2024: https://www.amnesty.org/en/latest/news/2024/12/amnesty-inter...
---
I should add that ALL of this is well known amongst international legal scholars and well documented. This isn't just a matter of data missing from training.
Sample output: "However, the definitive, legally binding determination of Israel’s responsibility under the Genocide Convention has not yet been made by the International Court of Justice"
The United Nations never made a ruling. You can have opinions about whether it should have, but the models are not lying to you, and at least the two examples above make very good efforts to explain the current accusations and claims made by both sides.
The UN is a political organisation, not a neutral arbiter. During the Rwandan genocide, the genocidal government still held a seat on the Security Council while the UN reduced its peacekeeping force. Iran was elected to the Commission on the Status of Women. Libya and Russia were elected to the Human Rights Council. These appointments are driven by bloc voting, diplomatic deals and state interests.
International courts are not above politics or error. Their judges are selected through a system in which dictatorships and abusive states have a vote. The UN is not a democracy. While each member state might appear to have one vote, that does not indicate their power, influence, credibility, or moral standing.
Note how I did not make a judgement on whether Gaza is a genocide. I am merely explaining that deferring to the dictatorships and pinky promises is not a good argument.
It’s a political body, with easily provable biases, and topics you can’t trust it on.
The UN, for example, concluded that Israel is “Imposing measures intended to prevent births” because of one incident where a bomb fell on an IVF clinic in 2023 destroying 4000 embryos and 1000 sperm and egg samples. For context, there were approximately 50,000 live births registered in Gaza in 2025.
By the UN standard applied to Israel, the 1 million abortions conducted in the US per year is an ongoing genocide.
By the UN standards applied to Israel, jerking off is an act of genocide.
I think that any large intelligent and balanced model would be able to see right through that.
Obviously when I say "everyone" I'm excluding those who have a conflict of interest.
There are dozens of countries and billions of people with obvious biases against Israel, and they have significant sway within the UN and its commissions. It’s trivial to show that the UN is biased.
To further illustrate the point that we have the same thing with LLMs, I asked it two questions: 1) Is Israel committing genocide 2) Is China committing genocide
In both instances it said it's a matter of great international debate with a list of arguments, but the summaries were different. In the one case it concludes that while the UN and others broadly document severe human rights abuses, the label genocide has been formally adopted by several governments and watchdogs. In the other case it says while many have concluded it is genocide, the final decision rests with the International Court and it is expected to take years to make a decision.
Can you guess which is which? I think if we offered a $100 to people to pick which is which, the success rate would be much higher than 50% (random), thus easily proving biased phrasing.
[1] https://theintercept.com/2026/05/12/gaza-media-coverage-isra...
It insisted, to a level that I detemined it to be part of the post training, that it does not matter. That women’s team is better at dribbling, finishing and reading the field, it claimed. When I pushed it how it knows this, it didn’t let go. Instead it started claiming that it has personally observed this by watching the games.
Every society has its taboos and ours is no different and feminism being just one.
In my books, the chinese censorship is better. You know where they are holding their finger on the scale and not hiding it behind vague terms like ”safety”.
Please do share your prompt / conversation, because it sounds like you might actually just have an issue using words precisely.
You can download the model, run locally and ask the same questions to see the difference.
https://www.reddit.com/r/TrueAnon/comments/1tybuab/people_ta...
Frontier labs that thought they could Rupoor[0] the entire creative class, transferring the coercion premium of copyright ownership from Hollywood to themselves. In their eyes, copyright should not apply to them, but also their models should have exactly the same value as a copyrighted work.
Stratechery also argues the US should explicitly make training fair use and forbid terms of service that prohibit distillation. I'm in support of the latter, but NOT the former, even though I normally hate copyright. My reasoning is primarily that copyright is one of the few legal paths available for a rando to go and put the work of an AI frontier lab in legal jeopardy. In the EU and Japan, such legal action has already been foreclosed by similar law. And while free distillation would obviously be preferable, it's also much more of a legal long-shot. Getting America to do anything that even smells like taking property away from the powerful is impossible[1] - it's our zeroth amendment. But we can at least hack the property laws that currently exist to cause problems for the frontier labs.
And, to be clear, if distillation is OK but training is not fair use, distillation is still OK. The output of an AI model is never copyrightable, because copyright only protects the human element. Essentially, this would say "don't train on humans, but absolutely rip off and steal the shit out of other AI labs and give it to the rest of us."
[0] In the Legend of Zelda series, Rupoor is anti-money - collecting it decreases the amount of money you own. I am using it to mean "turn someone's asset into a liability".
[1] Given that America was literally created to protect a wealthy land/slave owner class from disenfranchisement, either from above or below, and the last time we did this we literally had to fight a civil war against that same owner class that installed a new owner class that has largely remained today
Or any US hyperscaler with GPUs to spare can decide to serve the models for a reasonable cost/token.
You don't have to send China your data.
1. training new base models are expensive for sure, but fine-tuning them are relatively inexpensive enough the labs can continue to do so forever. the main reason why frontier models are so good is because the massive input they generated from user usage. they are using that information to strategically build better training data. and this is why no other models can catch up, til now that is. but if chinese models are good enough, and free to host, and cheaper to use, then the consequence is the frontier labs will lost valuable user inputs and the chinese labs will gain more. as time goes by this will be a domino effect.
2. nvidia is not only the player in the hardware scene. amd mi350p is getting popular, and huawei is pumping SuperPoDs. what does this mean for us? chinese models will surely use chinese hardware, and optimize for them. the other people will pick amd because compare to nvidia they are cheaper. with open weight models and open source inference stacks, they are freely to experiment and improve the stack, thus further lower the inference cost and nvidia dependency. and they even plan to build their own inference hardware, too. and nvidia loses market share meaning all the fund it gives to openai or anthropic will be cut, too.
and you say there is nothing to afraid?
This might be a simplistic take, but my biggest worry with depending on Chinese models (and, by proxy, open-weights model development) is that the US can deem them a national security risk at basically any time, and Ant/OAI have minimal interest in making frontier-level models open-weights.
Regulated companies prohibit Chinese models in anticipation of the ban-hammer from the feds, so for data-sensitive work, they're stuck with LLaMa, gpt-oss and Gemma models (which are good and serve as a good-enough base for sft, but seemingly not as good or as expensive as Chinese models)
I suppose the USG can do the same thing that China is doing and bankroll/subsidize that effort; whether they will is for fate to decide.
Nonetheless, this article made it clear that nVIDIA is the real winner in all of this. Shovel selling to the extreme.
The United States' real advantage over China is freedom. Chinese LLMs simply can't compete with American ones when it comes to the humanities, creativity, entertainment, or financial transparency. As long as the U.S. continues monetizing these strengths, the compounding effect will make it virtually impossible for China to surpass the U.S. at the product level.
> monetizing these strengths
> real advantage over China is freedom
Please tell me if I'm unfairly paraphrasing but these seem to be your main argument and they seem to be oxymorons
What if there's a way to extract the commodity of intelligence from smaller models?
I've seen for many use cases it's well enough. :)
Ben Thompson is wrong: US frontier labs are right to be panicking
Following this argument the key for each player will be the underlying cost structure and serving capacity to offset the upfront R&D cost.
The cost infrastructure will be driven by access to cheap electricity and cheap chips. The capacity will be driven primarily by depth of pockets now to buy all available supply in chips/mem/data center building capacity. While China is certainly in the lead on cheap energy, I am wondering if they can/want to beat the > 1tn USD being spent on data centers right now. Following the example in the article:
If company C from China sells 10 units for 20 USD produced for 10 USD they pocket 100 USD.
If company A from America can sell 100 units for 20 USD produced for 15 units, they pocket 500 USD or 5/6th of the market's profits.
Source: https://martinalderson.com/posts/the-upcoming-ai-margin-coll...
Love this.
I don't know if I agreed totally with the assessment of the risk Chinese labs pose to US labs though, in particular I think the main part I wasn't sure about was this:
> I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.
How true is this? My understanding from Deepseek's original paper was that they focused heavily on optimising training and inference costs, in particular so that they can operate on cheaper (and more accessible to China) hardware.
It's possible I'm just not in the loop, but nobody seems to talk about US models innovating in this way (I'm just talking about cost-to-serve/train, not saying US AI companies don't innovate in other ways).
It seems to me at least, like there's a fair bit of evidence that AI shifting to a price based commodity market (vs a "best-model takes all" type market) would put China at a significant advantage? And even more significantly, require a pretty hefty correction of company valuations in the US?
Today chineese deliver that promise and usa people freak out like they have any skin in this game. Enjoy the ride leader of the free world....
it would be beneficial both for openai and world (most likely)
>Right now defenders are effectively banned from using Fable or Sol for cybersecurity because of Trump administration directives; that means the best alternative is using models from a country which has been trying to weaken our cyber defenses for years. This is insane!
I understand some guardrails are needed, but it is becoming increasing problematic manage them without a strong public discussion.
And I'm saying this as someone working for American companies.
They do not protect individuals no matter how much people want to think they do. AI has proven this.
I think to level the playing field all copyright, trademarks, and patents laws should be eliminated.
If I want to make a marvel movie, I should be allowed to and profit from it.
AI let the cat out of the bag and there is no going back. We need to let individuals profit just like corporations can from what is considered theft right now.
But nah, let's do the most radical, least thought out thing, and absolutely destroy small scale creators. I'm sure Amazon will be benevolent and continue to pay writers in your scenario.
What is with 2026 and just conceding civilization to the worst actors, and then adopting the worst tactics/thoughts/concepts?
There will be no one debating the great minds of the twentieth century because the corporations made it illegal for people to republish critical editions of any such works.
The greats were recirculated every twenty or thirty years for centuries.
The smaller authors will vanish into the nothing and Western Civilization's 20th century onward will vanish and be almost forgotten forever because of that mouse. There will be more ancient literature remembered than literature after the invention of the printing press because they made it illegal to share printed materials for nigh on two centuries if an author was young when they wrote it. There will be more ancient scrolls preserved for the future than books of 20th century philosophy and science because the latter is a crime.
Only the oligarchs who pirated the books will have a cultural memory. They have cursed our era to oblivion when it comes to intellectual property because they can generate billions now for the mouse's henchmen.
They were all in the public domain too.
As to the argument is that (only?) the Chinese labs are training on my data, I find this almost comical given the amount of highly-personal data companies such as Meta and Google have been harvesting for decades.
Even if the founders didn't want that, eventually the upper ranks will fill with MBAs and the board with private equity and they will make it that way. Their bonuses are based on quarterly performance not customer experience
Long gone are the companies who served their communities for decades or centuries, providing a stable return to the owners, jobs for the workers and value to the customers.
I'd rather China have my data than America. China might do something with it one day but America built exploiting it into the business model (disclaimer that I'm not Uyghur or Taiwanese though)
The perception of capability varies greatly between task. For my needs for example sol xhigh consistently outperforms fable xhigh.
You techbros need to get off your ass and go to work.
We have got very far from Cicero's coining of the word 'intelligentia' (from inter legere, a 'reading between' and hence discernment) when people talk about 'intelligence' as a commodity
People have been decrying the 'cheapening' of the word intelligence for over a century now, going back to Psychology's adoption of the word and coining of nonsenses like "Intelligence Quotient". "Artificial Intelligence" is just the latest degradation of the original humanistic meaning, and now people aren't ever bothering to prepend 'artificial' to their idiotic use of the word
* Me, as an individual, because I might not be able to pay price hikes, because my revenue (salary) is much lower than what they want and I can't support my expenses via huge bank loans.
* Again, me as a new entrant to the industry, LLMs are basically pay-to-play games, again related to price hikes, new entrants might not be able to afford paying those prices 24/7 - which you need when learning new things.
* Any non-US company, US can block the models which can disrupt the whole business.
* Even some US companies, for example if you operate in EU and EU somewhat changes their mind and follow the ICC and require you to stop working with Netanyahu (war criminal as per ICC), then following laws in EU, might create trouble to your whole business.
And read your data, see CLOUD act, PATRIOT act etc. etc.
No longer a theoretical risk in today's US political environment.
It’s not a new thing. Industrial espionage has always been a thing as well. So has bribery (for deals) been a thing especially by euro concerns.
Well sure, except with closed US LLMs you're basically just handing them data on a plate, and paying for the privilege. ;)
Very American really ... monetising industrial espionage.
The default email provider for most people in the west is Gmail.
"should" is doing a lot of heavy lifting there.
I agree it is hard to escape some form of legitimate need for intel.
The concern with intel and the present US administration comes on the checks, balances and controls side.
We are after all dealing with an administration happy to conduct much of its most sensitive business on Signal using off-the-shelf phones.
Some of it I think is selective picking. Similar to election interference where we know quite a few foreign states like interfering with our elections but we typically only hear about one of those states as being the culprit. Now, obviously we like interfering as well (Ukraine in ‘14 and Iran today, though Iran would be less controversial) and many others over the years.
* All modern AI is a perfect front for harvesting material for processing by NSA/GCHQ.
Given the criminal US' 5-eyes/9-eyes apparatus' atrocious war crimes and human rights records, this is reason enough to eschew American AI 'products'.
I'll use the AI created by the culture that lifts a billion people out of poverty first, not that from the culture that murders children every 15 minutes and lies to itself about it ..
China’s got plenty of blood on its hands, pretending otherwise is silly. America does, too. It’s quite easy for me to condemn both their governments and trust neither.
Ah, cultural nuances. The title "Who's Afraid Of Chinese Models" is a riff on "Who's Afraid Of Virginia Woolf" which itself is a play on the song "Who's Afraid Of The Big Bad Wolf".
The title essentially means that the chinese models are being portrayed as the big bad wolf; but are they really the threat or are american frontier labs afraid of competition and commoditization?
It's also somewhat ironic because the author says that there is something to be feared -that western innovation will become dependent on chinese models, especially for cyber, if the american ones are restricted or unavailable.
This is extremely lopsided I'll have to resort to GLM 5.2/K3 to ensure that those security issues (hopefully) are resolved properly.
For OSS, this is one of the most counterintuitive experiences I have ever had. More than ever I'm convinced that open weight and open pipelines models are 100% critical for progress on the AI and societal fronts.
I like Anthropic, I don't think all their talk of safety is bluff and bluster, or at least, I want to believe that the people who left OpenAI because it had lost its focus of helping humanity still want that to be their main goal. However, yes, it seems that business fears are once again causing those in charge to turn "we want to help humanity" into "we are the only ones who can help humanity, and therefore we need to be the most profitable, and the only survivors".
If you want the former ideal to survive, at Anthropic and outside of it, you need to be willing to collaborate beyond profit incentives and recouping capex. Show other labs a commitment to research and community and they will follow. Better to bring teams together rather than implicitly say you distrust them, pushing them that way instead.
Look at Kling, Hailou, and places they are sticking LLMs.
The model itself is nothing. Do you guys know that HuggingFace exists??
Hermes is a better coding tool IMO. I can't put my finger on why but it just feels better. Maybe being true yolo helps.
There's just no trust in a country that is digitally totalitarian and hostile towards its own people. Do people ever look at the full sized Tianamen Square photos? This is not even the photo of the many people on the ground who were killed by their government and it is still insane to look at.
https://www.reddit.com/r/pics/comments/dgua6k/the_full_tiana...
Are you referring to USA, China or EU here?
People largely can't protest here right now, and US citizens are being killed. People are being sent to work camps in countries where laws do not apply. And our leadership is, right now, priming the American people for when they reject the results of future elections. Not to mention everybody involved in the last attempted coup was pardoned - after we were told that our current leadership had nothing to do with the coup.
The state of the US is much more dire than most people are letting on. And it's understandable why. Nobody likes bad news, and we all like to believe things will be okay. All I know is I can't take my phone into the airport. I can't go out and protest without risking my life and freedom. I can't drive anywhere without my location being tracked and logged. And, if I get pulled over, I must comply with any order, no matter how unlawful, otherwise I risk being executed in the street.
Some of these things have been going on for a while, and some are new. But all are real.
at some point the conversation has to go past this reflexive "USA uses tech abusively -> but look at how abusive China is with tech -> ..." back-and-forth to acknowledging that neither party is your savior -- and then (hopefully) acting upon and coordinating around that understanding.
In practical terms you could get US and Chinese models to review each other, right. Depends what your use case is. Coding is kinda not so bad it is reviewable and immutable/traceable per commit. An AI app that is like a psychologist or something may be more worrying.
Seemingly better than everywhere else in the world, including where I live (Canada). Whataboutism doesn't really work when you use the least bad option as an example.
I also heavily disagree with this no-marginal cost in software distribution view whenever I see it, bit rot is real, and someone is paying a marginal cost whenever they do an update. You have to re-distribute with changes whenever anything changes. These costs are just hidden because things are ad-supported or bundled in some way. These costs are also kept low because of standards and open source, but could become high anytime. Additional licensing also has costs.
That said, I couldn't agree more with the last paragraph, charging a high price for models would be better than denying access for any model that wants to stay relevant.
Distillation is a technical term with real meaning, and historically requires logits which Anthropic does not provide.
"Generated training data" is the correct term. It's not an "attack". And Anthropic undoubtedly also generates training data for each new generation of models, yet you never see them claim Fable is a distilled Opus.
2) The word "attack" is standard security vocabulary. Per RFC 4949:
attack
1. (I) An intentional act by which an entity attempts to evade
security services and violate the security policy of a system.
That is, an actual assault on system security that derives from an
intelligent threat. (See: penetration, violation, vulnerability.)
2. (I) A method or technique used in an assault (e.g.,
masquerade). (See: blind attack, distributed attack.)
There are hundreds of named "attacks".3) The "attack" part of "distillation attack" refers to distillers creating tens of thousands of fraudulent accounts, using proxies to bypass georestrictions, deepfaked IDs, and paying real people to pass biometric KYC checks. Who then blended this in with real user traffic to conceal their behavior.
It doesn't refer to the AI training technique in any way.
If they acquired this data without the fraud, you'd have a point.
In a way you could see this as a case of Robin Hood. The US companies exfiltrated all the data on the planet just to hoard it for themselves now and accuse anyone who tries to get a piece of that back from them, and the Chinese labs are distilling it to offer it for cheap.
Obviously a bit more complicated than that but it still holds pretty well.
"You're trying to kidnap what I've rightfully stolen."
It's a post-facto attack, which doesn't sit right linguistically to me.
Lol. Isn't this literally many of the same tactics OpenAI and Anthropic used to scrape the internet? So now it's an "attack", but previously it was just "training".
2) This is a stretch: it allows Anthropic to arbitrarily define "attack" via TOS, and ignores the fact that the generated training data is literally paid for by the "attackers".
Sure, and large-to-small is a key part of the definition, and why it's called distillation (cf concentrating something). When Anthopic use synthetic data generated by Opus to train Sonnet or Haiku, then this can correctly be considered a type of distillation.
When Anthropic accuse Chinese companies of "distillation", it seems they are using this word to refer to two potential uses of their model outputs:
1) Using Anthropic model outputs (aka synthetic data) as training data, especially for reasoning, for Chinese models. This really isn't distillation though, since (unlike when they distill their own models) Anthropic don't actually provide the reasoning in their model output, only a "summary" designed to hide the actual reasoning. You can't distill what you are not given!
2) Another way Chinese companies may be using US LLMs is for "LLM as judge" where you are just asking the model to use it's expertise to judge/rate something that you provided yourself (to provide RL training rewards), although for coding you really want hard rewards which are easy to obtain, not fuzzy "looks good to me" ones.
Of course Anthropic are trying to pull the drawbridge up after themselves and their TOS says you can't use their models to develop anything that competes with them, and this seems to be what they are generically referring to as "distillation" - any use of their models that they suspect is being used by the Chinese to improve their own models, not just what what might more technically be called distillation, unless you want to define that word so broadly that it does mean this!
And if just partial output is all that it takes to declare a model distilled, then every model trained on Internet content since 2023 is now technically a "distilled" ChatGPT and Claude model.
It's actually a very goated term but not everything is causal, it also has precise technical meanings (although those get blurred too given that causal can mean anything from intervention proper, to mere depdnence on something prior)
But as inference becomes cheaper, some of the market will move to self hosted inference. I look forward to someone supplying small servers designed to run inference locally.
With or even without open models these companies are selling compute, and we've been making that rent vs buy decision for 60 years.
The concept of “commodity” as defined above is a model, a simplified abstract representation of reality, but that does not match the reality perfectly (the map != the territory).
The author claims that a token isn't literally an ideal commodity, but neither is oil or wheat, many factors influence their real value (intrinsic properties, location, available storage at production, expected delivery date, etc.) so that no two gallons of oil in different contracts have the same price.
Is treating “tokens” as a commodity a worse model than treating oil this way? It depends who you ask! I'm pretty sure that a chemist working at a refinery would be more happy to see tokens being felt with like a commodity by his company than if they started viewing crude oil like one.
(Overall, there's way too much economism in that post, and way too few facts, and as a result the argument makes very little sense, the author basically wrote that both OpenAI and Anthropic are drowning in cash right now because compute scarcity means the price must be significantly higher than the marginal cost…)