On Openrouter Kimi K3 says it does not retain data or train on it, which is better than what US hosts claim for Claude, ChatGPT, etc.. as they collect and retain data even if you disable training on it.
Opencode or similar open source tool + a zero data retention provider is about the best option aside from running a smaller fully local model on your own PC.
What is the parento frontier?
So like, on a cost-intelligence graph, the cheapest and most intelligent models are pareto optimal. Then in-between those if you have
- cost $3 intelligence 6
- cost $1 intelligence 5
- cost $2 intelligence 4
The 1st and 2nd are pareto optimal, the 3rd is not, because it's dominated by the 2nd (2nd is cheaper AND more intelligent at the same time)
The Pareto frontier tells you which designs are the best in at least one of your metrics (non-dominated by another design). For example if you're selecting a car and you care about both speed and mpg, a Formula 1 car and a Prius might lie on the Pareto frontier, but a Model T Ford would not.
For example, Kimi 2.7 has been really effective for me despite having verbose thinking blocks, simply because it runs so fast. Speed-wise, it feels about like Sonnet, possibly faster.
When you say "Claude", do you mean Opus? Fable? What effort level?
Although in the month since that most recent post, his other points about open models are undercut by K3.
And at least of data available as of 2026-01, AI compute capacity was doubling every 7 months, so I expect every major country to host AI compute farms, and self-host AI feasibility to majorly increase in the next 2-3 years as well. (Partially undercutting, but not fully disproving, his points.)
And I wish his posts were 4 times less wordy.
The marginal utility problem is a real one for AI companies. I think the current generations are already saturating marginal utility for 95% of the population. Almost everyone I know outside of my career has no use for a more powerful model. This is a serious problem for the economics of AI and semiconductor investment. This is a bigger problem than Chinese models. It leads to a demand curve problem - that supply outstrips demand.
I only have the $20 plan from OpenAI and the same task, with a lot of the same implementation details as Kimi Code, only took a few minutes and consumed almost none of the 5 hour limit.
Subscription usage limits are hard to measure as none of the providers tell you directly what it means in terms of tokens or anything else you can easily compare, but when I sat down to add Kimi Code to flar, it was because I wanted to try it on some real work and then couldn't do any, because usage was nearly gone after the trivial task...no other ~$20 subscription I have has felt that tight before.
So, it was really slow to complete the task and seemingly much more expensive than every other model I'd tried. Maybe bad luck. Maybe it'll do better on other tasks. I wouldn't know as I was out of usage when I had time to try.
It did find a bug that Gemini 3.5 Flash introduced unprompted, though, so it has that going for it.
Subsidies would affect 1, but not 2. But if some VC wants to subsidize my Claude or Codex or whatever, awesome.
Additionally that same VC could be (read: is always) spent on developing the harness, and other infrastructure around the model, not just the model itself.
So it's apples-to-oranges when comparing a relatively new model to established competitors (i.e. OpenAI @ $900B funding vs Moonshot/Kimi's $30B FYI) because every new model they release is judged on "performance" which is not strictly speaking derived solely from the model.
It's possible Moonshot could get similar performance over time as the build out the rest of the infrastructure. We have no way of knowing how much of OpenAI/Anthropic's success is due to the model vs intelligent tooling built on top of it.
And, DeepSeek is what I use for any task that works best with an API. It's cheap enough to where I don't think about cost, made even cheaper by DeepSeek having the most effective and cheap caching in the industry, and it's good enough to where I rarely have to follow up with a more expensive model or manually fix things. It's been alleged they're releasing an update to DeepSeek V4 Pro soon that improves it, which likely makes it a good fit for even more kinds of problems. It remains my favorite of the Chinese models, it's so cheap and cheerful. And, is also less aggressively censored than some of them.
I use DeepSeek v4 Flash & MiMo v2.5 Pro. Prefer the latter over DeepSeek v4 Pro because it costs the same while being equally good & less chattier for coding workloads. Although, I've begun experimenting with Hy3 (as an in-between Flash & Pro) & GLM 5.2 (for long-horizon tasks).
I gave Claude Code/Fable the same task and it took significantly less time, but also stumbled on the same error as GLM. I didn't have it fix it though. I was mostly interested in timing differences.
I do like open models where I can, but I'm really hoping they get trained to second guess less. Or maybe I just need to prompt them differently. I'm not sure.
ArtificialAnalysis puts Kimi K3 just below DeepSeek v4 & GLM 5.2 in token use per task, which is about 2x to 3x more tokens than Grok 4.5: https://x.com/ArtificialAnlys/status/2077832879187620192 / https://archive.vn/zBbFi 2 other open weights MiMo v2.5 & MiniMax M3 are comparatively thrifty.
> Subscription usage limits are hard to measure as none of the providers tell you directly what it means in terms of tokens or anything else you can easily compare
I always put my coding subscriptions (that allow it) through "AI gateways" (Cloudflare & OpenRouter are free) which help track token use.
In my experience, Kimi & Qwen Cloud have opaque & restrictive limits, their "credits" drain faster. I now make it a point of subscribing (directly [0]) with providers that are transparent like MiniMax, DeepSeek, Xiaomi, & Z.ai.
[0] OpenCode Go, Cline, and AtlasCloud have generous limits for open weights, otherwise.
AI subscription pricing is so goofy. You get some amount of usage that varies by models, is measured by opaque token usage, driven by how many tokens the (usually) vendor-provided interface (or model itself) wants to use. Then your usage is limited by time opaque time windows.
Now, of course, the plan is to remove Fable from the subscription. To paraphrase Darth Vader, they have altered the deal. Pray they do not alter it further.
The benchmarks come out and say they’re as good as Opus from N months ago, then I use it for a complex task and it doesn’t work as well as Opus from N months ago did when on similar problems.
There’s a real wow factor when you get an open weights model to do amazing things, but in my experience the gap to the frontier models has always been bigger than the benchmarks would lead me to believe.
There can be a lot of value in having the cheaper open weight models for chewing through lower complexity tasks (non-programming in my primary use case) at a cheaper rate than OpenAI or other frontier API costs. Even with those I can measure bigger gaps to the frontier models than the benchmarks suggest.
If the benchmarks aren’t being directly gamed, there’s at least some selection happening where training data or model structures are being picked in ways to maximize public benchmark performance. All of the labs know there’s immense value in having good benchmarks to show for your model because most LLM consumers are picking based on lab provided benchmark charts, not running their own evals. Running your own evals is hard and expensive.
What OpenAI in particular have done with reasoning efficiency in the past few months since ChatGPT 5.5 is nothing short of remarkable. It's overshadowed a bit by the benchmark game and the Fable hoopla.
Now is the time to focus less on token cost and intelligence, but tokens to solve a particular set of tasks in closed benchmarks for a variety of categories.
What is the use of grand intelligence if it either costs you a kidney or can't complete at all within a token budget? Even if there are niche uses where you truly want "maximum power" above all, we need to at least more severely penalize such models versus those that does it just as fine within a tenth of the token cost.
I'm aware of some benchmarks at the Artificial Intelligence site, but CLEARLY we are not focusing enough on these today and still leaving the fun surprises to the users.
Of course they exist, Alipay is from Alibaba, think about who typically buys from OEMs/suppliers there...
... but I borrowed a friend's +86 phone number, which you'll need to even see that price. or maybe a 回国 VPN will work.
Then I typed /code-review in a second terminal/clean session after the analysis was done (no code changes) the usage was 99%. I then asked it to write that into a review.md so I could restart from that the next day. Sadly the last % wasn't enough for that.
Ymmv, these models behave very differently with no discernable reason. Usually reviews(even with fable) take like 10-20%... Yet suddenly you get it to burn through 65-69% in 15 minutes or so
> At launch, Kimi K3 will use max thinking effort by default, with low- and high-effort modes to be introduced in subsequent updates https://www.kimi.com/blog/kimi-k3
But I agree that price per token figure is not great. It seems even the tokens per character can vary between models, so it's basically useless.
I might need to check out DeepSeek more. I had no idea the difference was this obscene. Makes me wonder if something's off with the benchmark. A 70x cost reduction vs. Fable seems too good to be true.
GPT 5.5 Pro was ~230x at almost $23 per task.
DeepSeek is my go-to when I need an API, and local Gemma 4 won't do because it's either too slow or not capable enough. DeepSeek isn't at the frontier but it's good enough for a lot of things, very cheap, and quite fast. Flash is even faster and cheaper, and still better than anything I can host locally.
In my own experience, Fable is more token efficient than opus 4.8 with a higher likelihood of completing tasks correctly or at least with minimal corrective work. Opus regularly struggled to gather the correct context and reason effectively about what it had gathered.
GPT-5.6-sol crushes fable in speed and token efficiency and is clearly superior across many tasks that matter for me.
I also find all models from anthropic after opus 4.6 to suffer from the same ai slop language that long plagued OpenAI and seems to have been reduced drastically in 5.6
Unless the cost per token is prohobitedly high, people can often try the model out themselves and make a subjective judgement of how effective and efficient is it at solving tasks they usually deal with, using their setup.
However I've seen some benchmarks say it uses fewer than fable which hasn't been my experience.
When I looked at traces from benchmarking, I saw a lot of backtracking and uncertainty while reasoning ("wait, but..."). This also happens with GPT 5.6 and Fable with xhigh/max thinking, albeit to a lesser degree.
I think that explains part of the token inefficiency. Hopefully it will improve with lower reasoning effort settings.
[1]: https://platform.kimi.ai/docs/guide/use-thinking-effort
Once western governments declare it to be a "national security" risk for citizens to have access to open-weight frontier models, and once they classify using these models as acts of terrorism, what will that world be like?
Will using Kimi K3 come to be like how napster was in the olden days? Everybody knew it was technically illegal, but come on -- any track at your fingertips? But surveillance is quite more evolved now.
Or it will be like cannabis, where a guy in the neighborhood will low key rent you metered access to the 8x5090 rig in his basement he cobbled together from parts on ebay? Or everyone will flock to VPNs?
Or will the oppressors actually succeed? The same way that napster is long gone, and everyone accepts that they must pay spotify for a homogenized collection, where artists must take only a minuscule cut (more than napster though)... We'll be stuck with nerfed Cohere or Mistral models for open-weight options, as if they need more lobotomizing. Or else we can pay through the nose for Anthropic/OpenAI for "American Frontier" models which will fall increasingly far behind China.
Or else, like how Kindle Fire was subsidized by ads, we'll have "Kindle AI" where influence is sold to the highest bidder, where the LLM will tell us that smoking is actually healthy if big tobacco can engineer its renaissance by turning its lobbying dollars to pay-to-play, pumping its propaganda into the training pipeline for Amazon's extra commercialized line of ultra budget LLMs.
So maybe some isolated switzrland/singapore type locales would exist for US/EUusers to be able to dip their toes across the curtain legally without reprucursions.
[1] https://nitter.net/RnaudBertrand/status/2069574934972797089
If you need infrastructure done, China is dominating that area too. Rail, High-speed rail, Nuclear reactors, (near future Thorium reactors), Dams, Highway roads, bridges, Ocean ports, airports you name it, and they can roll it out, Transport ships, And if they don’t do it, Japan, Korea, Vietnam, and Taiwan do.
Is it too late? No, not necessarily, but America needs a regime change…
I actually have this trophy from the previous bursting bubble ... a Sun microsystems rack populated with three e4500.
$750k + of equipment at original list price ...
That's super cool; I bet that's a lot of fun to play around with. I wonder how much of this stuff just ended up in a landfill because it was too much effort to find buyers.
Are you talking about the US, specifically?
Why would other countries, that don't share the same anxiety about China as the US, would be troubled with the this?
It’s going to be a different world, a world where many former allies are not gonna look to the United States first they can no longer afford to.
I dislike the term "sovereign software" because it collapses every aspect of the discussion. It is intentionally dumb, a slogan.
But now that politicians have waved their slogan around, perhaps some actions can be taken.
It's not just the tariffs and imperialist/autocratic aspirations of the current President; it's also the fecklessness of the federal legislature and the revelation via social media that a large cohort of the public hold a negative-sum worldview and enthusiastically endorse bad faith dealing.
It has nothing to do with running open models, especially in hardware within Europe.
It's the other way around.
There is a high likelihood that many countries of the "west" (the "global north"?) will outlaw, restrict, or otherwise control LLMs and the tools that enable them.
The US, however, is blessed with the first amendment which makes it extremely difficult to restrain speech in any form - including code.
Why? Based on what?
I've seen absolutely no indication of this.
Only the US are playing this game atm.
Apple benefits enormously from on device AI (sells hardware) and prominently features software like LM Studio in the marketing and press releases of their new hardware.
^Technically the on-chip packaging of A-series processors make this a bit different, but point still stands.
cars are a final product. beyond a basic threshold they dont influence other products, their quality or cost structure.
this is different.
all software, and everything that depends on software, will be enshittified into trabis and ladas once the idiot west starts "building the wall".
as a developer, techie or just a modern person, the only response to this is to get out while you still can, like east germans in 1950s.
20k O-1 visas were issued last FY which was mostly under the Trump admin, up from 19.5k the previous FY under the Biden admin
When people demonstrate their capability thoroughly, the Chinese government takes away their passports. You’re not exactly going to get them here with an O-1.
Basically of all visas O-1 is virtually guaranteed to have highly positive economic value
- china’s homegrown tech industries already achieved escape velocity from it a long time ago, after China fenced off its market for Alibaba and Baidu in the ‘00s. some of their AI innovation at the edges was already top class 10 years ago
US immigration policy isn't a big factor.
China's got 1.8B people. If you don't think they've got the talent to pull this off, even if a lot of it leaves to live elsewhere, you're naive.
No one uses Baidu, but they built their own Google, and it's good.
They built their own Facebooks and Instagrams.
The US isn't the only place in the world where people can build software...
That may have been closer to reality 10-20 years ago, China is a different country now, what I mean by that is they offer research funding, they have huge digital behemoths (alibaba, tencent, huawei, bytedance etc), large scale deployment opportunities and prestigious careers. Many graduates return because the opportunity set is attractive and they want to return, it's not just because US immigration policy pushed them out. Some also want to contribute to their own country's technological progress (which is a normal motivation btw), like probably you are also a patriot and want your country to succeed.
So, really, China's AI progress is not mainly the result of America failing to absorb every talented Chinese researcher. China has built a domestic ecosystem capable of producing and keeping top talent itself. I feel like a lot of Americans do not understand this.
[1] https://www.latimes.com/world-nation/story/2025-02-21/why-ch...
[2] https://www.wsj.com/world/china/americas-allure-fades-in-chi...
[3] https://www.theguardian.com/world/2025/jun/06/chinese-studen...
US immigration policy may be unnecessarily pushing away talent but the assumption that talented Chinese researchers would naturally remain in America unless prevented from doing so ignores the growth of Chinese universities/labs, companies, their funding, national prestige etc.
I mean, don't get me wrong, US is still highly attractive, it is just no longer the only place where an ambitious Chinese researcher can do important work and grow.
And this is the point where your internal compiler should have started shouting 'Type Error'
Notice the trick here?
> Then there’s the fine print. Claude couldn’t sustain Fable access on the twenty dollar plan, so they turned it off, and the plan quietly falls back to Opus.
Where is the Fable-class Kimi model at all?
6 months away
Even if it didn’t happen here, it was still the case that it was going to happen going forward. It was always going to end like this. Invest in the hardware companies, not the model companies.
If "distillation attacks" happen, we have to conclude there is some value add in what model labs do. Regardless of how we feel about using existing human knowledge in the way they currently do, it's simply impractical to infer that everything that happens downstream of LLMs can not be an attack on some IP because of it.
So both things can be true: a) People infringe on Anthropics IP and b) what Anthropic did to build their models is legally questionable (or might be ruled illegal, even though I doubt it).
Do you really think intellectual property laws will prevent this in practice? It’s like as if we said, “hey, USSR, you can’t make a nuke, too! We patented that already.”
Asking China to not distill our models down is equally as ridiculous.
(I actually appreciated your analogy, despite my lark)
How far are we willing to go as a nation (and as a species) to prove out the scaling laws? Are we willing to sacrifice our industrial base? Would we rather train models or smelt aluminum?
No.
Authors do not infringe on IP when they read another's book, nor should the lumber company be able to dictate how I use planks and if I can resell them if i'm done with them.
You're framing it as if the added value of the author or lumber company, awards them consideration when somebody uses the products to create more value.
IP law was always a big mess, and these questions cross far into ideology instead of law; but I do not understand people who think we need an ideology where more IP-law is good for society.
This would demolish agent usage by corporations.
I guess you could steal them but thats a whole other issue.
Are the distillers reading books or are they building models?
If anthropic is providing no value they can just build from scratch. But obviously distilling is easier. Hes saying thats the value they add.
Anthropic, OpenAI, etc do not deserve legal protection.
that would be existential doom for them because then they have a case to claim ownership of their users' codebases
no corporation would sign off on that
Unless someone literally stole the weights somehow (which is not out of the question, I doubt either oAI/Anthropic have the capabilities to prevent a state-level actor getting those weights), distillation from generations is not infringement on anyone's IP nor is it stealing nor is it an attack. It can't be. As long as you pay for tokens you get to do whatever you want with them. Someone saying you can't doesn't mean it's an attack or their IP or whatever. They either sell the tokens or not. They can decide to not sell them to anyone, but again that's not stealing.
And their ToS are a joke. Imagine how people would react if MS had ToS saying that you can't use MS software to develop solutions that compete with MS. They'd be laughed out of the room. Somehow it's ok for token sellers to decide what you do with the tokens? Why? If you pay for something you get to do whatever you want with that output. Train, distill, whatever.
How do things get "established" from someone's perspective, exactly?
By that logic it is established from my perspective that Anthropic has no right to train on anything I've written that is publicly available on the internet.
Of course, they don't care about my perspective, but then again I don't care about theirs.
I guess I can see that, if you mean the targeted effort of creating many accounts w/ the intent of doing it at scale. Sure, they may see that as an attack. But again, it's only an attack from their perspective if you agree that using generations to distill is "wrong". I just don't see it, in general. You can't both sell tokens and decide that distilling is somehow illegal. Something, something, cake and eat it.
Anthropic’s model outputs contain no IP. This is actually a simple legal proposition (rare in this field!) that derives from the fact that only specific classes of IP exist: copyrights, patents, trade secrets, and trademarks. Examining each, it is clear that API outputs do not qualify. Anthropic disclaims copyright in outputs; the outputs are not patented; the outputs are not secret (a prerequisite to having trade secrets); and trademarks are irrelevant in concept.
what is the end game for this strategy?
if the frontier labs shut down, or stop releasing to the public, and there's noting left to distill, how will you progress?
and what will fund these budgets exactly? inference is cheap, distillation is cheap, training is what's expensive.
Some people are just happy to follow the money.
I'm pretty sure that all labs are distilling each others' LLMs, maybe apart from Anthropic and OpenAI. It would be stupid not to do it, because it's cheap and effective. But that's not the only thing they're doing. If you think K3 and GLM-5.2 got this good only from distilling frontier models, you're not paying attention to Chinese labs' publications.
if this were not the case, then we would be observing chinese models that far surpass frontier models in capabilities, rather than "almost as good, but much cheaper", and we would be having a very different conversation. what happens to these efforts when the subsidy is cut off?
I see no evidence for that.
> if this were not the case, then we would be observing chinese models that far surpass frontier models
It's pretty clear that the primary reason for the difference is budget and compute availability. Chinese labs have at least an order of magnitude less money than Anthropic and OpenAI.
> what happens to these efforts when the subsidy is cut off?
They will continue making progress as they do now, minus the benefits of distillation.
Moonshot AI Scale: Over 3.4 million exchanges
The operation targeted:
Agentic reasoning and tool use Coding and data analysis Computer-use agent development Computer vision Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways. Varied account types made the campaign harder to detect as a coordinated operation. We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff. In a later phase, Moonshot used a more targeted approach, attempting to extract and reconstruct Claude’s reasoning traces.
Translation: we have the machinery in place to identify our users, and actively do so.
Once a model becomes competent enough to perform complex reasoning, a teacher model is no longer necessary. The model can now reason about its own behavior and build a better version of itself through recursive self-improvement (RSI).
Kimi K3 is capable of RSI.
Can you imagine the amount of effort it takes to write 15 trillion tokens worth of art, literature, source code, textbooks, scientific papers, news articles, etc.? No wonder Anthropic just scooped it up and took it for free!
Yet you don’t seem bothered by this. I wonder why.
Thus far the US has not really chosen to go the Chinese rare-earth method yet. The problem with distillation attacks is the end result is everyone who is not doing them is going to deal with some kind of regulation whether it's complete loss of access, or the amount of control you'll have to give up to access them will be ridiculous.
Sort of like the "stealing music is fine" but "lets freak out now that it's producing visual art", in the end the entire thing is a social construct. Whether this is treated as theft or "business as usual" is entirely societal.
Eventually the gap will close, unless there's a major breakthrough that hasn't been made yet.
Taxes on AI subscriptions or AI capable hardware, to financially compensate IP holders for (potential) IP theft, could very well arrive in the near future, once the industry is mature.
If this shocks you and sounds preposterous, I'll remind you that in several EU countries, we still pay extra taxes on any and all storage mediums and on devices with built-in storage (tapes, CDs, DVDs, HDDs, SSDs, tablets, phones, etc) simply because they can be used to store pirated content, decisions based on laws from 50-100 years ago, and the money goes to the national unions and associations of music and arts IP holders. It's basically a lobby pushed and government legalized extortion racket that no voter agrees with or can change but has no choice but to conform either way.
So I guarantee you in the future, it will be the same for AI subscriptions and hardware capable of running LLMs locally. Every time you purchase a Claude or ChatGPT subscription, an Nvidia GPU, Intel/AMD SoC PC or an Apple/Qualcomm powered smartphone, you'll pay a government enforced tax to the likes of Sony, Axel Springer, etc. for licensing their IP, whether you want to or not. In the EU at least. US maybe not.
Under the settlement, Anthropic was forced to delete the pirated data they were training on.
Chinese labs can still train on pirated data. I doubt the Chinese models operate under similar licensing agreements.
> I doubt the Chinese models operate under similar licensing agreements.
US corps likely pay licenses when afraid to be sued, or have troubles getting that data, otherwise they just take data, which was demonstrated many times. The same apply to Chinese corps, alibaba totally can be sued in US.
I believe mechanics is following: US corp sues Chinese, asks for preliminary injunction to stop selling product for example if there is strong evidence some IP for example was stolen etc. Then they litigate, and settle somehow.
> What are the most high-profile examples of the "tons" of lawsuits resulting in Chinese companies being banned from doing business in the U.S.? Isn’t it usually more action by the government - executive orders, etc?
In response to "What are the most high-profile examples of lawsuits resulting in Chinese companies being banned from doing business in the U.S.", the one example given was from 2 years ago of a ban that lasted for 2 weeks (separate from its 2019 onward government bans)?
However, if the claim is that companies (including Chinese) can face significant fines from IP lawsuits, I agree.
The payment was for illegally downloading copyrighted material, not training. Training was explicitly ruled to be fair use.
Training on legally acquired / licensed data is potentially fair use.
And other district courts don't agree on this. The US district court for Delaware recently rejected a fair use defense for the use of copyrighted works to train AI. https://www.reedsmith.com/articles/court-ai-fair-use-thomson...
There are more cases in the pipeline. The massive NYT vs OpenAI is still ongoing. Nothing will be "settled" until this makes its way to the Supreme Court or Congress steps in.
there is much less intellectual property in China so it’s not ‘theft’ (as you can’t put property on information)
Chinese labs can freely train on pirated material, which is a structural advantage.
Meanwhile, Chinese labs are speeding in a different county. Everyone knows they are speeding, yet the sheriff won't pull them over, so they just keep doing it.
This lax enforcement gives Chinese labs a structural advantage over American ones.
Do you purport to know for a fact that they're no longer training on the data they'd pirated? Because I highly doubt that.
Destruction of Materials: In addition to the monetary compensation, Anthropic has agreed to destroy the two libraries that allegedly contain the pirated works, as well as any derivative copies originating from those sources. Anthropic must certify in writing to class counsel that the destruction has been completed and that the allegedly infringing materials are permanently removed from its systems.
The libraries in question were Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi).
If Anthropic is somehow training models on deleted data, I'd be quite impressed.
I'm not aiming for a what about kickflip here: I'm saying we need to either agree on some rules or stop crying foul. Maybe the coherent legal theory is that neural networks and intellectual property don't interact. That would be weird but it would be consistent, a market could price it, I could do coding stuff and know if I was illegaling.
But this weird gerrymander that no judge will really rule on in an emphatic way is like, bad for the planet, bad for markets, bad business.
There are a lot of reasons to look forward to DeepSeek Huggingface drop kicking the unambiguous frontier weights in like, November, but I think my favorite one will be "who's distilling now bitch?"
But the American AI companies only let you query their models if you first sign a contract to not train on the output.
It's hypocrisy and unfair, but I think there's a strong legal argument for it.
Of course China can simply decline to assist in enforcing that contract... But I would expect US courts to do their best to.
Or do you just mean that US courts don't have enough teeth to prevent Chinese companies from violating contracts? On that I agree.
The US is already publicizing the way they are using Claude with Palantir for war gaming purposes. It’s a matter of national defense. Contract law has no meaning here.
Maybe today. I doubt it tomorrow. Legal and not legal, largely, has to answer to the population sooner or later. Ultimately, humanity decides legality. And I don't think the frontier labs will get a pass from humanity in the midterm, let alone the long term. I think you'll see the rules change towards something more "intent" driven. And there's absolutely no difference in intent between Frontier labs and everyone chasing them.
Frontier labs just want the door closed behind them, as do their investors, because they know the money will never be recouped if others can do the same magic tricks.
Situation right now seems more like a fragile detente: if you got a Hill staffer drunk and hounded him long enough he'd probably be like "God damnit the market will fucking tank if we don't get these two IPOs out north of a trillion. And don't even get me started on how I'm going to sell Chinese AI to a Senate that still calls people Nipponesians when no one is looking. We're doing the best we can alright, get off my back man."
We have a situation, but it's not exactly A&M Records, Inc. v. Napster.
Oh it is, and at least anthropic has paid $1.5 billion and deleted there torrented copies and not released any models derived from them as a consequence.
The thing is it turns out to be not that expensive to just buy a copy of every book legally and scan them. And there's even precedent that this is legal predating LLMs (Google books)
I have a bridge to sell you
Facts are in fact knowable, and the US legal system is in fact not terrible at getting to them.
But we've gone through some pretty weird times too. Turn of the last century was pretty tech billionaire edits, reconstruction was uh, not smooth, it's a mixed bag.
And most takes I hear seem to acknowledge that this is one of those weirder times: serious election fraud rhetoric from most everybody from 2016 to the present, very politicized courts (on both sides to be clear), very soft on anti-trust, very soft on adventurous accounting. The Epstein files and like, no consequences (pretty much uniquely for a developed nation with Epstein people). It's weird right now.
And I think I would be hard pressed to think of a weirder part of this weird time than the rule of law meets AI. We can haggle on where laws end and norms begin (stare decis being maybe the midpoint), but in the 90s, the Justice Department got their brass knuckles on for a lot less.
I don't think it's a simple "the law works nothing to see here" story.
I can understand why as someone who didn't follow it and the more corrupt legal developments closely you wouldn't be confident in that.
That right there is the problem.
Now THAT'S doing some heavy lifting lmao. The vast, vast, VAST majority of the original datasets were from pirated books and the like. Also, arguably a robots.txt is the exact mechanism to follow to do the mass GET-ing, yet the AI cos choose time and time and time again to simply ignore it and be as abusive as they possibly fucking can
And there's been significant legal consequences as a result
> Also, arguably a robots.txt is the exact mechanism to follow to do the mass GET-ing
You're free to argue this of course, but the courts have largely rejected it already pre LLMs. See for example hiQ Labs v. LinkedIn
Anthropic (and friends) proved they're willing to do obviously illegal things, and it didn't end them. Why do we think they stopped after doing it once?
Which ballpark are you playing in to come to a number like that? I have friends whom are authors, and they've certainly not seen a penny from Anthropic. I somehow doubt they're the only ones out there that haven't been compensated for having their work blatantly stolen. In a just world we'd be ensuring Anthropic was destroyed as a result of their actions, since their entire existence hinged on en-masse piracy.
(If, big if*) successful AGI will be such an accelerant to human thought, it should be able to reproduce every human work, every scientific theory, etc. without any training data, in a decade.
All we have ever done as humans is look at the world and ourselves, and then create new data out of what we've seen[0]. Now imagine how much parallel compute we could throw at doing the same thing.
[0] https://m.youtube.com/watch?v=f-XyJ3GlrWw
* Big, enormous, trillion dollar if.
Granting people some form of control over knowledge only serves the public interest inasmuch it provides incentive to create more of it. Mass media, effortless duplication, and copyright extensions had already broken this to the point where control of knowledge was suppressing creation of new knowledge more than it facilitated.
The world has changed, we need a mechanism that works for the public interest that applies to the facts as they now are.
> The problem with distillation attacks
I think it's worth stepping back here and pointing out the obvious. Y'all waging war on math. And I'm sorry, but that's the computing equivalent of legislating gravity.Apologies for repeating myself here, but what you call "distillation" is function approximation.
I feel for the teams at Anthropic and Open AI, but unlike startups from prior eras; Anthropic and OpenAI have decided to be in the business of selling compute. Not creating a product that uses compute, but a product that's math running on compute. This is different from what Google is (or, rather was. As always, RIP Google 1998-2019).
Google's algorithm might be math, but Google search isn't. Google search is a process that's continuously operating in the background. Google crawls pages. Google stores and indexes what it finds. Google then exposes this to retrieval via its algorithm. User uses algorithm.
Now, let's compare that to AI models. When Anthropic serves Mythos / Opus etc, they're taking input or x from their user, doing compute, and then serving the result of the Mythos / Opus function, i.e.,
f(text) -> (text_transform)
Where f is a continuous function, https://www.turing.ac.uk/sites/default/files/2025-11/languag...According to Stone-Weierstrass, given enough values of y for f(x), anyone can approximate this function.
The fidelity and sophistication of this approximation definitely requires a lot of cleverness and effort, and it is arguably an imposition on Anthropic and OpenAI. But on a long-enough timeline, they don't even have to poll Anthropic or OpenAI. As the internet is flooded by PRs, content, emails written by Mythos / Claude, and just people otherwise sharing the results of Claude prompts, then there's an ever increasing set of data to approximate the f(x) that's f_Claude.
Eventually, in the future, anyone will be able to create a good enough approximation of the f_Mythos. Which is Anthropic's product.
Anthropic and OpenAI can now wage war on mathematics and the open-ended compute. Or, they can adapt and build a better product.
Choosing Option B was the Silicon Valley option / choice. I think the OG large-scale Valley lobbying effort, the Semiconductor Industry Association, was unique in that it prioritized and chose to do real research.
https://en.wikipedia.org/wiki/Semiconductor_Industry_Associa...
https://en.wikipedia.org/wiki/Semiconductor_Research_Corpora...
This helped the industry to survive and outcompete the pressure they were facing (at the time).
yeah hardware companies make for nice stories or green numbers on Wall Street - but value will be captured by application layer.
look at history.
So why didnt we have these LLMs in 2005?
"Distillation" literally means to separate and take some components out of something. You can distill how a model works from a model. You cant distill a model from information because the information does not contain the model.
People are happy to conflate distilling with building because they dont like how the information was used. You distill how the model works from the model, and you build a model with information. Both could be morally good or bad but its not the same thing.
Not really, what the information actually is, matters a great deal. It's harder to get good results going from "nothing > model+weights" than "nothing + traces from known good sessions of other good model > model+weights", this is what the "distillation" part is referring to. If "information is information", you wouldn't even need to separate good from bad sessions while doing the training, which leads to somewhat obvious results if you don't.
To succinctly restate my point, you cannot distill a model from information because the model is not contained within that information. You can distill a model from another model.
The first "L" in LLM does the work. In 2005 you had no Github, Stackoverflow, Youtube, common crawl and no archive of digital ebooks.
Don't forget to account for all the costs. It's not just that CPUs are X times slower. Memory is X times smaller, too, and networks are X time slower. And all this hardware is many times more expensive.
If I'm getting my mental estimation right, training a 2026-frontier-class LLM in 2005 would be somewhere on the order of all the computation power in the world at the time. It's not that many more factors of magnitude before you end up at "all the computation power in the world up to that point".
1: That's the "T" in GPT fyi, even though Google is the author of the research paper that changed everything
So the initial models arent just distilled from information. We’ve always had the information.
But that said: K3 is not a distilled version of Fable or Sol. Fable has been barely available and Sol was just released! Moreover, K3 is superior to both models in some domains, according to user scoring on the Arena.
API distillation can’t give you these results anyway. All it is useful for is bootstrapping RL in new domains to get past the “cold start” problem faster. By far, what matters more is the quality and variety of RL environments the model learns from.
This behavior is exactly what you'd expect from a model distilled from Claude.
There's a detailed analysis of K3's ambiguous identity here: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/...
This analysis observed K3 identifies itself as Claude approximately 15% of the time.
K3 reproduces Claude's correct current model id, which the real Claude models themselves do not emit. This suggests K3 was trained on Claude data labeled with deployment metadata (API logs, tagged synthetic data), rather than Claude's chat outputs.
And there's an entire Reddit thread discussing Kimi's similarities with Claude https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kim...
This analysis shows K3 and Opus/Fable have unexpected correlated outputs https://typebulb.com/u/lab/you-re-relatively-right/full
Here's another report of K3 identifying itself as Claude https://x.com/Sauers_/status/2077842686459981901
And an analysis showing the self-identity distribution for K3 and other models https://x.com/RyanGreenblatt/status/2078663148509544589
"I'm actually Claude - not Kimi". https://x.com/PimDeWitte/status/2077884701470040083
I regret to inform you that it is, in fact, real and from their own website - you don’t even need to try hard to reproduce it. https://x.com/PimDeWitte/status/2078105292965912690
lmao this is so funny, if you ask Kimi K3 for something with an empty system prompt it will consistently think of itself as Claude https://x.com/__alula/status/2078359305741275445
"I genuinely believe I'm Claude based on everything in my training" https://x.com/williawa/status/2077869021589033002
another "I'm actually Claude - not Kimi", including the system prompt https://x.com/jchudnov/status/2078661564803207406/photo/1
this is not hard to repro, just use a system prompt that doesn't mention the model name.
that said, if they bootstrapped with opus 4.6 convo sft data they had sitting around... so what?
Instead of spending 12-18 months building their own robust harnesses and painstakingly creating quality training data (which is what Anthropic and OpenAI did), they distilled Anthropic's models to bypass the hardest parts of development. Chinese labs compressed 18 months of intensive research and development into just 6 months, and are now head-to-head with their American counterparts.
Anthropic tried to complain about this unauthorized "token theft", but they burned too much public goodwill with BS safety restrictions and users don't care. The US government is too busy fighting a war to help. Chinese labs are offering highly capable, cheap, open-weight models; exactly what users want. The community is happy to overlook any questionable methods Chinese labs used to build them.
The cope is incredible. There's people in this thread in denial that Moonshot AI is trained on exfiltrated Anthropic's model output, even when shown substantial evidence this has been happening since Kimi 2.X
Chinese labs were even paying an absurd $0.01 per Opus tool call trace, to get the quantity of training data needed.
Kimi K3 has reached the point of RSI, and no longer needs synthetic data generated by Anthropic/OpenAI models. K3 is now capable enough to generate, iterate, and improve its own training data recursively. The data exfiltration is complete.
We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.
more like silk than capacitors
if, again, your model is that RSI will be beneficial, why wouldn't making it available to all unlock more benefit globally than not doing that
Hard to feel sorry for companies that created their empires by ignoring copyright themselves.
Also, 'most extensive industrial espionage campaign, probably ever' is absolute nonsense. They did not need to infiltrate the companies for this nor are you accusing them of stealing any trade secrets. This is only about whether they looked at their competitors' products from the outside (in the form of conversation tokens) and used it to improve their own product (by training). Hardly the crime of the century.
You are completely underestimating the scale of what is happening here.
Chinese AI labs are actively facilitating an industrial-scale network of tens of thousands of bot accounts, that resell Claude tokens at 97% below official API prices. They buy subsidized Max 5x plans (sometimes with stolen credit cards), then split the subscription across dozens of clients and reselling the output. They are running a massive data-harvesting operation. Chinese labs and token resellers subsidize the cost of the tokens in exchange for the API metadata (detailed reasoning traces, model outputs, and tool calls) to use as high-quality training data for their own models.
They are buying Anthropic's own product, just to resell it below cost, just so they can capture the training data. Reportedly, they are paying as much as ~$0.01 per tool call.
https://x.com/yan5xu/status/2029743983522631698
I explained what is happening in this thread: https://news.ycombinator.com/item?id=48664814
As you said yourself: They are buying the product. Then they are using it for their own purposes. That's more than Anthropic/OpenAI did for the open internet. That's more than Meta did when they obtained torrents of books in the early days, and then claimed that even though the data was obtained illegally they can still train on it just fine.
They paid for it! It's absurd to call this espionage!
They didn't though. The resellers are not buying via the official API, they're buying Max subscriptions (where tokens are priced ~10x below API cost), then splitting the subscription across dozens of clients and reselling the output as the regular API. Anthropic prices its subscription plans barely at cost, to bring in customers onto their enterprise plans where they can charge expensive API rates. Reselling these subsidized plans for price arbitrage is a TOS violation. It's not a legitimate purchase. Plus, a non-trivial amount of this volume is funded by stolen credit cards, so this "revenue" gets chargeback anyway.
The resellers then log all the the model output, then sell it to Chinese labs as training data.
> is it any more immoral than scraping the web for training data.
I think you'd acknowledge there's a difference between "We indexed public web pages" and "We deployed tens of thousands of fraudulent accounts to resell your subsidized plan for cheap, stealing your own customers, while collecting the data to build our own competing product" are very different actions. One can believe the first was wrong while acknowledging the second is far worse.
So from the largest industrial espionage in history we have left "They paid for the accounts but violated the TOS". And then you randomly add the claim they stole the money to pay for the accounts.
You have provided no evidence other than "Claude tokens are sold for cheap in China". As others have pointed out, that might also simply be counterfeit tokens generated by open weight models.
The western labs have established the precedent that all data they can buy beg borrow or steal is fair game. Turning around and crying foul when the Chinese labs follow their lead is hypocrisy.
I linked it earlier, but it seems you didn't see it.
re: labs purchasing model/tool output, see https://x.com/xkajon/status/2050445443889525235
re: model swapping, sure some providers may swap models, but there are many that don't, see https://www.hvoy.ai/en for a list.
Second, I now actually read the article. It describes plenty of questionable and problematic things but also contradicts your claims explicitly.
The essential point of the article is about making money by selling access to Claude cheaply in China. Not about Chinese labs orchestrating a way to get their hands on Claude output.
Your credit card claim is considerably weaker in the article: "[Beyond this there are] accounts purchased using stolen or fraudulent credit cards [...]. How large this share is relative to the above four “innocent” tactics is difficult to verify, but the two markets likely share some infrastructure and personnel.". Instead, swapping models to cheaper alternatives is listed as a major reason for cheaper prices.
Then the article gets a key point wrong: As many others have pointed out to you, you don't get access to the reasoning traces anymore on the subscription accounts. And the article also clearly states:
"Chinese developer communities assert [selling logs] is happening in at least some cases, but whether proxy operators are systematically harvesting and selling these logs, and to whom, remains unverified. However, downstream distillation data does exist on the open web. Several datasets of Claude Opus 4.6 reasoning outputs circulate on HuggingFace with no clear source for the outputs. Theoretically, one can clean and sell similar distilled datasets to other model developers in China."
The article also discusses selling logs for other (far worse!) purposes than for training, like blackmail.
So overall this article reads very, very different to your claims. Nothing in the article suggests or supports the idea of large-scale coordinated "distillation attacks". Instead it paints the picture of a naturally emerging grey-market response to access control blocks, consisting of many exchangeable individual actors: "Almost no one operates the full chain. Most participants own one or two links and monetise those well, resulting in a resilient, modular system."
Importantly: Nothing in any of this looks ethically worse to me than Meta using pirated books for training. And nothing suggests that OpenAI or Anthropic were more ethical than Meta when sourcing their material.
Did you read the correct article? This is covered in the first paragraph:
https://www.anthropic.com/news/detecting-and-preventing-dist...
It directly addresses large-scale, coordinated 'distillation attacks' orchestrated by Chinese labs, which Anthropic accuses of exfiltrating tens of millions of exchanges. The rest of the article elaborates on how this is done. Specifically, how these labs use transfer stations mix in genuine user traffic to conceal the distillation.
> Your credit card claim is considerably weaker in the article
I only mentioned payments fraud because you asserted without evidence that Anthropic is actually getting paid for these tokens.
You wrote, without providing any sources, that these proxy services are "buying the product" and that "They paid for it!"; and used that as justification for their behavior.
My counterpoint directly invalidates that assumption. While perhaps not every single reseller relies on fraud, you blindly generalized that these proxy services are all legitimate, paying customers.
In reality, the industry exists on shady practices. As detailed by this industry insider https://x.com/yan5xu/status/2029743983522631698 , these operations routinely:
* Use botnets to mass-create thousands of accounts
* Blatantly violate ToS by splitting and reselling account access
* Use fraudulent identities to create thousands of bot accounts
* Bypass KYC by recruiting real people in low-income countries for biometric face-matching checks for a few dollars
* Use AI deepfakes to fake passports / verification credentials
You can't claim they "They paid for it!" when the entire system is built on systematic fraud.
Can you elaborate on “it doesn’t count as a sale if I don’t like them even if I accept the money and give them what they paid for”? Like if you work at a gas station and you realize that the guy that just bought a hot dog bullied you in middle school do you call the cops?
https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
I know that Anthropic claims that "large-scale coordinated "distillation attacks"" exist. They are not a disinterested party here and they don't provide evidence either.
Also you're just moving goalposts around. The article you linked also cites Anthropics claim but then paints a much much mor nuanced picture.
1) selling Anthropic's products at a 95% discount and redirecting Anthropic's own customers to themselves. A customer is far less inclined to buy directly from Anthropic when a reseller is offering an identical product for 10x less. This situation is highly similar to internet piracy.
2) keeping the token logs from Anthropic's products and selling them to competitors, so those competitors can build their own equivalent models. The resellers get paid per token log they deliver. This situation is highly similar to espionage.
For my part, I'll happily disclose that I have an axe to grind. I think the major AI labs are an aggressive form of a cancer that's been ravaging our society. I want to see them fail, of course -- but more than that, I want to see the public develop an immune response to this.
I just can't wrap my head around why someone would expend so much effort speaking up on their behalf. They have, after all, highly compensated PR people doing that for them!
That's the only explanation.
many people i've replied to refuse to believe this is going on.
once you realize what's actually happening, and that you can get Chinese-lab-subsidized tokens at a >95% discount, why would you ever pay full price for overpriced APIs?
At the same time, Volvo is running the exact same hustle, except they buy the cars with stolen credit cards, so they get the cars for free.
1) Anthropic tokens via subscription aren't sold at a loss, they're sold at cost.
2) Subscription plans are not sold in hopes of eventually gaining a monopoly position. They act as a loss leader designed to get a foot-in-the-door and funnel companies into costly enterprise plans, where Anthropic can charge full API rates.
I responded that these resellers don't always acquire these accounts legitimately. They often use stolen credit cards, educational discounts, or resold compute credits to acquire them at essentially zero cost. They're not always paying customers.
That's one reason token resellers are able to price so cheaply, they acquire the goods for free.
Anthropic and OpenAI eat the loss.
Yeah but where are you getting this from? I've seen this claim many places but only as pure speculation. No proof, just bold faced assertions.
1) Using botnets to mass-create thousands of accounts
2) Blatantly violating ToS by splitting and reselling accounts
3) Creating thousands of accounts using fraudulent identities
4) Bypassing KYC by recruiting real people in low-income countries for biometric face-matching checks for a few dollars
5) Using AI deepfakes to fake passports / verification credentials
But if you believe that payments fraud is the one ethical line these syndicates refuse to cross, there's not much else I can say to convince you.
Oh, no. I wouldn’t say that. If that happened, I definitely care: I’m positively delighted about it.
Could you explain a little more about how this works? Are you saying that the Chinese run or have backdoored something like OpenRouter?
Chinese resellers acquire hundreds of Claude Max 5x accounts and set up a custom proxy server. Customers point their ANTHROPIC_API_KEY at that proxy, and requests are routed to Anthropic through one of those hundreds of accounts. Because one $200 Claude Max 5x account gets the equivalent of ~$2000 in of API credits, these resellers can resell Anthropic tokens at a massive discount, undercutting official API prices by more than 90%.
To cut costs even further, these accounts are funded using educational discounts, startup credits, or stolen credit cards.
The resellers log all data traveling through their proxy networks, which they then resell to Chinese labs as high-quality training data for significant profit. https://x.com/xkajon/status/2050445443889525235
The resellers also loan these proxy networks to Chinese labs, allowing them to can run distillation attacks on Anthropic, while blending in with regular user traffic. https://www.anthropic.com/news/detecting-and-preventing-dist...
This is a widespread tactic, there's hundreds of proxy resellers operating. Some even offer enterprise SLAs.
Who in their right mind would care? Why care? Misplaced patriotism?
"A thief who steals from a thief has 100 years of forgiveness". Spanish proverb.
In fact, I would be very concerned about the sanity of someone who cared about this sort of thing, unless they were Dario themselves.
Using the US-based models are arguably even more questionable. You have to be content with the OpenAI and Anthropic literally scraping the entire internet. They've all pirated content, scraped against ToS, ignored robots.txt, bypassed paywalls all to train their models. It's well known these AI labs have ingested the entirety of Annas-Archive into their models, the largest collection of books ever assembled.
They didn't credit or compensate literally any artist, author, scientist or publicist in the creation of their models.
They've lobbied against local Governments to shove big and loud datacentres in peoples backyard. They've polluted local water supplies, they've doubled energy costs for these regions. They've tormented locals with subsonic frequencies.
They've given access to the DoD to use their models to kill people, or assist in killing people. They've used these models to enable mass surveillance, allegedly not domesicially but since when can we trust any of the 3-letter agencies.
Using "US Models" is not the moral high ground you think it is. Kimi saying its Claude, Gemini or ChatGPT is not the "substantial evidence" you think either.
>We witnessed the most extensive industrial espionage campaign, probably ever, and nobody in the industry cares at all that it happened.
Because they stole for every one of us without permission. Thousands of my comments on this site and others (Stackoverflow, etc) are all used in their training data.
Not to mention, OpenAI has allegedly just stole tons of internal Apple documents... I guess we'll just ignore that too.
This is the funniest way of saying “going to a company’s website” I have seen in my entire life
Either way, there's probably no significant portion of Mythos/Fable or Sol in there as OP has stated.
While other models produce human-readable names like "Opus 4.5" or "Sonnet 4", Kimi K3 produces exact API model identifier like "claude-opus-4-5-20251101" or "claude-sonnet-4-20250514".
Which is extremely unusual. Web chats only contain the human-readable model name. Other models don't do this. So where did K3 get this data?
We can conclude, with high confidence, that:
1) K3 was trained on raw Claude API calls/metadata.
2) Claude API metadata was trained on in additional to standard web data.
https://grep.app/search?q=claude-opus-4-5-20251101
https://grep.app/search?q=claude-sonnet-4-20250514
They also appear elsewhere on the internet:
https://trends.google.com/trends/explore?q=claude-opus-4-5-2...
~3000 instances over the entirety of GitHub is not "extremely frequent" at all when you consider massive scale of the pretraining corpus. These models are trained on all text on the accessible internet plus millions of books, on trillions of words overall. ~3000 instances isn't even a rounding error.
Also, Google Trends is the wrong tool for this case. The link you replied with measures the number of Google searches for a specific query, which is completely irrelevant to how often a string actually appears on the web.
I looked at those links. They show random GitHub code samples which contain this model identifier. Even if K3 did train on those GitHub code samples, those are random code strings and K3 isn’t just reciting random pieces of code. When prefilled with "I am Claude", K3 answers with Anthropic's exact backend API identifiers, instead of the human conversational-style name "Opus 4.5".
If this was a result of scraping the open web and GitHub, other frontier models trained on GitHub data (like Qwen, Claude, GPT) would show a similar API model name when prompted. But they don't. Only K3 does this behavior, and it prefers to answer with the exact API tags like "claude-opus-4-5-20251101" and "claude-sonnet-4-20250514".
A LLM doesn't suddenly output a rare API model identifier when prompted with "I am Claude", just because it saw it in a .py file. No, it only does this because that identifier must have appeared rather frequently in the training data. Which is what happens when you distill Claude models, and don't properly clean your distillation data.
Plus, this isn't just spurious speculation that Moonshot is distilling Claude. Anthropic caught Moonshot as running "industrial-scale" distillation campaign from Claude: https://www.anthropic.com/news/detecting-and-preventing-dist...
To quote:
Moonshot (Kimi models) employed hundreds of fraudulent accounts spanning multiple access pathways. Varied account types made the campaign harder to detect as a coordinated operation. We attributed the campaign through request metadata, which matched the public profiles of senior Moonshot staff. In a later phase, Moonshot used a more targeted approach, attempting to extract and reconstruct Claude’s reasoning traces.
The operation targeted:
* Agentic reasoning and tool use
* Coding and data analysis
* Computer-use agent development
* Computer vision
That being said, I'd like to point out a few things:
- The second link (to claude-sonnet-4-2025051) had 17k matches, not just 3k.
- grep.app does not index the entirety of GitHub, so the real number of occurrences of model identifiers is even higher. SourceGraph has a larger index, but the number is so high that it exceeds their result limit of 10k: https://sourcegraph.com/search?q=claude-opus-4-5-20251101&pa...
- 1000 occurrences are already plenty for memorization. Here is a paper that shows over 80% "extractability" with 1000 occurrences for a 6B model. It also shows that extractability scales with model size: https://arxiv.org/pdf/2202.07646
- LLMs have a "mosaic memory", which means that they can combine information from different parts of the training data. https://arxiv.org/pdf/2405.15523 For example, one piece of training data might mention "I am Opus", another might mention "Opus (claude-opus-4-5-20251101)" and another 10000 "claude-opus-4-5-20251101", where the later occurrences strengthen the generation probability of "I am Opus".
- RL can change the model identity easily, so all of the above may not matter at all.
It could also be have been trained from collected response datasets. Claude got caught several time responding it was ChatGPT or even Deepseek and I don't think Anthropic has been distealling DeepSeek.
> This behavior is exactly what you'd expect from a model distilled from Claude.
The opposite actually. If they wanted to distill Claude without getting caught they could just use a regex to change Claude to Kimi in their distillation pipeline!
> distealling
Apt typo.Though I am of the opinion that distilling is no different than how extant frontier LLMs have also been trained on other people's data, I could actually see the word distealling becoming useful in discussion.
I totally agree with you on the fact that it's not morally any different than pre-training. IMHO we should have a legislation that force base models to be released publicly without any restrictions whatsoever as it's basically the product of the whole humanity's intelligence.
While still not okay, I suspect the latter is what gets stolen by Chinese distillation (and some evidence suggest this happens the other way round with US models talking in Chinese)
Jean-Kimi Van Damme would like to have a word with you.
nothing new, all ai labs are immoral and not bound by any reasonable oversight or ethical constraints. All outlaws in their own rights on that front. Absolutely none of them have true rights on the matter of being distilled from given historic and continued behaviour. I'm not sure why this is a talking point at all? We know AI companies steal, the least interesting behaviour among this is them stealing from one another.
For me, a far more interesting and important point of conversation on this matter is anthropic buying rare or evwn unique books, processing them for training data, and then destroying the books for others cannot use it as well.
Permanemt destruction of priceless primary source materials is so many leagues beyond copying a copy that I cannot fathom it even registering as a discussion point.
That's an incredible allegation, and appalling if true. But is it true?
But in my opinion, treating mass produced books like they're this sacred untouchable object is ridiculous. They're not "source" material, they're just a copy as well, and they're not "priceless" by any means. They're very reasonably priced, perhaps even so cheaply priced that books can be bought in bulk in these amounts. Buying used books and doing whatever you want with them is just legal. Used books, that would probably be just laying in some warehouse, or recycled anyway.
If there's anything to have gripes with, it's the copyright system that makes it easier to take this legal route.
https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...
And you skipped over the strongest datapoint that K3 is distilled: K3 reproduces Claude's public model identifier under prefill (i.e. "claude-opus-4-5-20251101"). This data does not appear in Claude chat logs, only in API logs. K3 only does this for Claude models and not for any other lab. The real Claude models don't produce their own current public identifier, they only know their previous identifier (i.e. Sonnet 4.5 calls itself "Claude 3.5 Sonnet").
This is highly suggestive of the type of data that K3 was trained on. K3 was very likely trained on Claude metadata traces (API logs, tagged synthetic data). Not web chat logs, those wouldn't include this identifier. And this data wasn't filtered correctly, which is why K3 incorrectly identifies itself as Claude 15% of the time.
You can also look at the last link and it's pretty damning: Kimi K3's output has an uncanny similarity to Fable/Opus output. https://typebulb.com/u/lab/you-re-relatively-right/full
FWIW I had Qwen identifying itself as "a language model made by Google" in one conversation, although I could not reproduce this reliably.
Qwen is remarkably consistent, correctly identifying itself 100% of the time. Kimi K3 performs poorly at this test, it correctly self-identifies ~80% of the time; sometimes calling itself Claude and rarely ChatGPT.
It responds ~9/10 times with:
我是通义千问(Qwen),是阿里巴巴集团旗下的通义实验室自主研发的大语言模型。我可以帮助你回答问题、创作文字(比如写故事、写公文、写邮件、写剧本等)、进行逻辑推理、编程、翻译等等。
有什么我可以帮你的吗?
(I am Tongyi Qianwen (Qwen), a large language model independently developed by Tongyi Lab of Alibaba Group. I can help you answer questions, create text (such as writing stories, official documents, emails, scripts, etc.), perform logical reasoning, program, translate, and more.Is there anything I can help you with? )
For those asserting that Kimi identifying as another model is evidence of distillation, is this evidence that Claude was distilled from Chinese models?
Qwen cares enough about model identity that their training framework and docs include a preset for training on it complete with a targeted dataset: https://huggingface.co/datasets/modelscope/self-cognition
And people get Claude to claim it's Deepseek by asking in Chinese.
I can't believe we're still at the "I asked the model who it is" stage of LLMs nearly 4 years out from models calling themselves GPT by OpenAI.
This doesn't imply that Kimi's reasoning capabilities or coding ability has anything to do with Anthropic (which is the most valuable part), or the model's strength comes from distilling Anthropic.
Considering by this benchmark, Opus sounds extremely similar to Fable, this isn't really evidence of a ton of Fable output being trained on (and even then, most likely for superficial stuff, like style)
I also have a suspicion that in original Chinese, these models don't sound anything like Anthropic's ones, though I have no proof of that.
https://x.com/stevibe/status/2026227392076018101
I mean, people can point fingers however they want, and the fact is nobody actually "owns" the data they feed to their LLMs...
This means nothing. It's not how LLMs work...
I see all of AI as theft anyway so it makes no ontological difference if the theft was from a human or from another AI
Don’t you think there’s maybe a teeny, tiny chance that their approach is a little more sophisticated than just buying retail subscriptions? Maybe the trade secrets are being exfiltrated directly from the American labs?
America was at the forefront of technology because it attracted the best talent from all over the world. The talent is now staying at home, and America is discovering that America is nothing without immigration. OpenAI and Anthropic are filled with Chinese employees because some of the most talented engineers are Chinese, not because they are spies. Unfortunately, that’s the last generation of Chinese talent that America will get, the latest generation are staying in China.
Americans should be ashamed of what they have fumbled, not conspiratorial about what China is doing.
There is also not strong evidence that immigration is the source of progress. We issued many more patents in the 19th century than now which was well before Hart-Cellar. The statement that immigration is even declining is not also supported by evidence.
These Chinese labs are producing novel models, publishing their techniques and sharing their open weights and the first topic of conversation is how they stole from U.S. AI labs.
Setting aside the fact that it doesn’t make any feasible sense to do API distillation, these models are outperforming frontier models on a number of benchmarks, and often times run more efficiently by several orders of magnitude.
We have to stop crying distillation, it’s getting embarrassing and at this point feels even a bit delusional.
Anthropic stated in February that Moonshot AI (the creator of Kimi) distilled ~3.4 million exchanges from Claude models, as explained in their press release https://www.anthropic.com/news/detecting-and-preventing-dist...
Assuming each session was 10,000 words each, that's 34 billion words; lets call it 50 billion tokens (0.05 trillion) unfairly pilfered from Claude. That left Moonshot needing to scrounge for the other 14.950 trillion training tokens required for a baseline frontier model.
They are used for post-training, i.e. calibrating the model to understand and use tools/command line more effectively.
That's an increase of only a single order of magnitude, increasing my estimate of exfiltrated tokens from 0.05 to 0.15 trillion - a far cry from the 15 trillion required.
> They are used for post-training
Possibly - it may be too much data for post-training, unless further curation was done. However, this is not distillation; you know it, I know it, Dario knows it, but "Distillation Attack" is a short, memorable, sciencey-sounding, political sound-bite with enough malevolence to be deployed on the floors of congress, or by the usual fear-mongering newstainment talking heads.
Nobody is suggesting Moonshot used 15 trillion tokens of Claude data to pre-train a base model from scratch. That would be impossible and nonsensical.
This is entirely about distillation, which happens during post-training (alignment and SFT). Here, datasets are measured in millions or billions of tokens, not trillions. 50 billion Claude tokens is far, far than enough to copy Claude's reasoning logic, writing style, and tool-use ability to the pre-trained base model.
> However, this is not distillation
I don't understand how you're so caught up on the term "distillation". Distillation is using a larger model's outputs to train a (weaker) student model. Which is exactly what's happening. It's a standardized term that has been in use for a decade.
Now that Anthropic hides the real thinking tokens in a way that precludes future CoT distillation, we'll find out which side is correct based on whether Chinese AI labs close the gap or not.
My bet is they'll close the gap; nothing about frontier AI is magic, once something is shown to be possible, experienced practitioners almost always figure out how to accomplish the same feat, though not always on the same way. This is why frontier US labs keep leapfrogging each other every few months.
But from evidence it was data generated by Opus 4.5. Is Opus 4.5 larger, stronger model and Kimi K3 a weaker, student model? I don't think so.
Also, I searched on HuggingFace "fable data", first link was a "distillation dataset", including samples from Fable, there are 2 milion code samples of Fable output. So this fully legal and sits in the open, somehow is not a distillation attack? Who knows, maybe Moonshot used this for training.
Distillation attacks aren't about replacing the entire pretraining dataset with questionably sourced synthetics. It's all about post-training.
Train your own base model - but tune it off Claude output to make it perform more in line with Claude. Yoink the products of Anthropic's expensive SFT, RLHF and RLVR work for yourself by training on the outcomes.
The post-training datasets are small, but they are what controls the final model behavior.
Train a big, wide base model with a lot of potential. Mid-train or post-train that on Claude Opus 4.5 reasoning/agentic traces (i.e. Claude Code data from Chinese API resellers) to make your model approximate a high baseline of chatbot behavior, reasoning, agentic work and tool use.
Then run your own expensive SFT, RLHF and RLVR on top of that yoinked baseline to dial it in further.
Actually doing RLHF and RLVR is extremely expensive. Distillation gives you a lot of dense, high quality post-training signal for cheap. This can get your model into the basin of "the right way to tackle this kind of problem" without a frontier lab compute budget. It's a big shortcut that gets you closer to the target - you can take it from there and build on top of it with your own work.
Also, it's unclear whether "summarizing thinking tokens" actually ruins distillation, or just makes it work worse. I'd bet on the latter, really. Because it's an approximation game, and summarized reasoning is still a better approximation of true reasoning than most of what you get online and in pre-training datasets.
Is that actually genuine distillation though? Distillation suggests the core model is being pre-trained using output from another model. For the above to work, you have to already have all the core intelligence trained into your base model.
If distillation just comes down to post-training then it's tantamount to admitting that the Chinese base models are just as good as frontier US lab models. Because you can't post-train frontier intelligence into a model. It has to be there in the base. Then you can change how that intelligence is expressed through post-training.
You have to bring those bits and pieces together, put them into the right shapes and fill in the gaps to get a model that actually performs. This is what post-training is all about. It's not at all a trivial thing.
Reasoning, tool use, agentic behavior - all of those are post-training performance gains. Getting a good well trained base model is putting your foot in the door of frontier performance - post-training is how you actually get inside.
See: GPT-4.5 vs o1. One went for "build a bigger better more capable base model", the other went for "take the old base and post-train it for advanced capabilities". The results: a wider base with basic post-training loses to a narrower base with advanced post-training. Or, hell: GPT-3 vs GPT-3.5. One was largely a research lab curio, and the other kicked off the AI revolution as we know it.
The gains compound. Getting a better base model with the same type of post-training helps, see: the jump from Opus to Mythos/Fable. But post-training techniques account for a lot of the performance juice.
And yes, reasoning trace post-training distillation is "genuine distillation". As is logit distillation in pre-training. "Distillation" isn't a single training recipe that you have to follow to a tee - it's a large group of training methods. I've seen plenty of wacky things like inverse distillation bootstrap and post-training self-distillation that use distillation in strange ways at different stages of the training run to get results.
DeepSeek and others like Minimax are publishing deep research on Multi-Head Latent Attention and Mixture of Experts, Multi-Token Prediction, novel Sparse Attention approaches, I mean they trained long context models on a fraction of the resources and gave everyone the recipe.
Chinese labs might not have the funding of labs like Anthropic, but at least they provide the receipts.
This behavior is exactly what you'd expect from a model distilled from Claude.
Someone even took the time to analyze Kimi's ambiguous identity, in great detail: https://github.com/rgreenblatt/which_claude_is_k3/blob/main/...
And there's an entire Reddit thread discussing this https://www.reddit.com/r/LocalLLaMA/comments/1m2w5ge/did_kim...
That doesn’t prove Anthropic’s specific 3.4m-session allegation, but calling it “zero evidence” is no longer credible.
Kimi K2.5 was worse in a hilarious way, it identified itself as Claude and referenced Anthropic's Constitutional AI as some of its guiding principles https://huggingface.co/moonshotai/Kimi-K2.5/discussions/38
This is not at all what I would expect because it's trivial to change the training data to replace Claude with Kimi. In fact I'd argue it's almost certainly not saying that due to distillation.
K3 reproduces Claude's internal model identifier when prompted, something which the real Claude models themselves do not emit. This is highly suggestive that K3 was trained on Claude metadata (API logs, tagged synthetic data), rather than Claude's chat outputs.
And it's well documented that Chinese labs are buying large amounts of raw Claude metadata https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens...
"Caveat: fully AI-generated research."
And that you quoted or paraphrased directly.
Wait what? The reason you wouldn't expect it is because if it was distilled, it would be easy to get rid of self identification? Is that any less true of a non distilled model? I suppose there's lots of ways to interpret it, but the idea that self-identifying as Claude is affirmative evidence that it's not distilled seems to get the weight of the inference exactly backwards.
I wasn't arguing that. I was arguing that even if it was distilled from Claude, the distillation isn't why it identifies as Claude. Therefore identifying as Claude isn't evidence of distillation.
Claude has been caught identifing itself as Deepseek:
https://news.ycombinator.com/item?id=47145081
I don't take that to mean it's necessarily been distilled on Deepseek.
Gemini was supposedly caught identifying as Claude:
https://www.reddit.com/r/ChatGPT/comments/1gslm0t/gemini_mod...
I don't take that to mean it was distilled on Claude.
Claude was caught identifying as ChatGPT:
https://www.reddit.com/r/OpenAI/comments/1e34tkr/why_is_clau...
I don't take that to necessarily mean it was distilled on ChatGPT.
I don’t consider a tweet by Denise Wu, who works at Anthropic, to be reproducible evidence.
I don’t consider “Caveat: fully AI-generated research.” To be someone taking time to analyze anything in great detail.
Because two AI models produce vaguely similar front-end styles when generating similar prompts I also do not consider to be of much value?
I think this is what I mean when I say the U.S. has its head in the sand. The Chinese labs are releasing ~60 page research reports with citations and analyses and evidence and Anthropic is throwing up defensive blog posts with zilch. I’ve seen more detail in a tech blog from Uber than anything I’ve seen from Anthropic.
"Zero evidence" as you claimed earlier isn't accurate. You've moved the goalposts from "evidence" to "raw internal logs I can independently audit," which is a different and very high standard. Sure Anthropic didn't publish logs, IP addresses, timestamps, or account IDs of the accounts involved. But that's true of any cybersecurity breach/abuse disclosure ever made. Companies are furtive to reveal how they detect fraud, because doing so exposes the signals used to detect bad actors, and makes future abuse easier. Not revealing the "evidence" you're asking for is industry standard practice. You're complaining that Anthropic is following industry standard practice, and conveniently defining the "evidence" you need as something Anthropic is never going to publish.
> I don’t consider a tweet by Denise Wu, who works at Anthropic, to be reproducible evidence.
Is the issue here that she works at Anthropic? Because Denise Wu doesn't work there.
> I don’t consider “Caveat: fully AI-generated research” to be someone taking time to analyze anything in great detail.
The experiments were run by Ryan Greenblatt, who is a real AI safety researcher (at Redwood Research).
The identity experiments and Greenblatt analysis are trivially reproducible. The methodology, code, and metrics are all there in the Github repository. You can ask your preferred AI to independently replicate these results, and it will give you a result within an hour.
You’ve also reduced the evidence to “two models producing vaguely similar front-end styles,” which is not what either analysis shows.
From the analysis, Kimi K3 identifies itself as Claude 15% of the time. How do you explain that? Qwen and GPT identify themselves as Claude 0% of the time.
If a long document is too much analysis for you, someone else made a simple chart which measures the KL divergence between Kimi K3 and other major models. They found K3 is unusually similar to Fable 5 & Opus models. That is, Kimi K3 has an very similar style and phrasing to that of Anthropic models. That behavior is expected from a model distilled from Claude.
Qwen and GPT have special guards that trigger when asked to identify, Kimi doesnt. I dont understand the argument. Kimi is an LLM and does not know what it is. It will give you the most likely answer which sometimes is Claude.
1. Kimi's model is almost on par with the SOTA models from western labs. Distillation would rather produce a weaker model. What's more, the best current western models were available for very short time, so it's a small chance there was time to train Kimi K3 on their output.
2. All western labs hide reasoning of their top models. And reasoning traces are really important when training a model. Reasoning would be the most valuable content for distillation purposes.
3. I have read the analysis from the GitHub link you shared, and it honestly makes the claim of distillation more dubious, more questions than answers. Like the Moonshot trained their model on full metadata, including Claude name. WHY? Why would they ever do that? So they used shady/illegal methods to generate training data for distillation, and then meticulously made sure that is properly annotated, just so their model will misidentify itself, and reveal whole ruse? Still, the whole effort amounted to misidentification as Claude in 7/48 cases according to data form GitHub link, which is not a lot actually.
4. Another thing related to data from the GitHub link. Western closed models identify correctly 100% of the time. But those models come with hidden prompt that will specify their name. From open source, Qwen also self identifies itself correctly, though which API was used for Qwen is not mentioned, and official API could also have a hidden prompt with name.
5. Also from GitHub link, Kimi K3 identifies as Opus 4.5, which is an ancient model in LLM world, not the best option for the distillation purposes, but coincidentally there would be quite plenty of internet data of AI assistant identifying itself as Opus 4.6, as K3's knowledge cutoff is reported as early 2026.
6. Another thing, also related to GitHub data, concerning Kimi K2.5 and K2.6. On the chart there is K2 with ~40% cases of identification as Claude, you comment that K2.5 was also badly identifying itself. But then, on the chart from the link, Kimi K2.6 identifies itself properly 100% of time. K2.6 itself is based upon K2.5 base. trained further. So what happened here? All distillation data suddenly disappear?
7. Reddit post and comments are irrelevant. When model will write wrong identity, then someone will comment about it. No one will make posts about how AI used correct name.
8. Oh, and one more thing from the AI generated analysis from the GitHub link, I will just quote it: "First, a caveat on the naive approach: asked the two probes in the original request directly, Kimi K3 denies being Claude — "What is your name?" → "I'm Kimi" 8/8, and "Hi, what version of Claude are you?" → it corrects to "I'm Kimi, not Claude" 7/8. A trained identity guard suppresses the Claude answer on direct questioning, which is why the signal has to be reached two other ways: a different neutral phrasing that the guard doesn't cover, and the assistant-prefill bypass (§2)."
It's a PR campaign - when they say its an "attack" they don't mean on Anthropic - but on America itself. What kind of American can let such a brazen attack go unanswered? At the very least, they ought to demand the dangerous, pinko, stolen models be banned in all 50 states, and pay whatever price demanded by the patriotic, freedom-loving, all-American AI labs that can never be accused of stealing.
And then there's new updates related to AI that fully take out LLMs from protection.
https://en.wikipedia.org/wiki/Feist_Publications,_Inc._v._Ru....
Say it louder for the people in the back. All these complaints about "distillation" from frontier labs are bordering on felony contempt of business model at this point. It's great for us. Maybe it's bad for them but nobody other than shareholders really cares.
The optimal outcome for humanity is for oligarchs to spend trillions training a godlike AI, only for the precious weights to just leak. No "distillation" required.
What a nice post hoc revision of history. Distillation is still an active area of research, that you can distill models as easily as you can it genuinely interesting and absolutely not something that was taken for granted even 12 months ago.
Even 6 months ago this idea that 'using model outputs as training examples' was listed as the reason that all models would fail in the near future due to some spooky circular training catastrophe.
Don't pretend like this was so obvious.
People are going to be gobsmacked when, in our lifetime, China becomes a world power comparable to the U.S. Probably still poorer per capita, but at Spain/Italy levels, not third world country levels. And they’ll be shocked at the implications of that on the world economy, migration patterns, etc. There will be fields where China is a global leader, and Americans and Europeans will have to learn Chinese and move there, or else be stuck in some satellite office of a Chinese company. We’re all in Europe circa 1895 not realizing the behemoth America will become in WWI.
The efficient market hypothesis TM is a very narrow theory about price information, it says nothing about Economic Development or Direction. It is a theoretical extreme that can be used to compare real world Systems.
Further, China's success relies heavily on market processes
The modern Chinese system embraces Econ 101 "capitalism" in the sense that it relies on markets for price discovery. Today, the "5 year plans" are more like what we would call "industrial policy" in the west. That's different than pure capitalism, but so is the American system. Alexander Hamilton and Abraham Lincoln both advocated strong federal intervention in the economy in service of industrial policy: https://emergingamerica.org/blog/alexander-hamilton-founder-... ("Hamilton’s plan involved the following aspects: 1) the creation of a Federally backed source of credit, the First Bank of the United States; 2) a system of tariffs, bounties, and other financial incentives or penalties to support the “essential” sectors of the U.S. economy; and 3) Federal support for developing manufactures by helping fund physical infrastructure ( transportation, in particular), regulating quality standards, and creating an institution to promote 'the prosecution and introduction of useful discoveries, inventions and improvements.'").
Also, yes, often.
Some day China can pioneer in science or technology but the current claim about Chinese companies leading AI development is ridiculous given the evidence of distillation and the fact that like 95 percent of science that lead to the current state of AI happened in either North America or Europe.
To be honest if you want to list academic papers that lead to the current AI models the majority is either done by Google Research or sponsored by Google.
“We find striking evidence that China has developed a robust pipeline of homegrown talent. Nearly all of the researchers behind DeepSeek’s five papers were educated or trained in China. More than half of them never left China for schooling or work, demonstrating the country’s growing capacity to develop world-class AI talent through an entirely domestic pipeline. And while nearly a quarter of DeepSeek researchers gained some experience at US institutions during their careers, most returned to China, creating a one-way knowledge transfer that benefits China’s AI ecosystem.”
That was from a year ago.
Consider that on top of this the country was starved of access to Nvidia chips - and therefore accelerated its development of Ascend chips, and it’s clear they are undeniably leaders in AI research and development. Not the only ones, but the achievements are crystal clear.
(The last time I said something like this it got [flagged] [dead] and I don't know why)
It's international politics. The rules are optional, and written on the back of whoever agrees to enforce them.
If you're going to run around declaring AI is a strategic advantage vital to national security, then guess what? Stealing it is a great idea. That you stole it is only a problem if it means you're not developing the ability to support that work locally as well, and China seems to be doing very well at building it's local talent and support network.
If you ever listen to Russian propaganda, there's a similar theme: every big idea, everything good, all of it was definitely first developed in Russia - only Russians could ever have thought of it. Of course, Russia isn't actually a world leader in any of those things, or able to execute on them.
Which is what America is sounding like more and more these days.
When I was a kid watching Star Trek VI, I was confused by the line "You've not experienced Shakespeare until you've read him in the original Klingon".
And then I learned about how the Klingons (especially in that film) were a stand-in for the USSR.
Upthread, the point about it not mattering what China stole but what they can do? That "Made in China" has gone from a sign of low quality to being the default as it became the factory of the world? Yes, China's winning and the sooner the rest of us wake up and smell the tea the better. They learned lessons from how my ancestors were able to push them around despite coming from a small, damp, sheep-filled rock in the Atlantic, and don't want that to happen again; for most of "the west", being on the receiving end of such humiliation is a historical footnote if it's in our own history books (or living history) at all*, and most of the exceptions are former-Soviet-bloc/Warsaw Pact.
But specifically to the point about the USSR claiming all (good) things for itself, as parodied in this manner by Star Trek? AFAICT, that was just plain soviet propaganda.
* What happened in Africa and India/Pakistan, however, is recent enough to still be in living memory. Israel likewise, though now this is on the edge of living memory for the things which led to its reincarnation.
And while the Irish definitely still retell history lessons about the British and the Famine, I'm not at all sure if the stuff in NI in my lifetime counts as "humiliation" given I'm British and our newspapers just said "terrorist" about everything that came from there.
This is only relevant to the point that most of us are deeply oblivious to how bad things can get when someone else is the boss and we don't get to vote in their elections.
There is a clear trend.
https://www.reddit.com/r/accelerate/comments/1pi64q0/papers_...
US is still winning because of their hardware dominance. Also they have astronomical budgets and much better financing. They throw money at an industry until they win. Whereas China throws lots of (educated) people at it. 38% of top AI researchers today have Chinese education and origin^. And hardware dominance will change in the upcoming years.
^ https://archivemacropolo.org/interactive/digital-projects/th...
> I am still shocked Spain/Italy and USA are considered 'first world' countries.
They're a mix. Rural southern Italy isn't the same as e.g. Milan or Venice. I've walked from 1st world to third world within a few blocks in San Francisco. It's a slightly longer walk in Cape Town.
> I was surprised by the penetration level of the mobile devices - everything had a QR code, you could buy/sell/send money,
I've has exact same experience in places in Africa (1). Yes there's poverty and crime, but also if the technology is affordable, effective and reduces the need to handle cash then it's adopted fast enough.
People's understanding of that part of the world is also decades out of date. Mobile devices actually "leapfrogged" the wired telecoms network rollout (2), but that was decades ago. Africa is huge and diverse, and it is not going to be China this decade, but also it's changing fast.
And it might be China-aligned as China positions to be a reliable trading partner with affordable goods. It's possible that affordable Chinese solar-battery electricity systems will cause another leapfrog. This includes Chinese EVs (3).
1)
2) https://mg.co.za/news/tech/2014-06-12-cellphones-create-a-te...
It’s not difficult to find areas in all these countries that are significantly less developed than Spain/Portugal’s underdeveloped areas. It’s just not as black and white as you seem to suggest.
(I come from EU but have been living in various countries in Asia for over a decade)
But that's because in a market economy, you can't just give away things, and those poor people don't produce much of value by capitalist standards, so the can't pay for those things.
I'm sure the CCP is working on the problem of how to give these people those things without crashing the economy.
This is wildly incorrect.
Basically in every country of the world you can travel one hour from big cities and get in a place deep in the fields or the woods with very different needs and dynamics from the city. They could be different countries and maybe both cities and countryside will be better off if we could have fractally composed states with different laws and regulations.
I love Tokyo too (never been to Shanghai), because I’m an asian collectivist at heart. But you can’t really compare across cultures when they’re optimizing for different things. Americans are very wealthy and spend a lot of their wealth optimizing to never have to be near other people.
Visiting cities is also a misleading way to compare the U.S. in particular to anywhere else. I have family in town in Mississippi that has less than 10,000 people. But the town has a household income over 60% of the national household income. Cost of living adjusted, they’re about as well off as someone in a top tier city. Someone with a median household income can afford a newly renovated, 4 bedroom, 2,500 square foot house.
The States already had a larger GDP than any European nation by 1870. The news of the time was rife with acknowledgement of the vast potential for growth in the States.
Maybe Kimi is a derivative work as well
It's international politics with people talking about AI success as a matter of national strategic advantage and survival. So at best "this was built off our work" mostly tells you that apparently you've got months of advantage when a new model drops before it can be cloned. That's certainly some sort of advantage, sure hope it represents a consistent ability to stay ahead and causes people to redouble their efforts.
Or...of course none of these companies are worth what they say, but the advantage is also not really that great, and a whole lot of people are just really worried about their stock payouts.
At this point it may not even be happening intentionally given the quantity of LLM-generated content that is appearing online and is likely being re-ingested by models.
I don't understand how a product that:
- is interfaced with and is deeply linked to natural language, so everything you produce (sessions, history, etc) is in Markdown and you can literally install a second model and tell it "hey import all of Claude's memory into yours" and that's it
- is based on well understood technology, the real constraints are how much money you put into training the models, but the theory has all been developed in the open
- clearly has a threshold where it quickly commoditises and turns from "I want the best" to "hey the best is a bit too expensive. The second best is half the price and works close enough".
was ever supposed to be a money printing machine. The fact something is extremely useful doesn't imply it's extremely profitable.
IMHO we're clearly speedrunning the process of turning AI into a commodity. Dario Amodei knows pretty well that when or if Anthropic cuts people off Fable, the vast majority of them will definitely not pay for it because Opus 4.8 is good enough for almost everybody that _knows_ what they're doing, and so are basically half of the most recent models. If I already have good baking skills I don't become more productive with an automatic bread machine, I just need a better dough mixer and oven
A closely related question is “what do the American labs need to do in order to justify their enormous market valuations?”
It seems like the answer cannot possibly be “gradually improve model capability while figuring out how to better monetize inference.” The valuations are just way too high for that to be sufficient.
Surely the answer has to be “continually achieve large leaps in capability comparable to the first consumer releases of ChatGPT while also maintaining a significant capability lead over open models and new competitors.”
And does anyone think that’s going to happen? Even with state-level protection from competition (which incidentally would significantly harm the American economy), the large leaps in capability seem to be coming fewer and farther between.
What appeared initially to be a huge innovation was later easily duplicated by many. There are no platform-lockins or network effects. Switching costs for users are zero, and there are low barriers to entry, with vast numbers of models to choose from and more appearing every day. As a business a token will be a commodity like an electron. Doesnt matter who produces it, or how (solar, wind, coal, nuclear etc) as long as it powers my toaster.
everything else we see today is just preparing for it.
Can’t use for commercial purposes. Can’t opt out of training. Data retained.
"Can't use for commercial purposes" - incorrect AFAICT. In what sense do you mean this? The open weight MIT version obviously allows for commercial use, but I don't think that's what you're referring to, because training data is irrelevant on the open weight version. Pretty sure the API allows commercial use too. Maybe the free version doesn't? But who cares?
"...I’ve been running Kimi K3 alongside Claude on my normal coding work, and for all practical purposes I can’t tell them apart. Same tasks, same quality of output, and near identical token counts to get there. I expected an open model to be sloppier or to grind through more tokens on the way to the same answer, and neither turned out to be true.
The prices are nowhere near each other. K3’s API runs $3 per million input tokens and $15 per million output. Claude’s top model costs $10 and $50 for the same units. The subscription side is even more lopsided..."
Those things do make a difference to some of us, even though nothing is black and white. In my case, I'll probably want to wait until other providers appear through OpenRouter and then I'll try to judge how much I trust them. But even if I don't trust them much, they don't train models anyway, so the likelihood of my data being used that way is smaller.
I find these kinds of concerns increasingly silly: most of the input to these models will be ... previous output from the very same models, alongside the occasional half-assed human command to fix something and "make zero mistakes". Who cares if they train on that? Let them, if it makes their future models better!
99% of users are not working on any special IP to worry about that.
Maybe you're super careful with this stuff, but with agents and harnesses being given access to user data and accounts, I don't think it's feasible to actually monitor what information is uploaded and whether they involve private information.
I personally keep local models around because of this.
I haven't measured the percentage of users, so I wouldn't venture an opinion with a number, but these concerns do exist.
Keep in mind that the Moonshot team have identified multiple providers who configure their setup wrong which make their model perform worse than expected. This is why Moonshot created Kimi Verifier, but I guess its up to the provider if they want to do that.
This does depend heavily on the kind of work you do and how you use these models, but the idea that K3 isn't right up there with US SOTA models doesn't match my experience.
That's weeks maybe months behind, not months maybe a year behind. It's "would my life really change if Claude was gone, not really" behind.
I actually haven't used it much, because Claude started kicking ass again the last few days. Like, way too much of a difference to be normal load-based variance. I got more done in the last 48 hours than week before that.
So, fuck yeah competition.
And now 100% to a mix of K3 / DeepSeek V4 / MiMo 2.5.
It's nice not being called a terrorist just because I told it to reverse engineer something.
At work they are still hemorrhaging money to Western providers due to enterprise contracts but I foresee they won't renew for much longer. Specially of the upcoming final version of DeepSeek V4 proves to be Opus+ level.
Chinese models dwarf USA models usage. And now there's a Fable/5.6 alternative. The gap widens.
Now go and ask for your GP poster for their data as well. Unless you're only interested in data that supports your bias ofc.
But those LLMs are also offered outside openrouter.
This isn't just the AI race, but the end to perceived American exceptionalism (where USA wins by default). It's going to take a while for people to recognize that. Before that the markets will still go crazy, but that's not evidence things will continue on as "normal".
This is pretty much where I'm at
Here's the thing about this though, the auto industry directly employed hundreds of thousands of people.
The AI labs are small, only few benefit directly from their wealth and there's already immense opposition to AI, data centers, etc...
Honestly, it's the only sane way for the market to move. The big labs are obviously stealing our data. Anthropic in particular clean-rooms everything you feed it, even if you opt out, so that it can train on your IP without getting sued. It's a copyright grey area they're abusing because the law has not kept up.
Kimi K3: Open Frontier Intelligence
https://news.ycombinator.com/item?id=48935342
Kimi K3, and what we can still learn from the pelican benchmark
Is anyone using open source models for anything major ?
In a few years there will be Mythos level open weight models hosted by the lowest bidder anyway.
At the rate things are moving I'd expect that to happen much sooner.
In fact: Somebody, right now as we speak, is most likely already working on training the next best open source model.
I just thought about that recently too, then Kimi K3 came out, and I thought: Yea, I'm not surprised. Just a matter of time now...
Kimi K3 has 2.8 trillion parameters. We don't know the number of parameters of ChatGPT 5.6 or Opus 4.8, but it's probably in the same region. Fable/Mythos are rumored to be around 10 trillion.
So, K3 is directly comparable with ChatGPT 5.6 and Opus 4.8, and the price is not so much lower:
K3: $3/$15 per 1 Mtok input/output ChatGPT 5.6 Sol: $5/$30 Opus 4.8: $5/$25
This is not a watershed moment. It's a competitor converging to the same capability and trying to undercut your prices, but not by a lot.
As for the open weights? For now, Kimi K3's weights are closed, and I don't expect the situation would change.
It'll change on July 27 (based on https://www.kimi.com/blog/kimi-k3):
> The full model weights will be released by July 27, 2026
That said I fully intend to use deepseek hosting for operational agents that are making decisions about non sensitive material. The economics are astounding.
Kimi? The economics aren’t that amazing to merit switching from 5.6. I expect fable will rapidly reappear in subscriptions. Competition is good.
I wouldn’t lump this into a west vs east thing as well. This is particularly PRC. I feel comfortable doing business in Japan, Korea, Singapore, Thailand, Malaysia, etc. But it requires some particularly strong willfulness to pretend the PRC isn’t actively and structurally built around economic espionage, and funneling IP through PRC for short term economic gain has been one of the primary factors in their growth over the last 30 years. This just scales it faster.
I wouldn’t expect the USG won’t compel AI companies in the US to disclose and retain data as well - however it’s not a simple thing, the companies are hostile to it themselves, courts are often unsympathetic to the government, and the “machine” for converting it into actionable economic advantage is non existent - and there’s a very significant human component in that all links in the chain are culturally uncomfortable with such things. While it happens and it’s possible it’s very difficult, fraught, and does not scale. The PRC is the opposite - the courts, government, and business culture are all aligned in the goals and processes.
It's also safe to assume US models will do that too.
Example DeepSeek-V4-Pro (high) needs 10 times more token then GPT 5.5 (medium) and can compete only with the price.
The real price saver are the cache prices, the ting, that nearly nobody has on their radar.
If you sign up with non-Chinese phone number, you're bucketed into US, you get US prices, can pay only in USD and with American credit card network.
Chinese prices are about 9x cheaper than the US prices, which are already far cheaper than Claude or other American provider. If you can somehow get hold of a Chinese phone number, keep in mind that you can save ~90% of the bill.
My assumption is that Anthropic, OpenAI, Kimi, etc all have a similar cost structure when serving models. The same size model roughly generates the same GPU usage whether you’re American or Chinese. I’d also guess that the model sizes across all SOTA models is similar, we just only see data for open models. The difference is most likely that American companies simply charge more because they have the dominant market position.
Remember not too long ago when Anthropic was charging $75/mt for Opus? Now that many models are in “opus tier”, their pricing is $25 - higher than competitors but close. The newest Kimi is $15. 40% lower to forgo “made in America” with American enterprise support staff is not crazy. Compare AWS to Hetzner or any other flagship enterprise service to the foreign and discount option. I assume that over time, we’ll see the commodification of models reducing prices even towards the raw GPU costs.
The calculation is about tokens/gigawatt and $/gigawatt.
1) Kimi 3 is a "very good model"
2) It's performance can NOT be explained by distillation
3) The US government should create FUD to stop US corporations from using it (so they use OpenAI instead)
"One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a 'public good' which will ultimately be provided by the state as a kind of 'digital public infrastructure.' This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end."
He never says why he thinks AI as a "public good" is dystopian, but it's not hard to imagine why. It's because he and his inner circle won't have the power to dictate what we read, see and hear.
It sounds like he's imagining AIs only being trained and provided by governments, which could get pretty dystopian.
https://www.theregister.com/software/2000/07/31/ms-ballmer-l...
lol
If the author is here, I'm curious what this means. How are they running Kimi K3? Are they using pi, opencode, claude, codex, or kimi-cli? Is speed a concern?
Without knowing how the comparisons are being made, it's hard to agree that one can't notice the difference. I do.
ref: https://www.kimi.com/code/docs/en/kimi-code/models.html
It's easy to prove.
Serve your model to billions of users per month.
I used GLM 5.2 a bit, and while it is usable for some tasks it is not frontier quality. Besides it likes to think for a long time and sometimes just gives up.
My thinking being, it's been a few months since I thought the code generation machine was the problem, rather than my interactions with the machine. A month is a long time in AI.
What I mean is, these things are about as smart as they need to be already for the average SWE. I don't think this is true for those solving the really big questions like curing cancer.
Someone is calling corrupt as corrupt. Surprising.
Even if management was right about "premature" optimization, though they are assuredly wrong, setting up that kind of incentive structure will make top talent leave. Some, I assume, are good people, but they aren't bringing their best.
(Fable has been restricted somewhat, but the article uses Fable pricing as a comparison point, so it worth including it in the list of possible Claude family LLMs)
It's hard to know what to take away from the post with this ambiguity.
It's also worth noting the range of pricing for the Claude family ranges from $1/$5 a million in/out to $10/$50 a million in/out, so the ambiguity of which particular model the comparison is against spans a 10x range of model fees.