Edit: miss me with the downvotes. These guys are clowns. I want alternatives. Thank you to everyone offering help!
OpenRouter isn’t a provider, they route to other providers, so you would need to specify which ones you’re comfortable with anyway.
[1]:https://docs.fireworks.ai/guides/security_compliance/data_ha...
For OpenRouter, you can setup an API key and limit it to only models that claim to not train on data, but that is just a claim. You can then use trust to judge which providers will honor that claim.
But if the data is really sensitive, you might want either a local model or a business subscription with some big name in the US that legally promises no data training.
So are we talking some app idea you are playing around with, or files filled with PHI/PII that you have legal mandates to safeguard? If the latter, I would stick to only provider with enterprise agreements to not store/train on the data. Even the ones who promise no training are likely storing the data for monitoring for abuse or such short term.
If it’s really sensitive then don’t use a cloud provider.
I'm very happy with Deepinfra. Less model coverage, but good prices and quite fast.
However I have to caution you about one thing.
No one will give you as many input tokens for so little money as Claude Max x5 (maybe x20 too, I use x5).
I tend to use 1.1B to 1.4B a week about 0.8-1B cached. Even with cache were talking thousands of $ in API prices a week. Hundreds if we're talking cheap cloud like Deepinfra.
However, local AI well setup is actually a good alternative for this if Claude Max was unavailable.
For example my system a ryzen 7950x 192GB ram, 5x rtx3090 plus an rtx5060 ti 16gb. (3 rtx3090 cards via usb4 egpu dock). Let's me run Qwen3.8-Flash-Next with 3slots (no rtx5060 used) at 55tok/s decode dropping to 50 at the end of a 260k context, 1200tok/s refill dropping to 950 at the end of context.
With RAM and ssd cashing and 80% cache were talking on the order of 4B a week could be ingested by this setup (roughly) if it was running 24/7. I found 6 interactive cloud code sessions are fairly pleasant with this 3 user setup.
If I include the rtx5060 in the mix I can bump to 5 users, but it slows down by about 15% (note the speeds are give are for one active user, multiple users at once see maybe 70% of tgat per user so aggregate is much higher in multi user setup).
So in theory I should be able to run 10 cloud code sessions. Although I'm testing CC alternative now (pi with own plugins) because this model, while multimodal has only 260k context 30k of which CC eats on the getgo.
Many people say local AI makes no sense financially. But in the event you process huge inputs that are often cached it does make sense.
I think your local setup would be a bit much for me capability / price wise. But maybe I can find a scaled down version. I don’t need insane tok/s. Oftentimes I just let things run and come back later.
They are near a third of the cost of an rtx3090 while the performance is much better than a third.
They cost $550 day before yesterday on Aliexpress here in EU.
The problem is there aren't any cheaper gpu-less setups that could run this at let's say around half of my speed including prefill. While some people reported 20t/s on a strix I never saw a prefill number. I think it would be pretty bad (like 150-200tok/s). And these 20t/s are probably single user only at small context. So not worth the money for me.
A Ram based alternative is a threadripper system with 8 ram channels, because it can have memory bandwidth comparable to cheaper gpus. But you need to use registered RAM and that is bonkers prices now.
So personally I think sticking to a desktop pc MB, 2 or 3 gpus inside, plus 3 via usb4 is probably optimal. Using usb4 leaves your nvme slots for nvme. I'd consider 96GB RAM minimum comfortable (to keep kv cache of 10-12 claude code tabs you may work in).
Don't forget the cost of the eGPU docks and psus. It's not trivial when you have 3-4 of them. I paid around $250 each.
I only have this system because I was lucky to buy 80% of it when prices were better.
If I was buying today I'd be hard pressed to justify even the 192gb of ddr5 (normal, not registered).
If you or anyone else does this mind you'll spend a couple days getting resizable BAR working reliably.
There is one more option. Tesla v100 cards. 16gb and 32gb. I would disregard 16gb cards immediately. Why? Pipeline paralellism allows you to split a model between cards for almost "free" (latency), but the layers are usually few GB big and you can never allocate it to consume all vram. You always have 0.5-2gb unused per card. 2gb is a lot for a 16gb card.
So 32gb v100 sounds good right? Maybe... But if I was going towards v100 I'd not buy pcie version but the datacenter grade (I forgot the interconnect name). There are big adapter pcbs on Aliexpress that take 4 of those v100s and they allow you to connect all to single pcie, but the 3 v100s are all nvlinked.
The pcb costs in the region of $500-600.
But it doesn't make any nvlink exit the board. If they did... I'd be buying two such systems. Linking 8 32gb v100s together and with fast interconnect you can run tensor paralellism which uses compute of all those cards at once. It would prebeat my system 4x at least.
Consider Qwen3.8-27B 4bit 8bit kv, q4 (if I remember correctly) run at 30-40 tok/s on a single rtx3090. Two cards in tensor paralellism and nvlink run it at over double at 90t/s and prefill, was amazing too.
OpenRouter allows to make an API key locked to certain provider, I do that myself.
This is crazy but honestly very tempting at this point.
You have to be the embodiment of hubris to believe the only reason a technologically advanced country of 1.4 billion people with several AI labs and government support can make competitive models is because they all distill yours.
https://www.heise.de/en/news/DeepSeek-orders-160-000-Huawei-...
Like I said, it will delay the Chinese labs for a couple years. These are not even top-line chips.
Frankly, China was not going to allow them to depend on NVIDIA forever, I don't think this motivates domestic production that much over the counterfactual. If NVIDIA didn't have export controls, China was going to set up formal import controls. They want to own their entire supply chain.
Not even that long. China can simply order its industries to stop supplying externally and go full-domestic. It has happened before in others parts of Chinese industry; it will happen again. As it stands China can easily build computational nodes at scale, and their LineShine supercomputer holds top place in the Supercomputer Top 500 at almost 2.2 exaflops. They have zero issues building performant hardware.
Couple of years? Their current technology level can make that less than a month.
Also as we have seen, the Chinese labs will just buy their kit through cut outs or do training in locations where there's no embargo.
They aren't patriotic and they don't really care about competing with China, they fancy themselves princes and are playing on the silents/boomers in charge being irreparably stuck in the Cold War.
I don't agree with all of their means for getting here, but to accomplish what they have in ~30 or so years [1] should blow minds far more than it does in the West. And they did it by convincing us to give them control of our manufacturing and selling us (literal) boatloads of cheap junk. Historians will look back on this era as one of the greatest demonstrations of "winning without firing a single shot."
This whole insular POV is tired, ignorant, and frankly just another indicator that America has lost the plot, shit-faced on its own arrogance.
I'm confused, what has America and the west lost the plot on? You would have preferred the west to exploit China more?
This framing of a zero sum game is counterproductive.
I agree but that's not at all what I was getting at. Quite the opposite. The mistake America made, what made it vulnerable, was playing everything like a zero sum game. We're in the early days of that behavior coming to roost.
> China played the smart long game and took advantage of America's...
Zero sum games imply that if there is a winner, there has to be a loser. That one country takes advantage of another.
“If you assume everyone only has hammers, you'll end up getting screwed”, or something like that.
I didn't bother responding earlier, and I'm not sure why I'm responding now because I don't think "because they framed things as a zero-sum game." explains anything, and it sounds like something that sounded like a retort but offered without any evidence means nothing.
If even the suggestion of considering different points of view feels like an attempt to disprove your own position, then that mainly says something about how you are approaching this exchange.
Hmm, sounds awfully sensationalist, designed to fan the flames of Sinophobia and thinly veiled racism.
Of course china isnt innocent in all of this. Their 996 work culture is inhuman, and basically indentured service. But, their government is able to leverage that workforce.
And why was it cheaper? Shipping things from the literal other side of the world isn't free. Cultural and language barriers to efficient communication aren't free. Etc. etc.
China's protectionism is emplaced to support long-term goals. Amodei's policy suggestion is explicitly in support of short term goals (1-5 years), a time window he claims to be "critical" without evidence. A cynic may link that window to forthcoming AI IPO's, and allowing current CEOs to secure ridiculous golden parachutes before leaving the mess of how to deal with the unpaced Chinese frontier models to their successors and future US legislature.
Letting AI companies set your foreign policy (voluntarily ceding global primacy to maximize domestic[1] profits) may have awful higher-order effects for the US. What's hilarious is the shameless about-face: for years, the same companies have been declaring that it is vital American AI innovation not be encumbered by new laws. Until the Chinese models glided to the frontier.
1. Maybe they hope to browbeat Europe too, like how they pushed out Huawei equipment from Euro telecoms. An alternative reading is that the loss of primacy is involuntary, so they may as well spin the reality, and salvage a win by colluding to eliminate better, open-weight models from the markets in "democracies" with the excuse that they are unpaced and therefore dangerous. What will predictibly happen is there will be `Qwen-8.1` for RoW, and a lobotomized `Qwen-8.1-paced` for Americans.
There’s a reason China has a GDP per capita somewhere around that of Argentina.
Of course, they also have futuristic cities, maglev, and plenty of brilliant engineers and scientists, therefore it makes sense to be pragmatic. But there are a lot people talking up China because they hate the US and are bought into a massive propaganda / influencer campaign.
Having visited the US a few times it felt that way too
(a) Sell many billions of dollars of fancy chips to China, thus bringing in billions of dollars of money and billions of dollars of trade balance improvement. And let China continue to build models quickly when they seem oddly happy to export the trained weights for free.
Or
(b) Decline to do so, thus reducing exports and very very strongly encouraging China to do everything in their power to avoid needing to depend on our chips.
In exchange for (b), what do we gain? A temporary advantage in model training and availability?
If you tossed the aviation regulatory framework out the window, you have maybe ten years before aviation reverts to high end hobbyists[1], due to the escalating risk profile.
[1] or proletarian desperation, of course.
It’s so obvious they’re playing safety card to get politicians to reign down on individual’s freedoms.
All at once in lockstep?
Who paid them to say this? These are all lies.
Chilling and suspicious. Of course, this will lead to a modicum of safety in some provable way to justify (in uncritical minds) the rest of the implications: fewer freedoms and more surveillance for everyone else.
"Anthropic Is Building a Huge Surveillance System to Spy on Anti-AI Activists and Predict Their Activities"
https://futurism.com/artificial-intelligence/anthropic-surve...