OSS struggles at being relevant when software is non-commodity e.g. office suites. In software domains like databases where the state-of-the-art computer science research is often unpublished, OSS struggles to be relevant at the higher end of the market on technical merits.
When deciding what should be OSS, it is useful to consider the preconditions that have made it successful.
See open weights gaining adoption, OpenAi talking about how 5.6 is cheaper than Fable, people are taking multiple approaches to reduce their token spend, expectations for progress in hardware and algos, and certain Ai leaders talking about how token prices should be 10-100x lower than they are.
Corporations have had many reasons to invest their money in open source software -- custom requirements, marketing / developer mindshare, commoditizing complements -- but as cutting edge LLMs get more and more expensive to train, you'd be hard-pressed to find corporations who will put in that kind of money if they cannot recoup their investments.
If you can not run the training yourself you can not contribute. So open source contribution model does not work. All examples you gave have a fairly low threshold of capital expenditure required to be a contributor (basically a laptop).
Even back in the 90s a person could get a standard, but powerful, PC to do these things. The one exception was 3d graphics which took quite some time to become affordable and even there it was a single one-time expenditure (a workstation) per contributor.
In an open-source LLM model contributors would compete with each other for computing resources for model tweaks and changes. The alternative model is that the contributor pays for the compute, but that increases the bar really high for contributions.
on the small chance that the four billionaires who currently have near-exclusive control of closed sota models, (that is altman, amodei, zuckerberg and musk), are not fleecing their investors and actually build AGI, closed source leaves a choice of powerful government or powerful oligopoly/monarchy.
further explanation of this list:
musk - structural command
zuckerberg - structural command
altman - de facto command after purging rivals and privatisation, loyalty of personnel
amodei - influential, could potentially overthrow current governance
Every 6-12 months, give out $200K to the first model to hit a min threshold on a set of ~5-10 hard benchmarks (+ perhaps one secret benchmark) using a total of 16GB / 32GB / 64GB / 128GB of VRAM (at a min context length of 200K), then move the threshold up. Quantization etc. is dealers choice, it just needs to nail the benchmark on a reference machine by using exactly that much VRAM (no mapping to RAM / disk etc.)
You could crowdsource the funding, and cross subsidize by adding targeted prizes focused on corporate needs (the classic one is PDF processing benchmarks), and say that 25% of each corporate prize funding also flows into the general prize pool.
For a lot of these open-source model companies, it's less about the $s (though $200K is nothing to sneeze at), it's the clear recognition that helps their model efforts stand out, gain usage etc.
Having it with clear hw requirements tiers is a nice differentiator. The only issue is that the benchmarks would 100% need to be closed, no other way around it. And then you have the issue of creating and curating good evals for every "stage" of the project. That's a hard task even for "honest" lab-internal evals. And you'd have to publish those evals after each round (for trust purposes), and start over for the next round. Doable, but it would cost a lot (probably more than the prize pools) and you'd have to keep doing this.
Yes, I hope the open model communities will someday be able to run current frontier models which will be able to handle autonomous tasks and the hardware to run it will be served at consumer level; however, like how recycling isn't profitable, no companies will fully commit to the movement. Don't get it twisted, I don't have a solution but maybe a global scale movement to liberate knowledge-library could suffice.
The library analogy in the scenario would hold true if LLM providers refused to answer any questions about RL or Transformers.
I am a big proponent of open-source open-weight models, but mostly because I think it's just a better product. We've seen that they are much cheaper to train and operate. Frontier intelligence might not be needed for most tasks. Just let the market decide. My bet is that LLMs will become analogous to programming languages, and big labs will make their money by fine-tuning models for very specific use cases or by deploying them for customers.
Op-ed alt link: https://fortune.com/2026/07/03/open-source-ai-same-fight-as-...
this is like saying "gov should invest in pyramid schem, because everyone is doing it". or btc. or web3 pictures of monkeys.
what i expect the gov to do is to add a 999% tax or tarif on top of GPUs bougth for AI, after the first 100mi that company spends on it each year.
> Even if the current approaches will continue to scale, this would be as if in the early days of computing, perhaps someone invented a bubble sort for sorting numbers (an n-squared algorithm), and the tech companies at the time decided they were going to build vast data centers to sort numbers and not bother to figure out that there's an n-log-n way of doing it <laughs>
...to which I have to say: yes, definitely! And he's right about open-source AI too.
Ask Amodei how he feels about going to spaceman bad for compute that he couldn't find anywhere else in the market.
To borrow a real example from a prior boom, railroads were congested during the initial build-out while people simultaneously funded and built too many railroads for future demand. Likewise, Anthropic et al can be compute-starved now while the industry as a whole is overbuilding expensive, depreciating infrastructure.
Yeah, wooho, new model found a bunch of bugs, now the bad guys can do it too so security expenses spiked! It's only good for shovel sellers.
I fully agree with this article - please let's skip the chapter of closed and enshittified AI and go for the good stuff directly!
Of course we do have basically open source research programs, including most universities and big projects like CERN. But AI grew up in universities until it transpired that sufficient capital could only be found in the private sector.
It would be possible to make a decent publicly funded AI research program. But it would look more like the Manhattan or Apollo projects (which frontier labs already model themselves after) than some extra research grants for universities.
The entire Apollo project at the peak of the cold war cost about $300 billion in today's dollars. That's approximately what OpenAI and Anthropic have raised together in total until now.
I don't think governments can supply this amount of money for AI in the current political and economic climate. The LHC cost less than $10 billion by comparison and it was spread out over a much longer timeframe.
I'm a believer in Keynes' "anything we can do we can afford". It could be afforded .. if there was a sufficiently good reason. And there isn't. This is way behind "governments, especially the EU, should have a sovereign cloud". It is also way behind "governments need to keep global warming below 2C by the end of the century" and "governments need to ensure affordable energy", objectives which the current AI buildout is in direct conflict with.
This is before we get into the question of whether AI has net positive social value in non-software use cases. Even in software the case for AI is explicitly job-destroying and raising electricity prices for everyone else.
It doesn't matter if a public effort is a year or two behind the curve ( and hence has dramatically lower costs due to Moores law and the ability to piggy back on research ) - especially as we approach the asymptotic phase of development.
History is full of these take overs if there is risk(usually happens after some catastrophe). See the finance sector(once upon a time private banks invented and printed money), nuclear industry, febrtilizer industry, crypto, a whole bunch of processes in biotech/synthbio. Classic textbook example is the East India Company. It was much richer that the British Govt or the King.
You understand how the system works if you’re thinking in terms of government/non-government. The current political and economic state is not a bug, it’s by design which serves a purpose.
Remember, the purpose of a system is what it does
But…the USSR and that entire model failed spectacularly? So not sure what you’re getting at here. Is there some fantasy economic model you believe you’ve innovated that will lead to utopia and the end of resource scarcity?
One way of looking at it. Another is that AI research progressed within universities, but it was only until recently that the private sector saw the profit potential when combined with modern CPU/GPU technology.
You could argue we're saying the same thing, but I think these angles are different. An academic research programme would not have spent billions on a datacentre to provide AI for free to the general public, for example.
_LEAN_ FOSS, including the SDK then the computer languages too.
All computer languages with an ultra-complex syntax are excluded de facto.
Then there is the stability in time.
developer/vendor lock-in on software, planned obsolescence, are much more common in FOSS nowdays.
But wouldn't this help china build models just as good as ours immediately? Wouldn't it make the investment in training a model worth a lot less?
Why? It's not like they did the hard work. It's disgraceful that this kind of commons enclosure has been allowed in the first place.
> The law locks up the man or woman / Who steals the goose from off the common / But leaves the greater villain loose / Who steals the common from the goose.
RL/post-training is now a much bigger part of it, with that being largely proprietary and expensive work that the open source model just won’t fund.
Agreed. Especially since now competitors have more difficulties getting the same advantage. They don't have to do so immediately and perhaps not their specific tuning. But the weights of the raw training data at least should be publicised.
What's the precedent?
All I see if a bunch of people who haven't loaded an ad in 20 years and have a 5TB collection of pirated movies and music suddenly decrying that LLM's doing next token prediction over a dataset is theft.
You don't get to change your mind 20 years into "I'm never going to pay for anything binary" ethos (cough look at what this post is cough) that has dominated the internet for decades. If people are genuinely upset about LLMs training on all available data without compensation, all I can say is "Reap what you sow".
Like all those unethical software devs taking salaries for turning stackoverflow into products? I suppose the blood still flows since now they just use LLM output?
Any way you try and slice LLM morality, you end up with "It's bad because they are not me" reasons. "When I monetize information I get for free, it's good, when they monetize information they get for free, it's bad"
"Reap what you sow"
It is kind of incredible that you're not focused at all on the copyright holders, instead focusing on random tech people you had online disagreements with.Artists and writers got screwed first by piracy, then by generative AI. They didn't sow anything. They just got reaped.
And the only thing the copyright hypocrites are "reaping" is a feeling of hypocrisy. Congrats for pointing that out. Your comment is simply myopic.
That's too glib. Artists don't have a right to money that people won't spend.
It's not all down to bad actors. Although it's true that musicians got screwed when their distributors switched to a subscription model. Writers got screwed by the Amazon monopsony that crammed down publisher's margins.
Mostly what happened is that media technologies changed, professional creators had less control over production, and the money flowed upstream leaving them with a smaller niche.
All of their output used to be gatekept by media companies who had a lock on publishing. The internet reduced distribution costs to zero. Suddenly anybody could reach an audience.
Now generative AI has dropped the cost of content creation to the basement. It's so much easier to write a blog post or make an illustration.
That isn't cannibalizing content. It's a new way of making. It takes a lot of the value out of creating those kinds of things. The customer experiences this as a reduction in cost.
These changes happen with every new gadget. Photographers have to compete against everybody with a phone in their pocket. Now they just do weddings.
My grandfather was a portrait painter in the 1940s, a profession that was already moribund from the proliferation of photography studios. He gave up and became an insurance salesman.
Maybe he got screwed. The photography studios are all gone too. Because anybody can make a portrait using the phone in their pocket. At zero cost.
Then we can verify that there's nothing nasty hiding in it.
For some kinds of risk (ex: walking people through on how to make infectious bioweapons) an open-weights approach would increase risk.
If you are motivated enough to assemble all the kit you'd need, and actually do it, then you should be motivated enough to find the knowledge to do it without chatgpt etc.
I'd imagine the main thing that's stopping people is biological weapons are a terrible tool to do what most people want to do - which is target specific enemies.
So while the materials you'd need to build such a thing are much more accessible than say nuclear material, it's much less attractive - but if already had somebody mad enough to try - it's already possible.
It's not a matter of a one thing standing in the way of people making bioweapons: there's a long chain of actions one would need to complete, and many places for the chain to fail. Access to expertise can reduce the chance of failure at many of these steps, and AIs can increasingly substitute for human expertise: https://securebio.org/benchmarks/
> If you are motivated enough to assemble all the kit you'd need, and actually do it, then you should be motivated enough to find the knowledge to do it without chatgpt etc.
If someone was motivated enough to assemble the kit you need for anthrax and actually distribute it then you might expect them to also be motivated enough to find the knowledge to identify an appropriate strain, but in fact this is where Aum Shinrikyo failed in the 1990s: https://en.wikipedia.org/wiki/Aum_Shinrikyo#Incidents_before...
Lack of knowledge is one of many factors that can lead to failure, and LLMs make it less of a barrier.
> biological weapons are a terrible tool to do what most people want to do - which is target specific enemies
Except:
1. There are also people out there who want to kill everyone. They're sufficiently rare that no one with the motive has also had the means, but as technological progress keeps lowering the bar the risk of motive and means intersecting increases.
2. This didn't stop the Soviets. They did an enormous amount of very dangerous research with minimal logical application.
Ultimately what stopped them ( and everybody else - let's face it it wasn't just the soviets ), is it's a bad idea.
I think you way over estimate how hard the knowledge bit is in a biological weapon ( if you just want to kill indiscriminately - specific/controlled targeting a whole different ball game ).
The barrier is motivation and kit ( though the kit is much easier to access that say the stuff you need to build a nuke ).
The idea that all that's stopping some disaffected Joe sitting on his sofa at home and suddenly deciding to make a biological weapon is he can't get a set of instructions from ChatGPT is to miss the point.
Ultimatelty it all comes down to are you motivated enough - if you are then the knowledge is out there with or without an AI summary.
That's a surprising claim: usually when I make this argument skeptics say that the knowledge barrier is so high that an LLM won't help enough!
I'd also be curious to hear what you think of the Aum Shinrikyo case, since that seems to me to be straightforwardly a failure of knowledge.
> The idea that all that's stopping some disaffected Joe sitting on his sofa at home and suddenly deciding to make a biological weapon is he can't get a set of instructions from ChatGPT is to miss the point.
I agree motivation is a huge barrier, in the sense that almost no one would do it. But to keep catastrophic bio attacks from happening as the knowledge barrier decreases it's not enough that most people wouldn't have the motivation. If 0.0001% of people have the motivation and 0.0001% have the means then we're probably ok, but if 0.0001% of people have the motivation and 0.1% have the means then we're not (0.0001% * 0.1% * population > 1).
— Governments, companies, nonprofits should invest in free, open source.
Open weights + deterministic orchestration feels like the only sane long-term bet.