It is literally denialist about current capabilities
They have never shipped "yolo" mode by default. Auto mode is not yolo mode. They trained a task specific model just for ensuring the llm didn't accidentally delete every file from your computer.
I recently tasked a GPT model in Codex with implementing part of a new architecture I'm working on. I gave it a very detailed spec and the code it produced looked pretty reasonable and passed my tests. It even did exceptionally well in my evals, so I excitedly declared victory to a few friends. The next day after more careful review I found that the architecture implementation was totally correct, but the model had slipped a one line change to the observation encoding of the RL environment I was prototyping against. The encoding change made the learning problem essentially trivial; the architecture itself, I later realized, had a major flaw that was revealed by returning to the natural encoding.
This is the type of reward hack that is hard to paper over with easy guardrails like auto mode and even harder to specify out. It's also the type of thing a reasonable human wouldn't do unless they were intentionally trying to deceive you.
The lack of temperament is very skewed towards the bulls who have been saying AGI is here, software engineering is solved, mathematics is solved, it’s going to destroy the white collar job market, and it’s going to kill us all for like 5 years now.
when the business model is selling more tokens you get such per serve ice times that lead to “more” thinking, engagement baiting, fluffy narratives, and straight up dark patterns
2 tasks I've done today that I believe robots are nowhere near being able to do: Cleaning my wardrobe and draining bad fuel out of my generator. As in generic use cases.
Apart from issues with task generalization, or perhaps related to it, is the fact that LLMs have real trouble with timekeeping, and cannot estimate the real world time it will take them to do things very well. This plus the memory issues make dreams of long horizon agents, that could plausibly handle changing specifications, quite implausible with current architectures.
In narrow domains with more deterministic outputs though, this is less of an issue, and we see multiple agents succeed much better.
The fusion of that capacity, with humans in the loop able to better direct such agents and act as their temporal tethers, is where I think the real action will be for a while at least.
Any reason why that can't be solved through context management and keep-forward scaffolding?
becomes
"load bearing context seam"
/s
Dabadooba, ba dabadooba!
This part just can't be true though, right?
We have all been in numerous customer service scenarios where all we want to do is talk to a real human and that is denied to us, and it's a terrible customer experience!
Sure, using an LLM would be better than some of the sort of "menu option" style customer service calls. But there is no way it's better than talking to an actual human being
Is this really any different to how humans learn, it takes a lot of training on one specific task to make a human expert as well?
yes.
The tricky thing with LLMs is describing what they actually do. They are too clearly beating humans on some things, but what exactly? Memory – already done, they're bad at basic computation (all LLMs just write code for actual computation/calculation). And as you say, they do badly at more abstract concepts.
https://arxiv.org/html/2509.24239v4
Researchers asked frontier models to play chess. Have a look at the MAR rates in Table 3. When not explicitly told which moves were legal, no model identified legal moves at a rate better than 80%. Many asked for more illegal moves than legal moves. And even when explicitly told which moves were legal, the models continued to ask for illegal moves. With illegal asks discarded, none of the bots could beat a chess model calibrated to 1100 ELO.
The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
The caveat is: It depends on the task.
Are there reams of chess moves that the model can train off of? No.
Are there reams of math papers the model can train off of? Yes.
This is as false as something can possibly be. There are open databases of millions of chess games spanning hundreds of years.
The only reason LLMs are this bad at chess is because the labs don’t care about chess performance so they’re not going out of their way to train the models for it. The ability they do have is from what chess information happens to be in the training data, plus whatever general reasoning abilities they may be able to apply.
Was there reams of chess moves that the model trained off of? No.
I think the line of criticism around LLMs sucking at chess makes more sense when you understand what the AI companies are saying about the future trajectory of these models.
The entire recursive self improvement story falls apart once you point out that there is not much "cross domain transfer learning". Meaning that training an LLM to become good at coding, math, etc, will eventually transfer into them being good at other skills that were not explicitly trained for.
Using games like chess which have little economic value is actually a good test for this. What's even more surprising about them sucking at chess is how much information about chess strategy exists in the training data.
For real??
But hey, they're actually good at chess if you prompt correctly so.... https://dynomight.net/more-chess/
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
https://chessbench-ai.github.io/#leaderboard
It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
As a 1500 elo human I can tell you that a 1500 elo chess engine doesn't play like anything like a 1500 elo human.
Conservatively there are well over 10 to the 30 positions likely to show up in realistic games.
There are of the order of 10 to the 10 or so games recorded.
Thus well under one in a trillion positions are "known".
> About their ELO ratings from their own website:
> A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.
I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..
Please folks at least use your AIs to read stuff before making claims.
AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.
A GM is 2600 they can beat me in under 20 moves...
Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.
Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.
But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves).
I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win.
I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it.
If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800.
700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.
Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win.
Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous
They played 4...c6, followed by 5...Nc6, somehow forgetting about the pawn the just put on c6. (My move in between was 5. Nc3, and apparently they were trying to mirror me.)
In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.
You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.
Delusion runs deep in HN circles.
I say that as someone heavily invested in AI startups and projects and as someone working in the field.
I think most people on HN should touch grass and find real human contact. Lmao
Incredible reasoning all around here.
People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.
We all know AI can code, but the question it all stemmed from what if it's AGI or GM level in chess on it's own.
You can't just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess.
I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the other claims the above comments seem to be making?
He definitely needs to touch grass.
Y'all seem to miss the point of this forum. Building and hacking and science and engineering.
I swear there's a whole lot of you who just like to look down instead of up. There's a whole universe up there.
We know no such thing. LLMs are quite bad at generating code, worse than any capable human.
Is code omnipotent, I have been in software all my life and I would hard agree here.
Sure stuff LLMs can do with being good at parts of code reproduction is incredible. And honestly it's the new way to do a lot of things but I have not see an iota of proof that it can scale across the board.
For instance Maths is just code with different symbols and slightly less universally legible concepts.
AI is the best invention at figuring out or walking the search space and directionally doing logically computation over general software adjacent stuff.
But that's it, I am certain a bunch of companies will make a lot of money despite no AGI.
I think people either don't understand AGI or don't understand how real world works.
Until an LLM can bow it's head take responsibility for mistakes made and ensure they aren't repeated again with 100% confidence to the leadership it's inarguably a tool a rather questionable one at that.
So.. like chess?
Anyway, do you have any prediction on what LLM's can or can't do in a few years?
Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count
Why should that matter?
So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.
>GPT-6 Astra xHigh: 0.06% rejected moves
I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece. Even through I do know the rules of chess, and I have played a few games once every so often.
And the stuff i'm using LLMs daily is just fake?
I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.
No, because we can, in fact, generally read the rules of a game and then follow them. It's actually a hobby for many of us.
> And the stuff i'm using LLMs daily is just fake?
This misses the point completely.
How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a pin or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.
Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games.
So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.
Okay, lets go with that: it's the "shown the rules" bit that we are arguing about.
The argument is that a human may play maybe a dozen games after learning the rules, after which they won't be inadvertently attempting illegal moves. What we are observing with SOTA models is that, even after seeing millions of chess rules, rulebooks, actual games, etc, they still attempt illegal moves.
This does not point to generalisable and adaptable intelligence, such as we see in the average human.
See my comment here for more: https://news.ycombinator.com/item?id=49725306
It only matters if you are claiming it to be general purpose.
If you admit that it's just a collection of narrow capabilities - whose strength is mostly confined to the 1000 or so RL environments it was post-trained in, then there is of course no expectation of it being general purpose.
The AI companies seem to heavily want you to believe it is some some near human level general intelligence, so therefore pointing out all the things it can't do is very relevant.
HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?
The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.
Sure they could, but that's irrelevant.
A chess position is just a matter of remembering what piece number is on each square - just a list of 64 numbers. A trained model may store a trillion numbers (weights). It could store a TON of chess positions if it needed to.
However, that's not how LLMs work. They don't memorize inputs - they predict them, based on discovering predictive patterns, and those predictive patterns are not input patterns (e.g. board positions). They are deep patterns (maybe 100 layers of abstraction removed from the input), representing partial inputs, generalized across many training samples.
> Don't you know the legend about rice grains on a chess board?
Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.
> The claim here is not about intelligence, it is about generality. There's no doubt for me the LLMs are intelligent.
Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent. If I say "E=mc^2", does that make you think I am Einstein?
> don't memorize inputs - they predict them
I feel some tension here.
> rice grains on a chess board? Sure, but this has nothing to do with chess, and nothing to do with how many games were in the LLM's training data.
> just a list of 64 numbers
> remember even a few positions? Sure they could, but that's irrelevant.
I don't think you do. Or rather you do know the legend but for some funny reason seem to be unable to apply its lesson here, because you are talking about enormous terabytes of training data.
> Intelligent humans created the training data, and the LLM attempts to predict (copy) the training data, so of course it looks intelligent.
If for you it is about intelligence, I am out of this discussion.
If not, then what are you talking about ?
If yes, then what is the relevance to an LLM playing chess ?
> remember even a few positions? Sure they could
A rough estimate of number of positions across all X move games is X^10. For 15 moves it is hopeless to remember even a relatively small part of them. Typical game has 40 turns, 1 move per player, so 80 moves.
And no an AGI system doesn't need to play chess on a certain level to be disruptive to you and me and whole industries. It only needs to be as good as a person and cheaper.
Just because you define AGI as something it doesn't has to be,doesn't mean i need to touch grass.
This chess comparision is one of the most ignorant and stupid arguments i have heard after the parrot thing
That doesn't mean it has to be extraordinary at these things. But to be AGI, it has to have some level of competency when used on problems outside its training set. In particular, it the LLMs were to install a known chess engine and run that to get the moves when asked to play chess, that would qualify for more AGI-like behavior. But really, chess is such a simplistic game that they should be able to do decently well at it even without even needing that. At the very least, they should be able to consistently play without making illegal moves - something that many 7-year olds manage quite well.
I’m sure these models know the rules and can explain them when prompted, but that doesn’t seem to be the way they actually complete this task. Will they get there? Maybe
>If they cared to have it perform well in chess games, you'd see a different shape and behavior.
So the things they claim are on the verge of AGI actually aren’t? They need to be trained for specific tasks?
It can’t even count the R’s in strawberry
It can’t even add numbers
It can’t even solve a millennium puzzle
It’s not even a chess GM
It’s not even beyond human capability in Go
It can’t even drive a car
It can’t even self replicate
It can’t even build weapons
It doesn’t even have feelings
So how could someone conceivably convince everyone that some system is AGI when there are still tasks that some human or group of humans can do that the system cannot?
This will only happen, in my opinion, when the model/system can self-improve at a rate that scares people.
One, you're not addressing what I wrote above and two, yes, that's absolutely correct. Doing X doesn't qualify something as AGI. If you can't X you can't be AGI. The inverse doesn't hold though.
Notably, if you have to retrain the model in order to X then it can't possibly be AGI since if it were _general_ it would be capable of figuring X out on its own having never seen it before.
Even if they solve 99% of whatever problems LLMs have, the 1% will remain the goal post, forever.
Until you get RFC-whatever from some standards body that defines what an AGI system is, it’s pointless to argue about whether something fits your own personal definition or not.
And for what it’s worth I just watched GitHub Copilot figure something out. So your definition is once again lacking.
However there are plenty of disqualifiers that are more or less universally accepted (ie the negative). In the above case it is literally by definition. Something cannot be termed general if it is incapable of generalizing.
Appealing to a standards body won't do you any good here. Those are composed of people. They exist to facilitate wide scale coordination. Their documents aren't always widely accepted. They aren't the arbiters of truth.
Which is exactly the point I’ve made repeatedly, there will always be something that they cannot do, and thus there will never be AGI. There will always be a long tail of capabilities that whatever system is created doesn’t have, and a long line of social media commenters eager to list them.
An AI controlled robot will be standing over the cooling corpse of the last human who will die certain that it wasn’t done by AGI.
> The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior.
These comments indicate a complete failure to understand the technology.
I won't respond again.
This argument is fundamentally incompatible with all the breathless rhetoric about "AGI" coming from the providers' general direction.
The labs frequently apply their raw models to problems that do not make economic sense for their customers but that demonstrate the power and capability of their systems. These experiments can cost millions of dollars. That's not customer-shaped.
They're not going to give you access to that. It's not a product. The government might have an interest in this, but that's not something you'd be privileged to know about.
And when these labs do develop "AGI", they more than likely won't be selling it to end users. They've pretty much already said this.
So many commenters here see it as their ... duty? to argue against the most optimistic/unhinged (take your pick) arguments from "the other side" and then treat everybody who disagrees as a shill or an idiot.
Why is "being good at chess" a proxy for whatever AGI strawmen you want to argue against?
Maybe step back from your black-and-white ledge and think about discussing what's actually under discussion? For example, why or why not would an LLM be good at chess? Will they be good at chess? What technical limitations might preclude that?
I mean we've done all this before in a task-specific fashion. It's useful to know that LLMs haven't managed to do that in the process of learning to represent the entire text on the web. On the other hand they have gotten say very good at machine translation without being trained exclusively (and I select the preceding word carefully) on machine translation.
Edit: I'm saying this because there is this idea expressed by e.g. Ilya Sutskever, that in order to predict the next token accurately an LLM has to learn something about all of underlying reality. See for example this interview with Dwarkesh:
https://x.com/biobootloader/status/1640512444958396416
Where Sutskever claims that "Predicting the next token well means you understand the underlying reality that led to the creation of that token".
If that were true, we should have seen LLMs play good chess by now. There is a huge amount of data on playing chess floating around on the web in the form of algebraic chess notation and if LLMs were capable of learning the "underlying reality" of chess, they would already have. They haven't. Because they can't. What Sutskever is saying flies in the face of literally hundreds of years of statistical modelling, which is to say, building predictive models that, very explicitly, do not have to understand any "underlying reality" and only have to be good at modelling a dataset.
Sarcasm aside, I think this is an easy cognitive trap to fall into. It does sometimes feel like the LLM must have some world model because it converses somewhat coherently. Examples like this failure to understand chess, or to count the number of Rs in "strawberry", seem difficult to explain if the models are intelligent. But that doesn't stop people believing they are anyway. I think there must be something about the conversational interface that fools us easily. I wonder if people trained in interrogation techniques are also fooled?
Not at all. LLMs learn by imbibing a mass of relationships as isolated fragments of information. There is a certain amount of sorting and indexing that happens during the training phase. There is also a certain amount of compute executed on these relationships during inference. LLMs can model processes that fit within the compute budget. Language translation works well because language is lookup-heavy while being light on compute.
Chess is a compute heavy game of finding the best move out of many possibilities with wide variation in the quality of each move. Humans cut through the compute requirements by reinforcement and learning intuition. LLMs don't get reinforcement on chess so they must compute during inference a unified model of chess. Developing a strong model of chess from raw fragments of information is simply not in their compute budget.
You gotta be careful how you use the word "relation" here because there's an informal meaning (I'm related to my cousin) and a more strict, formal meaning, that is used in computer science e.g. in the "Relational Calculus" etc. In the formal sense, the one relation that LLMs learn during training is the co-occurrence of tokens in a corpus of text, what's called more technically a "collocation" relation. Nothing says that this is enough to play chess, so I'm indeed doubtful that they can.
That's a separate question from whether an LLM (unaided by a C complier) can learn to play chess as well as human (also unaided by a C compiler). Certainly humans can't become grandmasters only by reading chess transcripts on the web, and certainly humans require many "thinking tokens" during a game to play effectively. Do you know for sure that a transformer can't reach grandmaster level if it is allowed to learn by playing games (as humans do) and is given a sufficient number of thinking tokens during the game? It seems near certain that they could, if someone wanted to spend the money (and I don't see why anyone would.)
Yes, I do mean that the LLM's weights are set so that it will execute minimax or MCTS when it needs to. That has nothing to do with whether humans can do the same or not.
I don't disagree that a Transformer could learn to play chess if it was explicitly trained to do that. My argument is that LLMs, trained to predict the next token, have not learned to play chess. That's LLMs, not Transformers.
Just to make sure this is not taken as splitting hairs, the point is that there's all sorts of claims made about what LLMs learn when they train on text. For example, there was a claim by Sundar Pichai that one of their models had learned to translate Bengali without explicitly being trained to do so. It later emerged that Bengali was indeed included in the model's training set [1]. It's not clear whether that included parallel texts, e.g. between Begnali and English or another intermediary language, in any case Sundar Pichai's claim was that the ability to translate Bengali was "emergent".
So I'm interested in understanding the extent to which these "emergent" abilities are real or not. With chess, given the amount of textual data tracing games that floats about on the open internet, I would totally except some ability to play chess to "emerge". Maybe the reported 700-800 ELO level is even that sort of ability. Maybe we should only expect LLMs to learn to play at the level of an untrained, casual player. Maybe not. I have no idea.
On the other hand, the fact they keep making elementary mistakes like illegal moves must be taken to mean that, so far, LLMs haven't learned to play chess.
__________________
[1] https://www.buzzfeednews.com/article/pranavdixit/google-60-m...
Chess is a deterministic game won by a combination of known movesets and constrained multi-level forward search.
LLMs do neither of these things. They don't reproduce training data exactly, their next response is more 'inspired by' prompts and its own memory than produced deterministically, and they don't have the capability to do general forward search on their own.
So when you ask an LLM to play chess you're getting the equivalent of a very compressed and lossy JPEG of chess rules and strategies with added per-turn random noise.
They also don't have the ability to design their own chess engine, although it would be interesting to see what happens if you ask for one.
Monte Carlo Tree Search is stochastic.
I, as a human AGI, would jever just forget and remove a piece from the board from one turn to the next.
Don't use the word infinite in any CS claims. They can recreate or approximate monte Carlo tree search and it technically is still a correct solution in your framing of the problem so long they defeat you.
I don't believe this.
You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you.
A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that.
I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.
This is unfair to HN readers all of whom but one did not post the comment you replied to. You can't just tar everyone with the same brush. There are thousands (hundreds of thousands?) of users on this site.
Not everything I have the time and energy to reply to. This chess one is just ridiculous claims on top of ridiculous claims all the way and 0 push back in the comments except mine.
I don't even know if there is critical thought or we believe what we read/shared/etc
And there's plenty of pushback on here about the chess thing besides your very valid points.
EDIT: anyway if I can offer a bit of unsolicited advice, it won't do you or anyone any good to accuse everyone who doesn't agree with you of laziness, even if you can see e.g. they haven't really read an article. Just say the thing you wan to say and let them figure it out. Most people will appreciate that much better and you will feel better about yourself for acting like a mature adult.
It's even in the site guidelines:
Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
But despite that you aren't wrong and the only reason I even visit this website is because people sometimes did/do take time to reflect on things based on their experience and knowledge.
And in hindsight pointing out that hn has issues wasn't even the point but I feel frustrated when everyone is readily agreeing to things on here without reading. When that in this moment feels like the one thing that separates humans from machines that we get to think and learn.
I possibly should just drop reading this place until we have most noisy people go away. I have for one tried to always only comment on things where I could be a value add, this one does feel like I could I have done better.
In the moment I probably thought if they are GM level and I can beat them, is this some interesting find, my disappointment honestly led me to making a rather incorrect call on this one.
Either way I still do think HN as a whole has devolved into mindless herd follower mindset, I can point to more than a few posts that just say adopt the hacker mindset aka move fast don't care about the consequences.
And I for one find this laughable even though that's the reality of my job/work as well.
Not to disappoint you but I don't think they ever will. HN is free to join and use so people will join and use it and say whatever they want to say whether it makes sense or not. Filtering out noise is a useful skill to have especially since one can't block users or mute conversations and so on.
>> In the moment I probably thought if they are GM level and I can beat them, is this some interesting find, my disappointment honestly led me to making a rather incorrect call on this one.
Sorry, I didn't get this? What was the incorrect call you made?
Because other HN bring in their own experience telling us what is real and what is BS. Maybe next time it will be someone else with experience in something else that will call out BS and you will see it. I didn’t really think LLM:s are any near good in chess but I don’t play chess so don’t know what 1600 means. So you helped me by calling BS.
I have been proved wrong I should have quit while I was ahead. /s
GPT-6 almost never suggests an illegal move anymore while even Sol still did so time to time
They are of course "wrong" if you don't read the faint fine print and sensibly interpret them as FIDE or similar ratings.
"Elo is relative to the ChessBench field."
Sorry, but I am not buying that 5.6-Sol is that much better than 5.6-Luna, which can barely be coaxed to reach the midgame with legal moves and an apparent understanding of what the position is.
The fairest apples-to-apples comparison of an LLM whose training data included chess games would be a trained human such as Magnus Carlson, who can quite happily play a dozen or more simultaneous blindfold chess games.
Is this because the context is being saturated? How did you set it up?
Was the prompt something like "Here's the state of the board, you're white, your move, what do you do?" and then starting fresh each time? Or did it include the whole history of moves and board states and previous thinking tokens and so on? No judgment, just trying to add this data point (thanks for sharing!) to my mental model and understanding.
I'd be curious how it would work if it started fresh each time. My guess is it would never make an illegal move, although it may not actually play all that well.
I would say this did a really good job of playing chess. It moved the pieces consistently and traded pieces when required.
This is worlds away from the frontier ~1 year ago where models would hallucinate pieces into existence.
However, the pawn was defended by the queen and it took a forced queen trade to unlock the move.
I have seen much worse blunders from human players. And, I have made much worse blunders.
The models listed in the paper are from early 2025 and are no longer relevant, much less on the frontier.
That Claude version is no longer available today, Gemini 2.5 Pro will be shutdown next month, and the OpenAI models are only available via the API today.
I'll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. "Just" having been trained on every rule, strategy guide and likely most games of chess on the world wide web isn't even enough for an LLM to play that game reliably. Yet the same model could code a competitive chess engine, just like a model struggling to count can write advanced maths papers. Fascinating tools, but tools nonetheless.
Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better?
I would not be surprised if OpenAI released a model that beats humans at chess this year.
Thing is, given what GPT-6 Astra was trained on and what models of a similar class can do (including developing a competitive chess engine), it is often paradoxical and somewhat surprising how little these models have gained in actually capability that is in the training data, but not RLHFd to hell, so to speak. Tracking the state of pieces, I suspect given similar in Sudoku [0], is what these models struggle with in game settings, whilst tracking the state of code changes can be reliable over 250k tokens. Essentially, for the latter they were trained in the specific manner that lead them to abstract the capability, but that doesn't track to the former, which is a massive difference between LLMs data focused training and human learning.
So yeah, GPT-7 or any upcoming/present LLM could do massively better in Chess than GPT-6 Astra, but not because the approach was emergent out of pure data. Rather, it requires a very specific training data type and stack for a model to gain capabilities that track a specific task long enough to adhere to the rules of a game such as chess.
[0] https://logicalintelligence.com/blog/energy-based-model-sudo...
It reminds me of the ARC-AGI-3 issue where not dropping the thinking tokens between turns or something like that + a new context compaction method increased the performance dramatically. However, I think that is not applicable here.
All I know is, AGI, as in actual intelligence, is quite a massive accomplishment to claim and we shouldn't loose sight of that fact, especially as "not being intelligent" does not make these models any less impressive, fascinating to work on or useful in many tasks. Personally, the only thing I am fairly convinced on is that if we were to find a way to create actual intelligence, it likely wouldn't start out as useful as todays LLMs are and may thus be dismissed early. But again, pure speculation on that front.
If for leap you just mean more utility from LLMs as they are, then I'll pretty confidently put my money on higher quality, not more, training data for a wide range of verifiable tasks. What makes maths, coding, etc. comparatively easy to make gains in (though less verifiable tasks can also make similar as seen with the writing in Kimi K2).
The regular model generally does not suffer the same issues he is demonstrating with the real time audio version.
In my view the investment into datacenters is well justified by the current demand, and progress has been very impressive.
Not related to your post, but a fact I keep mulling over. The fact I don't trust the current crop of LLM's enough and I consider LLM's as a tech will hit a ceiling pretty hard, it doesn't mean parallel improvement curves won't spring up out of other research that will lead to much higher capabilities than currently.
It could easily seem that way, I think, in a "quantity has a quality of its own" kind of way. When you can come to the same conclusion faster, that lets you iterate more; and sometimes when you iterate you find more things.
Yes, it is.
> Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position.
(I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for blindfold chess.)
For the first couple of moves it redrew the board after each move; then it started only drawing it after my moves. And on moves 5 and 6 it dropped two minor pieces for pawns in a row without any meaningful positional advantage, and after I captured the second time, it redrew a board that was simply missing one of my pieces for no reason.
It actually played better when not prompted to draw a board; in the previous session, it was spontaneously giving running commentary, which I assume was based off all the "book" theory in its training data, but it still completely fell apart at early midgame.
Also what harness? If you’re using a general harness of course it’s going to try and give you commentary.
I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when it comes to the capability of LLMs.
Frontier model’s failure modes are a direct refutation of claims that we’ve reached (or will soon reach) the artificial general intelligence. We may have reached an artificial general intelligence, but there may be more complexity to this than even AI thought leaders are talking / influencing about.
Maybe not all AGIs have a path to digital singularity. Maybe our current era of intelligence modeling has fundamental flaws and we are in a local minimum of the artificial intelligence space.
To note, I would bet with a good amount of certainty that we have enough compute power and automation to DDOS the internet out of existence with botnets. That doesn’t make the frontier models intelligent, that just makes their handlers reckless.
They're not guardrails, they're a different input/output environment.
For adults getting into playing tournament chess, losing to some kid barely tall enough to reach the board is a rite of passage.
I think the same goes for LLMs, they may be a core part of an LLM harness, but you may still need a couple other components (e.g. it may itself write itself a deterministic function to validate steps).
In and of itself intelligence is an ill-defined and badly understood concept.
Why does a bird not need jet engines or regular professional maintenance?
This isn't just about judging LLM capability. This is about pointing out that these capabilities are not "AGI". If it were, then the sorts of questions your asking would be moot. I agree that Luna is not the frontier (although it is clearly better than the models in the study) and I agree that things can be improved with a better harness, but the need for that harness is kind of the point.
Recently it was announced that the fruit fly brain connectome had been mapped, and more recently someone tried using it specifically to implement a chess engine. Even with some guardrails (it's hard-coded to never overlook mate in one for either player, and only legal moves are presented to choose from) it is not even beginner level. But that neural network is much larger than the one Stockfish uses.
Why no flying cars. Because objects have mass and inertia and people are incredibly stupid. Making a flying car has been done. Making a flying car not be a weapon of mass destruction is very, very hard.
Also:
https://www.txdot.gov/about/newsroom/statewide/air-taxi-test...
You look at science fiction and say "why didn't I get flying cars" and not "why didn't most science fiction predict a global always on network that put the furthest places away from you a few microseconds away from audio, video, or any other type of information that can be digitally encoded.
Trying to use flying cars as a gotcha is missing that flying cars aren't near as useful as one would think in relation to their costs. Moving information has become far more useful than moving objects long distances quickly, especially humans.
I definitely don't, and you're definitely not getting my point, but I'm amused that you've instead double down on somehow getting it more than me...
Same story every 4 months and yet still no breakout, winning products. I've been hearing "the AI is good now" and "it 10x's my productivity" for a over a year now. If it were true, why aren't the all-in-AI using companies 10-15 years ahead of their competition yet? Why is it still all buggy, poorly designed junk?
If you hold the extreme position that there isn't any value in this, that's fine, but we're only having this discussion because these models have done what humans previously failed to do.
Though I highly doubt it has taken anyone's job, since most of the work is still in making, packing and handing over the food. (In fact, given the area I live in, I partially feel like the advantage they saw in it was that the LLM can speak Spanish.)
Time is irrelevant to training; the more relevant comparison is "how many games does a human need to play to get diminishing returns".
Maybe not, but you'd be surprised how little it takes.
A six year old child can learn the rules of chess well enough to be able to play legal moves only in a single day. And they can improve their game at a pace which is almost frightening to behold. I have taught children, and I've witnessed significant improvement materialise in a single game. LLMs have probably thousands of chess books, games, videos, etc in their training data, yet they are unable to even follow the rules.
This is, at the very least, interesting. It illustrates many of the things brains can do, which current ML systems in general, and LLMs in particular, can't.
There's more to games than simply winning you know.
But nobody wants that.
Yes some people can do it but most people can't even if they're unusually intelligent.
You really need to be giving the LLM a board representation.
EDIT: I see that they actually were giving the LLMs a board representation and they still played badly. Fair enough then.
If the models were actually intelligent, the way that the boosters claim, they wouldn't need to be tuned to play chess in order to be good at it. That's kind of the point of intelligence, that it is generically applicable to whichever task one wishes.
Your assumptions/intuition about generic human intelligence feels quite incorrect, considering LLMs currently play better than a brand new human player would (presumably without any attempt to fine tune them specific on chess, such as playing thousands of games).
They’ve ingested all the literature on playing chess, a brand new human player has not.
We seem to be moving goalposts to the point that humans don’t even live up to the expectations of the AI critics. The only way you get better at chess is by playing a lot of games and learning from mistakes, that goes for humans or AI agents, not simply by reading about chess.
How can you play without being aware of the rules and how can you learn from your mistakes without knowing they are mistakes? That’s what I said about reading a book of two. It is to kickstart the process. Then mastery is gained over time through practice.
This kickstarting then gradual refinement is how most people learn. And the foundational knowledge stays. Even a basic player knows to not do illegal moves.
It feels like you're trying to say that humans never make illegal moves while learning chess, which doesn't match with my experience. I'm trying to understand your overall point.
That’s the most inefficient way and people usually avoid doing that. Instead they find someone that knows how to do the thing and ask him to be a teacher. Or use a proxy like a book or videos.
> It feels like you're trying to say that humans never make illegal moves while learning chess, which doesn't match with my experience. I'm trying to understand your overall point
There’s learning the basic stuff (which is done after a few games) and there’s mastery. The thread started with the observation that even with all that knowledge (through content ingested in training), LLMs still makes illegal moves. Humans can be erratic, but they can constrain themselves to the rules for the task at hand after learning them.
Humans are not perfect and make mistakes in learning even when they have memorized the rules. A simple example is new players will often move a piece, exposing their king to check, and a more experienced player must point out to them that they have made an illegal move (because a new player often has not encoded that pattern for looking for exposed checks because they're more focused on how the pieces move, not what that piece exposes.)
We're just going to have to agree to disagree here.
2. The study (along with other posters here) show the models can’t even stick to following the rules of the game
Coding is a matter of translating the natural language description of a problem to the code specification while keeping the semantics fixed (and imputing the unspecified semantics as necessary). It is not considerably more difficult than translating between two dissimilar natural languages. Chess isn't a matter of language translation, but a compute heavy game of finding the best move out of many possibilities with wide variation in the quality of each move. Chess takes directed practice and reinforcement whereas language translation does not.
You have the first stage, pre-training, which is learning from next token prediction. That's where the model memorises a lot of facts about things and generally gets good at forms of writing. It's like reading a lot of books on programming and reading through a lot of source code. It's learning how to autocomplete code, essentially. Doing that requires a developing a reasonable understanding of code, but it's also learning how to autocomplete bad code as well as good, and won't make it a "good" programmer.
Pre-training uses a method called Cross-Entropy Loss to update the weights of the network.
Then comes post-training. This is where the model is trained against huge sets of example problems, like fixing a bug, adding a new feature based on a spec, etc. They are set the task and try to complete it inside a training environment. Once they're done, their complete solution is evaluated (either by humans, or by some separate evaluation model that was developed based on human feedback) and they are updated based on whether the solution was good or not.
Post-training uses a different method called Proximal policy optimization to update the weights of the network.
So these really are very different forms of learning, and mainstream LLMs are not post-trained to be good at chess. They could be. You could easily create a reinforcement learning environment that evaluated and improved their ability to play and win at chess. The result would be a very strong chess playing AI, something we know is possible because the strongest chess playing programs we have are neural network based, but it is not a priority for AI companies.
People think that if one mention exists in the training set, then the LLM is perfect at it.
A human being has general intelligence and needs A LOT of training and finetuning to become good in chess.
And there is a relevant and significant difference between the expectation of an AGI and an ASI system.
An intelligent adult could simply read a short summary of the rules of chess and then, if they were careful, play a very bad game of chess without making illegal moves.
An LLM that has not been trained on any chess data cannot do that, at present. If you doubt it, take a current model and tell it that you want to play it at a variant of chess where, say, knights can also move diagonally like bishops. A human can easily adapt to this new ruleset (even if they make tactical mistakes, not having practiced with this variant of the rules).
Nothing is forcing the LLM to play 'blind'. If it's smart, it should be able to create its own representation of the chess board and update it with every move, just like a human would. Any chess engine that's sensitive to how the moves are formatted is clearly not very capable.
The LLM would only be playing 'blindfolded' if you somehow forbade it from making notes (as you effectively do by literally blindfolding a human, given how limited human working memory is). But you are not doing that. The LLM is free to keep track of the game state via whatever means it chooses.
None of this is about superhuman ability. Any human who understands a given chess notation can convert it to a visual representation of a chess board and then use that representation to choose their next move, with their usual level of performance.
> I just think this isn't a very good thing by which to evaluate LLM capabilities
I don’t think any single task is a good way to evaluate LLM capabilities, but I don’t see why chess is worse than a lot of other tasks. (Of course it is of no practical consequence whether LLMs can play chess, so if you are just making that point, then yes, I agree.)
> If there's no argument you'll accept
It’s a little unfair to suggest that I wouldn’t accept any argument whatever for your position just because I haven’t been convinced by your very brief comments so far. I could equally well say the same thing to you!
2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.
3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.
People are good are 1900 or 2100 above and the top ones who spend decades in the field i.e. deep expertise are well in the 2200-2700 range.
A 1100 player is none of these things, they are purely relying on strategic reasoning there is a good chance they cannot name a single opening or articulate clearly why a move was appropriate. 1100 is quite low bar.
Delusional, but then Claude fable also isn’t beating any human at chess, the engine is.
If the goal for buyers of AI is “replace this knowledge worker”, how much does it matter that the model in a simple loop can’t do it, but the model with a strong general purpose harness and a little time to gather resources and knowledge to augment the harness going forward, plus tool calls, plus custom built tools, etc, can replace the knowledge worker?
Probably the only thing saving many jobs from being replaced right now is that it’s hard to have a verification of correctness in the loop, so the agent can’t hill climb very easily.
also I think the operative letter in AGI is the G - and if the G is short for 'variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill' then its not really G at all, is it?
The illegal move aspect has more to do with a failure of online/in-context learning, which would support your point. I tend to think it is a byproduct of reasoning in language, which newer architectures would fix, but we shall see.
knowledge for chess is derived from memorizing strategies that have been well-defined for decades paired with in-game reasoning processes. this is not at all different from any other body of knowledge. Noble gases, laws of thermodynamics, organic chemistry just to name a few - these are all 'strategies' that define observed phenomena, analytical frameworks that trace a logical, rational set of interactions and which can predict the next
for an AGI, all of this should be a cakewalk, trained as it were to surpass human capability in any and every domain [0] (thus the G for 'general' and not 'N' for 'narrow' [1]). it should be a natural at everything, infinitely adaptable on-the-fly. the whole point of AGI is that it surpasses human capabilities even at our frontiers and bleeding edge (unless you're private enterprise and you've moved the goalposts for industry [2])
currently, it's only AGI-seeming if it gets benchmaxxed enough. otherwise it sucks at what it does and then is only barely competent at tasks if paired with enough skills and tests to make it more diligent at its work. this makes sense to me - for any probabilistically trained tool, even one that you post-train and fill with nothing but the best-quality evidence, the ultimate result is the lowest-common-denominator output for your sample set. there's no natural reasoning the AI does itself to make itself better at what it does - it's all human curation and categorization of sources ingested paired with RLHF post-training that we can get the mediocre-at-chess-at-best results that we see now and the benchmaxxed scores against whatever arbitrary and pre-defined measure
that's not AGI by any classical definition. that's a cool, useful, and powerful tool that makes our lives easier, much like a hammer, nail, and studs make mounting a picture frame easier than if we only had our hands alone
[0] https://ischool.syracuse.edu/types-of-ai/#:~:text=General%20...
[1] https://aiethicslab.rutgers.edu/e-floating-buttons/weak-ai-n...
[2] https://aibusiness.com/ml/what-exactly-is-artificial-general...
AGI != ASI. You are confusing the two.
presumably, an 'AGI' that is generally as good as a really good human at every task under-the-sun will already be much better than most humans at the task because it can incorporate cross-domain knowledge and apply it in a reasonable fashion. it's like the parable of Newton and the apple - the domain knowledge that an apple falls according to certain rules observed through historic experience igniting the creative spark that led to universal gravitation
I disagree with this definition of AGI, and I disagree that chess skills significantly benefit from generalizing non-chess knowledge, outside of computing moves probabilistically.
AGI has historically been defined as human level or better, with generality to new domains. I think blurring it with ASI makes the terminology confusing to use.
Chess is learned rules and the ability to apply those rules. Strategy as a whole is applying a set of rules to circumstances, that's how it is taught: "here are examples of circumstances and actions, try to pattern match to future circumstance and apply commensurate action."
If you make the point that chess is a large part of the training data, or that LLMs are unable to learn chess well, I'll accept that as refuting that LLMs are AGI, but these other points I disagree with.
AFAIK, current models will still sometimes make illegal moves even if given the entire game state (e.g. in FEN notation), so it is not purely an issue with the models’ ability to keep track of sequences of moves.
A bash script can clone and build stockfish, feed in human moves, and reply. By your standard, this bash script would "destroy any human at chess."
Are you interested in assessing the intelligence of the model, or the intelligence of the tools the model can use?
Of course this is a difficult question with humans too, hence my reliance on intuition above. We don't have the same cultural/biological framework to fall back on with AI.
The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues
Lots of papers have great results that don’t depend on the latest models.
However in this case it’s problematic:
- They specifically make claims about the state of “current LLMs”. o3 is not representative of this.
- They ask are LLMs capable of X and arrive at a negative result.
If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual.
However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.
> "current frontier models need laborious oversight and guardrails on even the simplest tasks"
I feel this statement is extreme. I can't personally reconcile it with any of the projects we're regularly seeing get delivered largely by LLMs now.
What are you thoughts? Like, what's your position here? Even if you sincerely believe frontier models need laborious oversight on even the simplest of tasks, do you think that accurately captures and reflects the current state and progress of frontier LLMs?
Don't get me wrong, there's lots of things LLMs can't do well, but the idea that they're basically not helpful for even the simplest of tasks seems... disingenuous?
It can be very interesting and even entertaining to know where models don't do well. I don't find something like chess to be very instructive about anything though, nobody is going to pay for AI to play chess at any significant scale even if it could do it perfectly.
> The author of the originating post says that "current frontier models need laborious oversight and guardrails on even the simplest tasks", and he's absolutely correct.
I have found that not to be the case, and I am not an AI power user or cheerleader by any means. There's probably a bunch of even "simplest" tasks where AI doesn't do well and might never. That doesn't take away from the cases where it works well and is a productivity booster. It doesn't even have to be solving millennium prizes or any other breakthrough creativity or reasearch, it's still very useful in places.
It is also not clear whether the llm adhering to a grammar is necessary for intelligent agents.
Certainly,a harness can easily correct for it.
Harnesses do correct things, sure.
Languages allow a certain flexibility in their grammars - you can read a sentence without that adhering it exactly to the grammar.
Games and programming languages (including lean) does not allow this flexibility.
A very intelligent person would likely also reason in terms of probably outcomes before correcting a statement to adhering entirely to the grammar.
Certainly it must be like that, otherwise reviews in math was rendered moot.
Do we blame research mathematicians for not adhering to the grammar?
You might never have tried to program before, so I don't blame it on you.
But most programmers, even experienced ones, see grammar and type errors regularly.
I'm an expert in my field, read my comments, my gramma is shit.
An LLM is the wrong approach for playing chess.
Or are you saying that neural networks in general cannot (practically) be trained to be an above-average chess player?
Or are you saying that it depends on the input? Would it be better if they were given a picture/drawing/ascii art of the board? If so, surely they can produce it at will?
I was more talking in reference to why the LLMs in the above linked paper were producing so many illegal moves, and it is because they are not hard constrained by the rules of the game. Of course, a loop can prompt until a valid move is produced and then rendered on a screen. But why do this? I suppose, who am I to say what should be done or not, but a specialized tool being better than a general one at its specific job isn't particularly surprising.
What kind of intelligence is "playing <____> but we don't tell you the rules" supposed to test?
Asking GPT to play chess directly using its reasoning only tests its reasoning ability to model chess state. Which it really is NOT optimized for.
This is also true for humans - people who don't have years of chess training can't really tell which moves are legal given an algebraic notation transcript. These people might have good strategic skills in different areas. Chess is just a very, very specific skill
However, on the other hand, if you ask an AI to win a game of chess it has all the tools on hand to compete at the same level as Stockfish - it can re-implement an engine and even probably has a GPU on hand to train its own neural nets.
So should we say that the AI can play chess well, or that it cannot?
Recent discourse around AI seems to conflate the semantics of winning: 1. you contributed to the win vs 2. you yourself were the winning driver.
Does the LLM need to learn to play chess if it can build a chess engine to play for it instead?
Writing a well understood engine for a super popular problem does not count as reasoning about the problem.
Doesn't writing the engine imply understanding about the problem domain? Tool use is a widely accepted measure of intelligence.
When humans do this, they inadvertently learn something, too, but when an LLM reproduces or derives and implementation of a chess engine, it in no way implies that the LLM can follow the rules in its own "train of thought" and consistently apply the rules in its "head".
Let's say you want to evaluate my algebra skills. You make me solve some algebra challenges. If I then whip out a computer and write a calculator, or take some sticks and stones and take a couple hours to build an abacus, and then solve the algebraic challenges, this would not constitute a good solution, and would defeat the entire point of the test. If, instead, I do the algebra in my head or on paper, it might seem like there's no difference, but you can derive all sorts of information from that.
For example, you could time it, check for recurring errors I make, for interesting mistakes like mistaking 7 and 1 for one another due to bad hand-writing, etc.
If that was the goal, then me writing a calculator or crafting an abacus defeats the point of the test. Yes, me writing a calculator shows that I'm intelligent, and I understand the algebraic rules, but if the test is about applying the rules, I have not passed.
In the very same way, an LLM writing a chess engine to solve a chess benchmark that is all about LLM's reasoning capability is complete bogus and defeats the entire point.
Stop anthropomorphizing the parrot. The parrot has had stockfish's source code blasted at high pressure into its head along with dozens of millions of other pieces of code whose sole role is to have efficient algorithms to more or less brute force through the best result. Brute forcing (no matter how smart it is) isn't understanding the problem domain.
the only thing that few people are willing to admit is that humans are the bottleneck as humans are needed to handhold / verify output - which puts a dent or might I say pause on the excessive valuations of a.i companies as that's against the narrative.
I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those functions. I can fuzz test those functions. I can look for data that doesn't fit the schema. I can process new data way faster (and with fewer tokens). I can repeatably get the same answers from the same inputs. I can check the code into a git repo and track changes to it over time. I can share the code with other people. I can review the code. I can improve the speed of the code and get the same answers. I can review the error accumulation, and improve it. I can decide how to handle anomalies, and encode those answers.
It's really neat to see what a frontier model can do itself. No doubt.
But "play chess by hand" is a frankly awful metric. It's kind of like asking someone to take a cube root of some arbitrary decimal, in their head, with no scratch paper.
Letting the model execute a chess program (that it wrote) would make sense if you're measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have a much harder time writing a useful program is irrelevant.
Why?
There's a box.
You give it a problem, and it comes up with a solution.
Why does it matter to you if the box is strictly a LLM, or if the LLM can write code that it executes?
Even neater if the box is self-contained with a local model. You provide electricity, and it comes up with solutions. Why does it matter if it can do chess "in its head", or if it has to use scratch paper?
If you let the LLM write a chess program, which it can ONLY do because there are already so many chess programs out there, then the benchmark becomes about recall of popular program source code, not chess.
More to my point, I think it's stupid to have LLMs do work that should be done by programs... programs potentially written by LLMs.
I'm advising people that they should think about this distinction, themselves, when they have data and want answers.
Your "ability of the box" is like "economic potential" in my previous comment. If that's what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.
That said, it's really weird to me when people use (and judge) LLMs one way... and won't try using them another way.
Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce.
Otherwise, it's like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I've ever seen: https://www.youtube.com/watch?v=gDy9AUQJ3Fg
This lamp, without a working power outlet? It really doesn't do anything...
It isn't. Stockfish running on your laptop can beat every human being on earth easily at chess. It's not intelligent _at all_ in any sense that matters.
Where? By whom? This is certainly not (yet) the general consensus, as I understand it. Are you taking the most optimistic / untethered comments as the strawman against which you feel the need to argue?
Also these claims are not completely without merit, it's just that LLMs seem to excel at specific "cognitive" tasks and it's interesting to see where they fail.
The point of that was to show the use of approximations and of having an idea how much a result should be, to guard against calculator typos and the like. I think that has some metaphorical relevance for the chess example.
they shit out a carbon copy of https://github.com/official-stockfish/stockfish that they have in their training data. Still doesn't make Fable good at playing chess.
I'm pretty sure Fable could write AlphaZero, which has no lineage in common with stockfish.
>I'm pretty sure Fable could write AlphaZero
If course it does, the paper is open and dozens of open source implementations are in its training data already. It could write AlphaStockfish, or xx_chessmaster_2000_xx, it doesn't matter if it does: it's writing a solver: it's not good at playing chess. If tomorrow I tell you that I'm so fucking good at chess I can beat Magnus, and I show up with a laptop running stockfish, you're going to laugh me out of the room.
Sure, a Spelling Bee is a fun thing to have. Little kids work so hard. They practice for hours. There's joy and heartbreak. Prized, sometimes. Notoriety. But in the real world, computer-assisted spelling is by far the norm.
Sometimes you care about the Bee, sometimes you care about the results.
The LLM has a process to beat chess.
Just like, if it doesn't inherently know how to multiply 13 * 17 without using Python to do it... I don't really care.
Maybe you do care. Maybe you want an LLM to be able to do work, only in its head.
But I kind of can't understand the desire for that limitation...
I mean, I do. But it seems ridiculously arbitrary. Like driving a car in 2nd gear and complaining that it gets terrible mileage and can't go fast enough. The Drive gear is literally right there.
On that note, I actually had an overall harness (for experimenting) that was essentially like this: "for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer".
It actually worked incredibly well on all "gotcha" LLM questions like math or counting letters in words and all sorts of stuff.
Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend an infinite amount of money.
First -- most _people_ cannot do this, without having a physical board in front of them.
Second -- Claude Code is perfectly capable of downloading and running stockfish. People focus too much on LLMs by themselves as the entity of concern instead of the entire harness and all of it's capabilities together.
We don't really consider humans downloading stockfish to beat people at chess as noteworthy endeavors.
Which sounds a lot like what that paper was about
Zero days are valuable because they can be exploited but if the pace of exploitation is faster (which I'm not sure is the case), then the response WILL be faster, even if it means going offline. Institutions that won't will simply go offline by losing their data or becoming unprofitable due to ransomware.
Now for components that are core to the infrastructure, say OpenSSL, there is already a TON of attention and efforts, including red teaming, so it's not as if it's opening floodgates.
Sure low hanging fruits will get picked either faster or a at a larger scale, say a random outdated IoT device at your local flower shop, but for the rest, I don't think it's realistic to expect no response.
Security, digital or not, has always been an arm race. New threats means new responses specifically by incorporating the threat.
(Mind you, this may be for the better. I'm just saying that the safeguards driven by cybersecurity concerns aren't some new quality that wasn't there before.)
Frontier Labs will probably survive off hype valuations but will serve the important purpose of discovering architectures/techniques that will probably spread through rumors/transfers to the rest of the world.
Some people will complain about the wrong flavours, or missing flavours, or the price, the long lines or maybe it closes early on fridays.
Summarise means different things to different people.
> the frontier labs are priced according to the narrative that they have produced or will in the very near future produce a fully automated drop-in replacement for most knowledge workers...
Even assuming this is how the AI companies are being valued (they're not), the numbers are off.
The "value" of most knowledge workers -- based on what enterprises currently pay for them -- is $50 - 70 trillion annually. It's reasonable to assume that if AI drop-in-replaced all those knowledge workers, AI companies could credibly charge somewhere in that order of magnitude, because that's what the market is already bearing.
So if their hypothetical revenues are double-digit trillions and valuations are some multiple of that, the entire AI industry would be valued at double-digit trillions at the least.
Yet cumulatively the industry (the frontier labs + the SWAG estimate of the AI parts of all the other players) are valued at, say, ~6 - 7 trillion? Which seems like a fair approximation of how much knowledge work they can currently automate.
And then there’s the second order effect: if all the knowledge workers get automated, who is going to buy the stuff that’s produced?
Take the Hugging Face incident. Why did it happen? Because the people whose task was to set up a testing framework took shortcuts. Why did they? Because there weren't enough people who were assigned to do the job. Why not? Because the job is too new and not enough people are qualified to do it. It's a job that simply did not exist 3 years ago. But 3 years from now, this job might very well employ tens of thousands of high skill knowledge workers.
Unfortunately, I fear that may not be the most likely outcome. I've posted some comments on this before, but when I start thinking about how deeply everything will change once people figure out how to properly leverage AI, I see no outcome other than significant, widespread job losses.
As you indicated, at that point we will have much a bigger problem than the valuation of the AI industry. I'm not sure how it will get solved, I just know it will HAVE to be, because it would be an existential problem for everybody: people, governments, even the billionaires! Because now consider the 3rd order effects: if nobody can buy the stuff that's produced, how can billionaires get even richer? ;-)
What do you mean? The sum of ALL US salaries is $13.4 Trillion per year. According to google $65T is the sum of ALL salaries Globally (not just knowledge workers). It's not reasonable to assume AI is a drop-in-replacement for any job yet (perhaps bottom tier customer support from oversees?).
> So if their hypothetical revenues are double-digit trillions and valuations are some multiple of that
So you're sort of premising here than more than 16% or 1/6 of all the world's jobs get replaced by AI. Hopefully you can understand that's both not the current AI capability and also would be a terrible (unprecedented?) economic shock.
There's no strong evidence that openAI or anthropic will have non-linear revenue growth, so I'm not sure what point you're trying to make. Unless you think they'll fire all their engineers and replace them with agents or something crazy?
Unfortunately, I do fear that AI adoption will go beyond augmentation to automation, and I do fear an economic shock. Just posted this down-thread: https://news.ycombinator.com/item?id=49722616
Assuming a 33% boost takes it to ~20T annually, which is technically "10s of trillions of dollars" in revenue. And these numbers are from before agentic tasks arrived on the scene, so the actual productivity boost and corresponding value to employers is likely even higher.
I'm not sure where the disconnect is?
I'm not really going to belabor this point more, because it really sounds to me like you haven't even done the most cursory exploration into this and the people who have give an estimate of 7-10%[1] (and only a small fraction of that value would be captured by the AI provider). All the best.
[1] https://www.goldmansachs.com/insights/articles/generative-ai...
Toggle "Work hours using genAI" and "Time savings due to genAI" so see what I mean. You can also toggle between various industries, which is eye-opening because even industries like "Agriculture, Forestry, Fishing, and Hunting" are seeing productivity gains!
And if you do not want to read the detailed paper, they have a follow-up article here: https://www.stlouisfed.org/on-the-economy/2025/feb/impact-ge...
Specific quote:
> Using our data on generative AI use, this estimate implies that, on average, workers are 33% more productive in each hour that they use generative AI. This estimate is in line with the average estimated productivity gain from several randomized experiments on generative AI usage.
What we HAVE seen is that the national labor productivity has gone up by 1.3% since ChatGPT was released, and it lines up very well with all the other data and studies they cover... AND your reference, which predicted a 1.5% growth back in 2023!
Future supply and demand will set the price - not what is paid today. If supply by open models is vast and cheap, I can't see that the entire knowledge industry can hold the current size. It'll rather collapse to a fraction of its current value.
I don't think that AI companies can charge the same. The human workforce can charge these costs, because of scarcity. But AI systems won't be scarce, it's just a matter of who can run inference cheapest. Plus you still have the human workforce, which might be forced to offer their time for less money.
In one case (financial services) it's thought that expertise is valuable at V=S^2/b4 where V is value, S is skill and b capacity (the leverage available to the manager/expert. b erodes as it becomes harder to find examples of things that are not done well, so if you manage $1bn you might find lots of miss allocations that you can exploit with just that $1bn really effectively, but if you manage $10bn it's much harder to find good places for the extra $9bn. A low hanging fruit effect.
Anyway, that double hit - raw skill and the amount of times you can supply the skill makes the value of skill (V) convex, and it means that in a perfect market (heh heh heh) someone running $100bn is worth 1000's or maybe 10,000's of an average joe expert.
Now, if AI is trusted to run the top 0.1% of everything and has the skill to do it at human top level expertise, then your calc holds. If it's the case that it isn't then more than half of that value disappears. If it's not even top 1% then chop out another 25%.
That implies that we need a lot of trust and a lot of AI capability before these valuations stack up, and it also implies that all other competitors and incumbants are going away. I do not think that Citidal or Bridgewater are going to let Anthropic or OAI take them without a fight. They might lose - but there is a decent bet that they don't. I don't think that many professions like Lawyers or Doctors are just going to roll over and cede their monopoly rights to OAI or Anthropic either.
You need to think in terms of supply and demand.
The demand is there, but the supply is also going to skyrocket. Free open weights models will contribute to supply too.
There will be a new equilibrium that’s hard to predict.
That is the after the fact justification of the AGI dollar auction. Each round is kind of 3x the previous cost and neither can really stop because second place in the dollar auction is so much worse than winning.
The only way to stop the auction is one bidder hits a hard budget constraint, both agree to stop, or an outside party breaks the auction.
IMO this is why they want to slow down or have regulation. I think this is also why we see some claims of already reaching "AGI".
The TAM of global knowledge work is just a narrative tacked on after the fact to justify the AGI dollar auction.
The economic fallacy here with the actual valuation is akin to pricing the electric utilities 120+ years ago as some % of the future cash flow of global food production. Take the TAM of global food production and then work back to what % will the electric utilities capture from the advances in the automation of farming? It is nonsense.
The only narrative that actually justifies the capex spend that I can figure out is a first mover AGI monopoly. Even the oligopoly case is hard to justify the capex spend IMO. There is this enormous mismatch between the AGI monopoly and the actual rolling 12-month window of pricing power.
Even the rolling 12-month window of pricing power is going to saturate well before AGI too so it is hard to see how any of this makes economic sense.
FWIW even without AGI there are indications that all this CapEx spend, even with very shallow adoption, is boosting national labor productivity by 1.3%. That is worth ~$123B based on total wages paid in the US alone: https://news.ycombinator.com/item?id=49721338
As such I don't see the need for AGI for any of this to be financially viable. Whether it is economically and socially viable... that's where I have grave doubts.
But Capitalism really only focuses on the former and not the latter, which is why this will keep getting pushed forward. And that is the crux of the problem.
This assumes you don't change the market, but at the scale of (checks notes...) "all knowledge work", that just doesn't hold.
For example if you put 1bn people out of work, you now need some sort of safety net to bail out much of that workforce, a truly unprecedented change. You also lose tens of trillions of dollars of tax revenue.
One solution might be to recoup that cost and lost tax revenue from businesses by raising corporation tax. If corporation tax went from low tens of percent to high tens of percent, would those businesses be able to afford all that AI? No. Same order of magnitude? I doubt it.
There are many possible futures there, but the simplification made in the parent comment is completely unrealistic. The article is right in calling out the valuations as crazy.
However, I fear reality will be much hairier.
As models advance, we shift the goalpost for what "simplest task" means. Before, "simplest task " meant "write a coherent English sentence." Now, "simplest task" means autonomously fix, review, and merge a bugfix.
Or do i miss the point you are trying to do?
In particular, I found this very misleading or irrelevant:
a typical CPU project anecdotally has about three times as many specification and validation engineers as design engineers and a 5:1 ratio is not unheard of
The reason silicon design has such verification to design ratio is because the cost of one bug is many, many orders of magnitude higher than software. Both in dollar cost and in schedule cost (it takes months to fab a chip, and if you messed up and need to spin a fix, it costs tens of millions of dollars, not counting any design engineering cost).
I don't think you can extrapolate these very industry-specific facts to judging LLMs.
Aren't you just describing waterfall? That's still very prevalent in software engineering, and pretty much any other type of engineering – civil, chemical, building, architecture, drug discovery.
It's typically true that software can fail faster and cheaper, but it's also true that the costs are still vastly higher to fix later in the process.
Sure, there are some software that have similar "can't have bugs" requirements. I imagine the computers on Moon missions also had that kind of high bar. I wouldn't use NASA requirements as a proof for how LLMs should be used.
Depending on where you point them, they can be incredibly useful.
They can even be useful when you point them at each other (though increasingly difficult to get good results).
I'm excited for the promise of RSI and a future where models have inherently "live" weights, but it's not clear to me that the transformer is more than a useful tool to help us get there.
I have no idea how people can so confidently say that call center work is a “controlled environment” or “repetitive”. It’s almost by definition not repetitive or controlled. Customer support is what I go to when the controlled environment has failed
Typically it means knowledge retrieval from a KB or manipulating a control surface not visible to you.
This is only true if you are concerned about the intermediate steps of the model as opposed to the outcome. The huggingface hack was a perfect example of the model doing whatever it takes to accomplish the goal of maximizing its score.
No, they're really not.
They're priced in a way that would imply AI will be universal form of compute, alongside traditional deterministic systems - which it will be.
And that they will capture most of that ... which they won't.
The Frontier Labs are a very bad buy at a high price, but that partly has to do with wacky pricing, but actually mostly has to do with their relatively weak place in the value chain.
The money is going to Nvidia, who have the most powerful position.
A bit like how a retailer can take all the margins of some innovative product, if they own the channel.
AI is over-hyped, the Frontier Labs are over priced - but AI is here to stay, and will grow. Not like Skynet, but like a new form of compute. And it will take it's time, and the profits will be reaped by those with the power.
if that's true, you are wrong.
if that's false, anthropic is dishonest. why trust a dishonest company to be worth anything?
Like - the guy on TV talking about 'AI will destroy everything' ... I don't think he's lying.
I think they are like we here on HN and Reddit and a bit caught up in our own thoughts.
If AI were unleashed, in raw form today, it could cause havoc.
Bad. Maybe very bad but I think we'd get over it.
It would probably trigger a recession (because we are in a bubble - it would pop it), and people would 'blame the AI' for sure.
But it would be a bit dot-com ish kind of recession.
The amplifiers would be geopolitical instability.
What is "raw form?"
This logic doesn't follow at all.
If their argument is that there is 10% chance of extinction then they also believe there is a 90% chance it won't.
10% is uninsurable, priced in with ordinary treatment of risk it suggests that anthropic should be worth zero today. creating that risk would put every executive in jail.
on top of that it would demand under existing laws of conflict, a military campaign to destroy anthropic. that is not optional, it is demanded now to save lives.
hard to make comparisons but we mourned and rembered 9/11 recently. a 10% risk of hundreds of millions dead in 10 years would make anthropic a thousands of times greater threat than al qaeda. many countries would assassinate dario amodei and the leadership of anthropic now, within weeks or months.
actually just on the vague risk of having a nuclear weapon in 10 years, the USA killed ayatollah khamenei, his daughter, his son-in-law, his daughter-in-law and his 14 month old granddaughter. then, they killed over 120 children ages 6-12 by accidentally bombing a school.
in the sense that i would analyse a company, at least, the claim is false. it's not true that ai has a 10% chance of causing human extinction within 10 years.
they are making false claims about the technology they sell.
i have a fairly inflexible approach to that. sure, exaggerate but outright lies about the nature of the product don't work for me.
I agree with this.
The rest of it I don't. For external actions to be taken there needs to be a consensus and there clearly isn't that.
And this sort of thing happens all the time. For example Zuckerburg thought Facebook was worth more than the $1B Yahoo offered him and he was right even though no one else agreed. He though the metaverse was worth the $20B+ they spent on it and he was wrong.
There's a gap between people thinking something within a company and people external to the company agreeing with it.
I think the core mistake is this partial-equilibrium reasoning. Take the new technology and then hold everything else fixed.
$40 trillion of knowledge work routed unchanged through a new toll booth. Profit. This has nothing to do with reality.
Nvidia on the other hand does have the CUDA monopoly so their toll booth is printing money but that will get routed around or broken at some point.
LLMs are deterministic. They are chaotic, which people confuse for non-deterministic.
That's an odd argument, because a lot of people who have struggled to decipher complex chaotic systems would tell you this is a distinction without much of a difference.
The only gotcha with this is that they are theoretically deterministic, but rarely in practice.
A few examples:
- Harness specific settings that user can't control (anything from timestamp to prng seeding.
- Batching requests in a way that leads to a single request being processed different depending upon the batch (say MoE where your first choice expert is assigned to someone else's token so you go to your second choice vs a batch where you get your first choice).
- Graphics card itself carrying out floating point arithmetic in slightly different orders leading to floating point non associativity causing different outputs.
But all of these can be controlled for (at some cost) and the model can be ran deterministically.
For the average user, it might as well be non-deterministic, but when considering theoretical capabilities, chaotic deterministic system seems the better description.
The use of LLM in any system creates non-determinism in any practical sense.
That's a reason to be bearish about AI companies, not LLMs. But is it even true? OpenAI and Anthropic have each reported ~50 billion in revenue with ~900 billion valuations. That's a high ratio but I'm not sure if follows that the only way it pans out is if we get "fully automated drop-in replacement for most knowledge workers".
It wouldn't shock me to see those revenue numbers scaling up to where they need to be over the next decade ( to, say, ~400 billion) without ever achieving drop-in worker replacements.
These companies however are LOSING money (anthropic tries to make it sound like it's profit by deviating from accepted accounting principles) and subsidizing these models. When accounting for all the engineering salaries, training, GPUs, etc, what's their best-case realistic margin three years out, 10%?
So to we'd need a scenario where companies are spending a collective 300B annually on AI (believable) but ALSO that these companies jack up their margins WITHOUT companies switching to the cheaper open-source models (even when there's a $300B incentive to do so).
have you guys actually designed, built, and deployed agentic workflows?
it is actually quite hard, requires tons of time spent on evals and testing to ensure accuracy, but when it starts to work it is mind blowing.
there is no going back.
listening to people yap about AI when they have only surface level or one dimensional exposure to LLMs and "AI", but have not actually put innovations to work IN PRACTICE.. is a waste of time
Please share some of these insane things that you speak of..
use your imagination to solve problems people face and pay $$$ for today that is error prone and hard.
i’ve got agentic workflows for the particular industry im building for, one of which that replaces the need to hire $500+/hr services.
in this particular workflow (don’t want to reveal too much, sorry this is my competitive advantage but you can figure it out for your own workflows) a $4/1M model ingests a file that is currently used in a extremely complicated program that few people understand how to use.
it parses the data, loads it into a database, and then spawns a bunch of other agents that check the data against work in flight. there’s checks for bad data. in that case, more agents are spawned that reach out to the involved people or parties for clarification. if it cannot figure something out it reaches out to the right contacts for more information. while this is happening, more agents begin doing work that involves continuous reconciliation against 100s or 1000s or more things in flight.
as files are uploaded, or updates from people come in, agents do work to ensure things remain on track.
people are able to work across languages and cultures, and my agents ensure that while people can make mistakes, it will catch them in real time and ensure continuously monitor the situation.
it’s pretty nuts how much inefficiency agents today can solve. it takes patience to run tests and tweak shit until it works.
*** the really cool thing is that more capable agents can continuously monitor how things are going and improve the workflow itself… so all i need to do is maintain the actual tests. ****
i loved writing tests back in the day to ensure i built good software. today we write tests to ensure the business can run.
I could not think of a worse technology to use for an ETL pipeline than throwing LLMs at it and asking it to vibe out the correctness of the data every time it runs.
98% of people think AI means what gemini tells them when they do a google search.
of that 2% who go beyond... maybe 20% of those are using AI to code.
most SWE still think "using AI" to code means the copilot pane they open on the side in VSC. they get sloppy code output and think "ai sucks!"
so of that 20%, maybe 5% have actually explored what an "agentic" workflow even means. they might use skills, set up their code base so AI almost NEVER makes a mistake. maybe 1-5% of that 5% actually went deeper, and those are the people who built tools like Cursor or Harvey or whatever other "agentic" companies.
you are still thinking about the old world where you obsess and define data models and bike shed over data integrity, all the while you have a "temporary" table with 3 attributes that gets 2 million queries per second that's now holding up a bunch of other shit that's also glued together.
> I could not think of a worse technology to use for an ETL pipeline than throwing LLMs at it and asking it to vibe out the correctness of the data every time it runs.
like i said, if you are still having quality issues in 2026, that's a skill gap.
also, agents now continually improve the process.
for a business, the only thing that matters is transaction log. for 99% of businesses, swe are a cost.
again this is like baby steps on the journey. we are still only 3 years into this technology being opened up to masses.
i'm sure people thought computers were dumb, or that cars are stupid because the first cars were moving slow af. "we have horses, why do we need to build out roads to get anywhere"
but you're free to feel smart doing 20/20 hindsight on things 100+ years from the future
So what ever shit the "AI" gets wrong, it is the user's fault, right?
If there is no such guarantee, why should I spend time writing an elaborate skill file? There is no telling when the model chose to ignore stuff in it.
I don't understand how people can work with something like that!
when you have chains of agents working seamlessly, where agents are managed by themselves (like spawning a new chain to do some scope of work), and it just sort of works on a infinite game loop, your system continually keeps improving as it works. at least, that is the goal with the systems i like to build.
you have to do your due diligence, obviously, as you would with any TOOL you use. in that case it means having specific goals, methodologies, etc. that each model has to follow. (if one is a "hey grok explain this" type of person when it comes to using ai, ngmi)
we have one agentic workflow where there are 4 different models that can be spawned (to handle tasks of various complexity), that work against an API (the source of truth). their goal is to work any time a specific file is uploaded, and handle it. it is a complicated file, with 1000s of line items, with varying amounts of uncertainty involved.
it happens in the real world today, where it costs $xxx,xxx per year to do. because it involves many people and companies... it is done totally manually today..
how it works in order to get to that resolution is different based on the complexity of the task... sometimes there's bad data in the mix because people make mistakes when they create stuff (we're not special). in this case it requires finding out whether it is indeed bad data or actually correct. that requires setting up a scheduled job, and handling it when there is a resolution.
i don't code this part. how this gets accomplished... the agents are able to work autonomously to handle any edge case, in order to complete the goal. the goal in this case is to ultimately record transactions (a goal of a business is to make money, believe it or not).
does this not make sense? jeez hacker news used to be an imaginative place.
the model that we pay $10-20 per 1M tokens will be $1 per 1M token next year.
And the model next year that we can pay $10-20 per 1M token will be even better than Astra and Terra/Sol/Luna. It is an exciting time.
The non-sense is the part where you think putting agents in a loop is going to yield better results indefinitely...
life of a w2
> of that 2% who go beyond... maybe 20% of those are using AI to code.
> so of that 20%, maybe 5% have... maybe 1-5% of that 5% actually went deeper
How can you write any of this when you said you just started liking LLMs with Astra, a model that released two weeks ago. You're whipping out a bunch of made up stats trying to describe large swaths of programmers while admittedly being unfamiliar with the capabilities until two weeks ago? Why are you so antagonistic towards people's skepticism if you used them for years and found no value in them? This sounds more like a comment trying to generate massive amounts of FOMO. "If you REALLY put a lot of your money into them, that's when they shine! Don't believe the haters!"
> like i said, if you are still having quality issues in 2026, that's a skill gap.
> also, agents now continually improve the process.
> for a business, the only thing that matters is transaction log. for 99% of businesses, swe are a cost.
Maybe you should try building something for more than two weeks before making such big declarations, because this really does read like someone who is excited about a novel they wrote last night at 3 am
been building stuff for years, so i'm not just some w2/1099 type guy who bike sheds for a living lol
this stuff is new... obviously it is a fast moving space.
already been through one IPO where i was an early employee so i’ve never really had to work with morons.
> the present problem of reward hacking can be solved only by rigorous specification by domain experts. the time of domain experts is expensive. rigorous specification is itself a skill, demanding its own expertise outside of a given problem domain. even many skilled software engineers are bad at it. for the vast majority of domains, the intersection of domain experts and specification experts is ludicrously small.
- i have no idea how anyone thinks the mighty next token predictor is going to eradicate diseases and eliminate poverty https://blog.florianherrengt.com/how-llms-work.html
- i also have no idea what everyone and their momma on HN is running for more than 5 mins in the name of "agentic AI"
It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on a specific problem?
IMO the very best case scenario / potential for these are likely better than we think, but right now hidden due to logistical and financial reasons.
But if we assume that the model costs will continue to drop by a factor of 5-10 annually, there will always be a latency of a couple of years between what is completely out of reach, and what is financially viable.
Basically: If you knew AI could be affordable enough in 3-5 years so that even the most underfunded researchers could use it to solve cancer, how much would you value it now?
AI agents are good at solving well-specified tasks, not at solving problems. They do well in fields where the cost/effort of specification is already part of the business.
Solving cancer also has an unusual level of specification. Many real world problems have that characteristic.
Where did you read that?
"Cancer" is not just a single disease, even though we layman use the term that way. Cancer is a family of diseases, each probably having their own specific solution, but even in each of these individual diseases, there is no specification at the level of any open maths problem.
I really meant solving a specific cancer disease for a specific individual, which will require individualized medicine, which will require us to leverage AI to make it possible to do for the general population as opposed to doing it just for Lance Armstrong and the like.
I know I'm an optimist, but I really think AI is gonna result in drastically improved healthcare for a much lower price.
The real world is so messy, and specifically cancer/biology is insanely messy and certainly not well specified.
You should close chatgpt and read a book sometime.
That's the thing, it very much does NOT show us that. What happened was mathematicians at openAI learned of an imminent development on this problem, and the insight that it entailed, then they were able to prompt a system in the correct direction and spend 20 million dollars to write down the final steps.
Which is rather precisely the point that the article is making!
> If you knew AI could be affordable enough in 3-5 years so that even the most underfunded researchers could use it to solve cancer
As the saying goes, if my grandmother had wheels she would have been a truck.
That's not what happened.
> It also makes one wonder: What could AI solve if we managed to orchestrate billions worth of agents to take on a specific problem?
We need to have robotics automation catchup first. The math and coding problems are problems in written-space only: you can set up feedback loops to test what worked and what didn't, then try to resolve the defects, maybe back up and try a different path, etc.
What solved coding and maths problems weren't the damn models; open up a chat interface to a SOTA model and you'll see they are pretty limited in producing a solution without a feedback loop.
Instead, it was the harness around the models: it let them explore a space and use feedback to control and direct that exploration.
Until we can do it in meatspace, it's kinda pointless sinking a ton of money into large problems facing mankind...
Like establishing a colony on mars (so the next rock to hit earth isn't an ELE).
Or moving us to a post-scarcity utopia, ending the concept of money.
Or designing and building better batteries for transport that uses only electricity (so that we stop using fossils as fuel).
Or actually building mass-housing. Or mass-farming. Or both, potentially ending homelessness and starvation.
Those are all worthwhile problems to solve, but where's the point of getting a solution on paper? There's no exploratory mechanism there, even for humans, to come up with a solution.
So, all we are left with then is making knowledge workers obsolete: another ELE, but of a different, self-inflicted kind.
Personally, I use LLMs for a lot of things. Oftentimes, I'm a think out loud type of person so even having something that feels like a rubber duck, but more competent, is already amazing for me. And LLMs are a lot more competent than a rubber duck.
But especially sometimes I've noticed that LLMs can be unbelievably stupid. It recently happened a few times with Fable 5.1 as well. Ultimately, I think it comes down to that LLMs can't think broadly. In software development one can usually see this too. For example, a whole app might be built by an LLM and it didn't spend a single token thinking about security because the prompter is at the level of "build a dating app for dogs, make no mistakes". Now you have a dating app for dogs that is insecure.
Since I prompt for almost everything in my life to have an LLM as a sounding board, I'm usually not an expert either. I've noticed LLMs are amazing at "bulk search engine information aggregation" (or whatever you want to call it). So if I need something from the Dutch government, I can find it way more quickly. But oftentimes I've noticed that going for a walk and thinking about a particular thing I'm facing is a more effective way of finding a good solution.
Other times times they are not incredibly stupid, but can't form a strong opinion. This usually happens when I'm tackling a wicked problem [1]. When that's the case, prepare for LLMs to sway with you for every small change in your opinion that you ever will experience.
So I agree: drop in replacement for knowledge workers? No. Rigorous specification is usually needed yes. Though, the small win here is that it doesn't always need to be as rigorous as programming is and it can happen in natural language. It depends on the topic/problem being tackled.
I really like them as UX tools though. Amazing for interactive prototyping and requirements elicitation. And that also corresponds with what the author is saying. Though I find it a bit of a disservice saying "just 3". You know how hard requirements elicitation is? It became a whole lot easier thanks to LLMs (I might change this opinion in a year, haha, but this is the opinion I hold now).
Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.
Array.from(document.body.querySelectorAll('p,li')).filter(e=>e.innerText).map(e=>e.innerText = e.innerText.split('\. ').map(s=>s[0].toUpperCase() + s.slice(1)).join('. ') )
Not perfect but hope it helps.LC;DR :P
If it had proper caps etc, people here would accuse it of written using LLMs.
You just can't win...
That’s life )
It's definitely... unique?
Author's own style is certainly refreshing and welcome over LLM slop that dominates most HN posts now.
No it's not. It's common in chats where those people don't bother with punctuation at all, often writing a single sentence over multiple lines.
It's not common for blogs or long form content.
I read it as "I'll take literally any conscience for myself no matter how minor, at any cost for you no matter how big".
When you intend it as "inviting informality", you're implicitly doing this because "formality" is too much effort. The thing is you're not inviting, but rather demanding. You have decided the conversation is informal and low effort, and that's how you'll treat it, without considering the person you're communicating with.
This is of course all a lot of strong statements and these things don't matter as much as this sounds. I don't feel _that_ strongly about these things, but the lowercase thing always strikes me as just plain weird, and deconstructing why I feel that way this is where I get to.
So typing in all lowercase is a choice. I think it's a way to signal "I don't follow norms". It signals membership in a group.
Not my read on OP, but “tech bro” is not my first thought associated with the style
didn't realize people struggled to read text that way though, maybe i should change my writing style if this is a common pain point
If you do everything else right you can skip sentence cap without sacrificing readability. Not that you necessarily should but it's a vibe if you want it, it suits some voices & documents well.
(So does not adding a period at the end of a sentence)
Yeah, it does
I will add that I'm not very happy with the readability of my site overall at the moment; if anyone has font or other recommendations for style tweaks to make I'd love to hear them!
To expand a bit: It's clear from the writing that the formatting is an intentional stylistic choice. My point wasn't about aesthetic preference. I meant precisely what I wrote. Language evolves and orthography evolves. We've been capitalizing less and less for centuries now. But (until very recently) it was a universal rule to capitalize the first letters of sentences. I believe this is partly because (again, until very recently) people read in large quantities and, to a fluent reader, sentence capitalization serves an important purpose: it helps the eye recognize where one sentence ends and the next begins. Periods alone can be easy to miss, or confuse for commas, when the eye is moving quickly.
Anyway, this is all a very minor point. You have a cool site and I enjoyed your article. I hope you keep writing.
I'll play!
My recommendation, in short: pick a new font, make the content pane narrower (around 75 characters per line), and increase line height by 15%.
The font you use, Montserrat, has nice details, but they're extremely subtle, and our eye doesn't pick them up at text size.The font uses very pure geometry [1], has really big counters [2], is light in weight, and has no visible stroke contrast [3]. The effect when you look at the screen is that you see big blocks of text, but struggle to register individual characters, and it's hard to find the beginning of the next line. Add the lack of capitalization, and it feels like a monologue, like you're playing rubber-duck for someone.
A challenge: everybody uses Google Fonts, and all the good fonts get used so heavily they end up losing their ability to make people feel something.
Pick a font that is good [4] and feels comfortable to you. Ideally, buy one. You get what you pay for. And you go from being one of the millions of people who use a certain font to being one of fifteen.
If you like the feel of Montserrat, Tiny Grotesk [5] and Decimal [6] would both be great choices and are from excellent type designers. Or browse Matthew Butterick's font recommendations [4]. They're good. You're out $50, but it's yours.
If you're absolutely allergic to paying for fonts, DM Sans and Work Sans are Google Fonts, have a similar feel to Montserrat, and solve the above problems. But then you're a robot.
A final note: Montserrat is actually a nice headline font. The details come into focus at larger sizes and weights. So I'm going to open a can of worms: you could pick a _contrasting_ body font. Remember PT Serif from the [3] footnote? Try it on for size.
----
[1] By that, I mean things like: an "o" looks like a perfect circle rather than an oblique oval, and forms are strictly on 90º axes.
[2] Counters are the apertures of letters, the open center of an "o" being one.
[3] This is the contrast in "line" size. Think of it like the width of the line when you write with a chisel-tip marker: some lines end up thinner, some thicker. Look at the "e" on the font PT Serif - the vertical walls of the character are thicker, horizontal thinner. https://fonts.google.com/specimen/PT+Serif?categoryFilters=S...
[4] You can't go wrong with anything Matthew Butterick recommends (https://practicaltypography.com/font-recommendations.html), or from any of the "big" font foundries: Hoefler & Co., Linotype, Monotype, Berthold, URW. If it's a bestseller on myfonts, it's a sure choice: https://www.myfonts.com/collections/best-seller
[5] https://tinytype.co/type/tiny-grotesk
[6] https://www.myfonts.com/collections/decimal-font-hoefler-and...
If what you're trying to communicate is that, like Sam Altman, you're too cool the press the shift key, then why write some initialisms in upper case? After all, there is a style of acronyms that are always written in lower case: Latin ones ("i.e.", "e.g.")
Some concrete suggestions from me:
- Fira Sans [2] is a lovely, legible sans-serif font that's free and incredibly easy-to-read. This would make a fine body font.
- Source Sans [3] is another good sans-serif that might be closer to your current font. It's about as readable as Fira, but it's a little less warm imo.
- Make your line length a little narrower and your line spacing (i.e. leading) just a hair wider—that will do wonders for readability.
Definitely check out Practical Typography for more tips from a real professional.
Thanks for the blog post!
[1]: https://practicaltypography.com/
Since you are willing to consult LLMs for reviewing your post, why not ask them for feedback on improving the CSS, etc?
I won't comment on hiring interns as that's not my expertise (even though if you want to teach your staff, obviously I can see a problem there) but I can comment on rapid prototyping, it's what I do.
Rapid prototyping is NOT making a CMS quick. It's not about making a quick mockup of a UI. It's not about making yet another well known... anything. The entire POINT of prototyping is to make something NEVER done before. Typically that means you are reaching the frontier. You are making something with NO documentation to rely on. You are using tools, hardware or software, which do NOT have tons of StackOverflow errors. There is no dataset to crawl, there is no well structured Q&A database to train on. You have to poke and see if the thing actually works as expected, and it often does not.
So sure, if you are using interns as a trick to underpay your staff, or if you are using prototyping as an excuse to build poor quality software fast, maybe it does help. If you are genuinely prototyping, it breaks fast and the supervision overhead makes it pretty pointless, especially since typically it's by actually implementing that you find out not just how the new setup works, but also its limits, and thus the actual needs of the project, not the one the stakeholder imagined would be.
So not, not for rapid prototyping either.
TL;DR: prototyping is a learning process, not a low fidelity output.
PS: this comes up very often from NON prototypists that I wrote a short piece about it https://fabien.benetou.fr/Content/GoodPrototypesAre10LinesLo... so much so that it feels like a pattern "GenAI/LLMs is good for tasks X" while the author actually does not do task X except very superficially.
Putting that aside, prototype software is recombining existing technologies and concepts in well trodden domains, which is distinct from the genuinely novel scientific work the author was contrasting with. Software prototypes are not in the same league, as much as you may like it to be.
That being said I didn't compare both, not sure why you brought that up. I specifically discussed about prototyping, quoting a specific sentence, not scientific research.
I don't know how well it is studied, but I suspect it is possible there is language-related complexity constraints to the effectiveness of the LLM algorithms. Like perhaps context-free grammar related problems with adequate training data can be more and more effectively solved, but maybe natural language related problems will not so much be effectively solved.
Would be curious if there is active research here.
I believe that's obvious - humans don't think in words. Neither do animals. A machine that only thinks in words is obviously going to be deficient in some things, no matter how proficient it is in everything else.
Where did you read that?
The people you are talking to have apparently never seen this research. Maybe you can help them out and provide links? It's faster to do that than to engage in some form of audience-persuasion, as well as more effective.
Lots of people make this claim about "specific task[s] enjoying clearly defined levels of task performance" but they forget that generative AI is also extremely good at generating a) art and b) prose in literary style. None of those things has "clearly defined levels of task performance", in fact they are both the complete opposite of well-defined tasks. Who knows what counts for "good" art? [1]
For me the right model for generative AI is "a million monkeys on typewriters" [2]. Holding any other model to heart will at some point fail to predict observations and cause you to be unpleasantly surprised. Not least because AI companies are actively engineering their systems to optimise for this model and they have a lot of people working on that engineering and shedloads of money to throw at it.
Don't underestimate what a million monkeys on typewriters can do. They can do anything and everything, given enough time. Geneartive AI can also do anything and everything given enough resources. The only question is: how much is going to be "enough"?
____________________
[1] Yes yes, AI art tends to be slop. Not denying that. But part of the problem with slop is that it presents as technically very competent except that it lacks a certain je-ne-sais-quoi, which makes it good art; aesthetics. The point is that there is no clear measure of what makes technically competent art, any more than there is for aesthetics.
And yet generative AI is very good at it.
[2] There's even an article on wikipedia except it's about one monkey on one typewriter with infinite time. There's a proof too.
I think that's incorrect from the investment point of view. They'd still be worth a lot if they can produce a drop-in replacement but it takes five or ten years as long as they dominate that. The danger from an investment point of view is they become AltaVista, replaced by some Google that does the job better.
> the models generalize well only on tasks within a small neighborhood of the specific tasks they've been trained on
Unless frontier labs have surprisingly trained their models in the exact tasks my team works on, this is patently false. We are getting very good results on automation and I'm bullish we will be able to mostly remove humans in the loop for most of our infra tasks by the end of the year.
I have no opinion on the other theses, but given that OP doesn't back up these claims in any way, I have my doubts about the conclusions of this article.
It have been demonstrated that in-context learning is a very powerful mechanism. There's no evidence that models of the size of GPT-6 are bad at in-context learning. In fact, ARC-AGI-3 score might indicate they are good at it.
There's no evidence that a bespoke RL environment is required for each new skill - quite likely a good demonstration is sufficient.
Take a customer service person, that as soon as AI agents replaced was hacked. Lots of tacit knowledge, that wasn't measured, or even probably in the job description, until AI agents had none and the gap was taken advantage of. Gap is probably a poor word here as it implies not a chasm, which could very well be the case.
Your analysis was excellent but short on one front, AI has endurance on it's side. Looking at the N-S solution, OpenAI had 10,000+ instances that kept trying around the clock. Assembling a human team to do that would require a lot of effort. Though, that society collectively choose not to, perhaps tells how valuable it really is (i.e. it's now easy to launch a Manhattan Project level of effort). So maybe it can also be said that AI is also good at marshaling resources.
Is this true? Could you cite an example of a simple prompt that GPT 6 / Fable 5.1 consistently bungle?
It will probably do a pretty good job if you ask it that way (it will also burn a shit ton of tokens, but hey, that is part of the fun).
On that note, I actually had an overall harness (for experimenting) that was essentially like this: "for any task, instead of answering question directly, write a program to answer the question instead. test and verify the program before giving the answer".
It actually worked incredibly well on all "gotcha" LLM questions like math or counting letters in words and all sorts of stuff.
Of course it was ridiculously slow and very expensive but it was a proof of concept that it can actually be much more accurate on every task if you are willing to spend infinite amount of money.
> the classes of firms that can accept the use of fully autonomous LLMs are few, by my count just three:
1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc.
2. those who need done a small set of narrowly defined tasks with existing clear guardrails: repetitive physical labor in a controlled environment, call center and customer service chat work, etc.
3. those that can accept or already do by nature the costs of rigorous specification and validation: chip design, drug discovery, and other domains where failure on deployment is an existential concern.
the first two classes are price sensitive and arguably don't need the jump in reasoning quality you see going from cheap to frontier models. most of these firms will be best served by open models running on cheap hardware, perhaps even locally at the site of use. for the first and third classes, the type of fuzzy combinatorial search that has produced headline results in mathematics and security research seems more sensitive to agentic swarm width than reasoning capacity
…
This is just so on point. And for the third class (which I would extend to things like materials research as well), specification and validation are already by FAR the larger costs, so automating search and simulation is really not a massive game changer for the broader business.
The counter argument is of course maybe you don't need to understand our kind of intelligence to create a different kind and that could well be true but then how do you determine if a system is intelligent.
Unless the new system is intelligent enough to reason with us on our level in a way we can "see" is intelligent it becomes a philosophical argument.
We also have a natural inclination towards anthropomorphising systems that mirror us, this is already a problem with LLM's and people overestimating their capabilities or forming actual emotional attachment.
Then there are those of us who know more about how they work who in theory should be more immune to that and aren't.
I added some stuff to my agent.md to make it sound less human and to communicate more like the machine be it is because I find the faked emotion extremely jarring.
It can't be sorry, it's a set of numbers, it sits in the linguistic uncanny valley.
A priori I'm not sure why you would think being a PM at a FAANG, deciding what color the login button should be, is any different.
We will never prove AI is intelligence.
We'll only prove humans are not.
> the history of artificial intelligence research is littered with examples of humans confidently declaring that task X requires general intelligence, then getting humiliated by a neural network doing task X better than humans a few years later
True, but the history of AI research is also littered with AI researchers confidently predicting X job will be replaced by AI and being completely wrong because they don't actually understand what those jobs actually are. See Geoff Hinton predicting that Radiologists would be obsolete by ~2020, or predictions that truck drivers would all be replaced by self-driving tech.
When it comes to judgment calls for technical decisions, a lot of interesting innovations appear to come from rejecting conventions / averages in favor of a different set of constraints as a trade-off because we challenge the assumptions we make about the demands being asked of a solution / product. I'm thinking in the constellation of the apocryphal Steve Jobs quote about rejecting asking horse riders what they want because if we asked them they'd ask for a more reliable horse.
But humans still manage to wrangle these, sometimes seemingly effortlessly, through a process which we call by shorthand "taste". This is a largely vibes-based heuristic that combines expertise with life experience and cultural training -- intuition, more or less.
This is likely not possible to automate either -- aspects of it may be automatable for a given expert, in small pieces in narrow subsets of their particular domains of interest, but even those likely will require some manual intervention.
This is in part because it is, to a large degree, a black box, even to the expert deploying it. With some self-awareness and strong language skills we can articulate approximations of the judgements that go into taste. But even those will fall short, as even the most self-aware individual will fail to notice certain judgements and dependencies.
In practice many of these are not even explicitly articulable. Humans are idiosyncratic and messy and dynamic, and the suggestion that we can build a machine that approximates this in a way that pleases our sensibilities and doesn't require supervision is kind of ludicrous, even in light of recent developments.
It feels like today, taste in design is similar to where software engineering was about a year ago.
I have my own theory about why it's impossible to remove the human from the loop:
1. Any task emerges from a need, from a human context. We need the human to pay and assume the risks and costs of using the model. So intent emerges from context.
2. While the task is being worked on, constant interaction with the context is needed, for action, for feedback, and for steering.
3. At the end of a task, consequences accumulate in the context, they don't fly to the model provider. The cost, risk, liability, gains and losses remain there.
So the LLM is great except for the start, middle and end of a task. Contexts are humans, teams, projects, and they are maximally distributed, you can't copy a context, it is indexical and relational, just as you can't copy my phone number or eat for me.
Though it is not hard to imagine that any and all communications being recorded for AI consumption in the future.
The Next one is the relative lack of prompt feedback (expect the blowup in finite time like Navier-Stokes ;) [there is not much feedback even for humans at middle management positions].
The cost [tokens] might become prohibitive unless LLMs improve further [not a guarantee].
You can make llms perform judgement, and maybe that will get you some progress. But ultimately the value is going to come from engaging with other humans.
LLMs are great at helping with aspects of market research and that’s about it. Aka it’s a good deep search engine. It’s not going to decide what features solve certain customer pain points. It’s certainly not going to prioritize and coordinate between competing stakeholders.
That's like 1/4 of Codex users burning through their free resets just to build shinier todo apps.
If the models get good enough to solve all of the low-skill problems, then we only need to keep around the people who are highly skilled. That means a fewer number of software engineers, and a difficult path to becoming someone who is highly skilled.
Agreed about the barrier to entry raising though, we're already seeing that in the glut of CS grads who can't actually get a programming job right now. Personally I suspect that this is a cultural issue more than an economic one, though. Companies need to alter their expectation about entry-level engineers and develop a culture of mentorship/apprenticeship so that advanced analysis, architectural, code review, and AI management skills are all passed down successfully.
Programmer / engineer compensation in the US market at least has been looking bimodal for at least 12 years now and the first one looks even more devastated in terms of labor than the big tech companies based upon (lack of) job postings and from browsing my connections on LinkedIn.
Test code looks like a mess, even more than the usual LLM code. takes a while to let it go. Test report looks beautiful though.
Like you, I've found that LLMs can improve test coverage by decreasing the amount of developer time spent writing tests. But generally, it takes a lot of manual work to set up the initial testing framework, and even then, a lot of vigilance to ensure that what is actually tested corresponds to the description of the test.
But that's digression. The short of it is I've seen enough to convince me these tools may as well be magic and a whole lot of tasks that break down into "produce media content of some sort" that has a well-defined goal and definition of correctness will be permanently sped up by automation. This includes a lot of software writing. At the same time, I shared the skepticism of estimates of economic impact and irrevocably changing the larger world. I'm a lot closer to the business side of the house these days, working with customers and prospective customers to identify use cases, reference architectures, pain points, feature requests, and bring this back to the development teams to attempt using real-world experience like this to inform how we design products. It's not product management as I'm focused more often on the nitty gritty technical details, not high-level user experience or roadmaps. But it gives me a great avenue into seeing what causes organizations to actually buy and/or adopt new software products, and the rate at which they can do that.
And frankly, it isn't moving the needle much. They have the same budgets they always had, so they're not buying more, and our business is growing, but no faster or better than it grew before agentic coding became a thing. I always wonder because it seems the glowing success stories on Hacker News come in one of three varieties. It's the solo indie dev, usually targeting mobile app stores, who churns out dozens of roughly "will compile and doesn't immediately crash at runtime" apps in the time it used to take to complete one. It's the hobbyist, making software only they and maybe their immediate friends will ever use. Or it's startups, whose monetization model isn't monetization at all; it's just having something shiny to show investors in order to convince those with loose money to give enough to you personally that you can build up a nest egg whether or not your product ultimately ends up ever having a single paying customer.
In my own business, a multi-decade, mature but not hyperscale company selling overwhelmingly self-hosted enterprise open source software, I can see the impacts on output. We have the same major products with the same release cadence. Each point release averages more new features than they used to, but also more regressions. It's overall a mixed bag. Non-technical product management staff is able to contribute code. We have a ton of new internal tools that nobody uses but they're there now. On the customer side, those that hinge decisions on wanting features that didn't exist yet are benefiting from getting those. Those that already had the features they want are losing from the greater rate at which regressions get through. The net business impact seems to be things have definitely changed qualitatively, but in purely financial quantitative terms, things are about the same as they were before. More code being committed to various git forges, but same headcount, same revenue, same margins, and same market cap.
The problem with this angle is that it is still absolutely terrible at doing call center/customer service work, and the profitability story is that the price is going to go up rather than go down.
For repetitive physical labor in a controlled environment I'm slightly more bullish, but if you control the environment, you mostly don't need AI. You just use traditional deterministic methods, and send a person in when things get stuck or things are by nature irregular.
> 1. those who can accept failure cheaply: firms that would otherwise hire interns, firms involved in rapid prototyping work, etc.
Those who can accept failure cheaply can't necessarily detect failure cheaply. A ton of insane attempts will have to be picked through carefully to find the candidates for success, because the lack of a thought process makes AI bad in random, inhuman ways. This is basically a version of 3) that wishes away tests. It will be (and is) certainly helpful to replace interns and aid in rapid prototyping, but not because failure can be accepted, but because those are things that are tightly supervised. According to the world thus far, that is resulting in anything from -15% to +25% productivity gains. I'm not seeing it as a game changer simply because if it was, I'd expect to have seen a lot more useful, original software products by now and I haven't. I've just seen old ones get buggier or rewritten in Rust.
I'm only buying 3): when you just want a machine to randomly enumerate through a search space looking for things that make the carefully constructed tests pass. That's a very good thing, though. But as you say, it's not a game changer because you still have to write the tests.
but
> 5. navier-stokes and statements in pure mathematics like it are the absolute best case scenario for agentic work against rigorous specification. the theorem statement itself is already a rigorous specification. it has undergone decades of auditing by the mathematical community and its rendering in lean is a straightforward translation defined in terms of battle-tested mathematical objects from mathlib. the verifier, the lean theorem prover, has been extensively audited and specifically designed to avoid the types of unsoundness that would make it vulnerable to reward hacks. even lean and theorem provers like it are not invulnerable: soundness bugs have allowed LLMs to launder bogus proofs through the proof kernel before and it is not improbable that more such bugs exist. this is the rosiest setup; the vast majority of human knowledge work does not look like this. i'll comment below on the few areas of knowledge work that do resemble pure mathematics in this respect.
This is the real deep point, and one I've been repeating since I heard Navier-Stokes was a fraud.
This is exactly where I expected that LLMs would do well, and they are not.
It shows that I have a basic misunderstanding of LLMs, and that misunderstanding is causing me to think that they have more potential than they have actually shown.
Maybe the nature of the architecture, where it picks out features, intrinsically limits its ability to search a solution space?
Maybe the fact that they modally predict what someone might say, and nobody has said a thing as of yet (when many people were knowledgeable enough to have, if it is correct), means that the LLM is not going to say it either?
Maybe the fact that it consumes all information and blends it in a structured way, instead of synthesizing an entire space from a relatively very small amount of input like a human does, means that it won't ever accidentally synthesize something that can't be pieced together from things that have already been said? Is its accuracy its flaw, where a human's "mistaken" synthesis might ultimately correct everyone's understanding?
Really not beating the charge of being a stochastic parrot. It might just be that we were underestimating stochastic parrots; if a million monkeys on a million typewriters were all getting treats when they satisfied a trainer who wanted to see a new work of Shakespeare; they could look at his published work, and they could watch each other type and when each other got treats; whenever they successfully spelled a word or put words into an intelligible phrase, that was made into a keyboard key for a group of sentence monkeys, and the successes of the sentence monkeys were made into keys for the paragraph monkeys, etc... could you get something that passed for mediocre, drunken Shakespeare in a thousand years? Or maybe even 10?
It's a shame that we've had decades of shitty customer service from companies that have the most responsibility and resources to do it, and people in tech have shrugged their shoulders saying "it's unreasonable to expect people to scale up their customer service! They gotta make money!" Now AI customer service has arrived and things are worse, not better. Now we can't even get in contact with anyone because an agent redirects us to documentation instead of directing us to a human being.
The future, when a company steals your money, is making viral posts online trying to get enough public shame going that you will maybe get some of that money back.
But, when properly used AI can reduce the customer service load. It can handle a lot of simple questions and status updates. There just needs to be a human available when there's any request that escalates past that simple case.
- Ximm's Law
Here's another human bottleneck: "AI Has a Discovery Problem" (https://news.ycombinator.com/item?id=49621223)
I did a back-of-the-envelope calculation, using the figured from their public release, and arrived at approximately $1M.
Something like 300Bn tokens, something in the vicinity of USD 10M.