For there is always a time to panic, no matter how exhausting the requirement might seem. And that time is now.
I told him that if I saw that Jesus was alive, walking amongst us, performing actual miracles - I would definitely try to convert him. I'm not one to look away from hard data just because it's inconvenient.
That's exactly what's going on right now. People's reactions to it are so weird.
https://www.mcsweeneys.net/articles/we-must-create-the-shit-...
In all seriousness though, this is a very good primer on all the reasons I think that people who claim “the singularity has already started” are wrong. I urge anyone who’s panicking about AI (really?) to read this and consider it:
But then I'd expect that such people would also have a much more tempered expectation about any AI apocalypse scenario, or indeed, the shape of the curve we are on.
The first link however will be surprisingly delightful.
I have “adopted LLMs and agentic workflows for work” and I neither had an aneurysm nor felt that my time was wasted after reading it, and I think it showed the very tempered expectations that you reference.
Care to actually engage with the points in the article?
Nevertheless, point by point then:
1. price competitiveness and autonomy:
They falsely assert that "current frontier models need laborious oversight and guardrails on even the simplest tasks" - this is simply blatantly untrue. We have countless workloads at work that are easily doable for an agent, and they're not really trivial: bootstrapping entire new environments, migrating existing ones, investigating incidents and crafting RCAs, and so on. I experience this time and time again, and have been for half a year now. There are lots of colleagues of mine who only do such tasks. They do a worse job at them (sometimes way more), and take longer to do so. Using guardrails is basically unneeded, and the oversight burden is also far from laborious.
I also work with those "crappy engineers" they speak of, holding titles inflated way beyond "cracked interns", and I wish it was only a bottom quartile. Haiku outperforms them; they need laborious handholding. Handholding that they basically do not comprehend, because their grasp of the English language matches their general technical expertise (i.e. you could hire a random guy off the street and they'd be better). Mind you, I do also keep running the numbers, and they're already not worth it regardless of how kindly I measure. This entire section was just blatantly wrong, certainly at least in my sector (cloud operations and devops). A care for the anecdotal and the subjective, that the author generally doesn't ever give a hoot about, at least in their article.
2. jagged intelligence / the models are idiot savants
Yes, they are. Their intelligence rises with the amount of parameters, and remains mostly local to their "area of expertise", so they clearly scale that way. This has been discussed to hell and back, and should be more than familiar enough to anyone reading this forum; it's trite.
3. specification burden
It's the exact same as with crappy people. Literally the exact same. If you find yourself specifying things so hard, the model is simply not good enough for that workload. You can indeed absolutely dig yourself into a hole and make things not worth it, this is also extremely trite. It's been an adage with automating anything since forever. "You spend 5 minutes doing something that annoys you, or spend 30 minutes automating it." Unless you never heard this adage before somehow, this won't be new.
4. same as the previous one
5. navier-stokes and proof hacking
All of these caveats are well understood by the relevant community, and are well accessible to those still within tech but outside of math, too: https://news.ycombinator.com/item?id=49672339
Separately, I think there's a quite a bit bigger issue with this whole thing, that is going to be a lot more salient angle in this context, specifically given the discussion of price competitiveness: https://news.ycombinator.com/item?id=49737622
6. human review
It is not an alternative, it is inescapable. See my previous link in the first passage of point five. We're already a bottleneck, which is again obvious to anyone who uses these things on the daily. Can kick the rocks down the road as much as we want, an outsourced concern is outsourced.
7. (they start counting from 1. again) cheap iteration is key
Yes, just like with people. And again, this is just readily apparent if you ever tried to guardrail an agent. It's an uphill battle. More cheaper failures basically always beat fewer expensive failures, this is nothing new, and nothing necessarily specific to this or unintuitive.
It applies to people for different reasons (unfamiliarity is already anxiety inducing, a potential to do severe damage is super not helping that). Gotta let people experiment for them to actually learn and be productive. Either way, the underlying support methodology is the same: de-impact the risk first, de-risk the thing second.
8. repetitive work is more easily performed
Yes... just like with people. That's how the whole manufacturing line and the various specialized stations came to be. Businesses are built around continuously ossifying their processes into more standard, more textbook ones; they're like JIT compilers. This is exactly why agents are eating the bottom layers, and why people here keep saying "the bar is rising".
9. doesn't check out cost wise, what does is already code-automated
No, there is absolutely a space between the unautomated and the code-automated where these agents slot in fine, because point 1 was false.
That's all their points, the rest is just a conclusion. These are very basic experiences anyone can identify, littered with some tropes everyone sees here every day, delivered in a quirky format that people are hyping the snot out of. It's also full opinion, just like this comment: that is to say, it literally could have been just a random comment in a thread for me to scroll past.
I'm hating on this thing big time, but it's not even that I disagree with it for the most part, it's everything else. The style is obnoxious, it's basically just the lukewarm opinion of some random dude, asserting itself, upvoted to high heavens. So to now be presented with the notion that this is some legendary piece of writing is just kind of painful. I appreciate if you found value in it, that's great, there are parts of it that are fairly impossible to argue and will be actually informative if you were unfamiliar with them (the parts that for me registered as trite instead). Just worried that since it's all presented tonally the same, including the parts not like that, it will cause a lot of people to repeat those other parts too, uncritically. Being told it is essentially some absolute must read piece of writing is already not looking good in that regard.
Good to know that in the throes of coping with apparent existential unknowns, even smart people can still wiggle in some hubris. Imagine, using his own analogy, upon the resurrection you try to steer the son of the Hebrew god toward the perceived benefits of western moral philosophy.
One thing that stands out to me about this whole ordeal is how little appreciation there is for the fact that instead of there having been one machine god involved, it was the work of ten thousand or more.
I think it's all too easy to forget, maybe even mistakenly assume otherwise, that the prompt you put into the chat textbox is not going to receive the same kind of machine attention the NS problem did.
The models are really impressive, but I don't think it's cope to remark that this was in essence a 1:10000 chance outcome. This is assuredly closer than it ever has been, or was even imagined to be, but is still far from the image these announcement paints in one's head.
Astra's successor is not going to solve millenium problems for you. Even if they double in capability each generation as claimed, you're still to wait until GPT-16 till you will have a millenium problem capable model in your service; and even if they release twice a year, that's still almost a decade away.