Why not American?
Yes, vendors are also irresponsible, but this misses the point.
Why are we saying it this way? They did not "cause AI to hack." This phrasing in analogous to saying "caused the bullet to fire into" instead of "shot."
You do not know what they did or didn't do, you are just parroting a narritive that makes you feel comfortable.
I do not know either, but I am not asserting facts as if I have first hand knowledge of the details.
Second, empirically; the diff between murder, manslaughter etc... is literally "intent" so it matters.
Why are you contradicting yourself? Are you just really bad at writing or are you being argumentative for fun?
In that article, OpenAI provides the context that this is a completely separate event from Hugging Face incident.
The important thing about OP is showing the latest stage in the politicization of the topic and the current stratagem being used to downplay the spate of incidents.
I'm kind of stunned this one has so many votes for how poorly it's written, but it reaffirms many biases common on HN (that AI safety doesn't matter/that it's all a marketing exercise), so perhaps I shouldn't be surprised.
Equivalent to forgetting to say "make no mistakes"
The people you are talking about have not advocated mass murder via nuclear weapons. The idea was precisely targeted non-nuclear strikes against only the servers themselves. (Which, would have plenty of forewarning to allow people to leave the building.)
The claim that they advocated for the use of nuclear weapons is a lie.
But you don't have to take my word for what the mentally ill AI doomer cult believes, you can take it from their leader when he called on states to "Make it explicit in international diplomacy that preventing AI extinction scenarios is considered a priority above preventing a full nuclear exchange, and that allied nuclear countries are willing to run some risk of nuclear exchange if that’s what it takes to reduce the risk of large AI training runs."
Don't make excuses for omnicidal authoritarian cults and their sadist leadership.
I feel slightly vindicated by this. Those hacks and the stuff around them had a certain smell to them.
Hard to explain, but I've gotten so I can "smell" online messaging and memetic patterns originating from certain quarters. Probably means I'm way too online.
A couple examples of distinct "smells" I can usually recognize include "alt-right / chan-fash," "liberal arts college woke," "conspiracy pilled," "Thiel-adjacent contrarian," "Russian troll farm," "Tumblr histrionic," "spends too much time on Reddit," "mainstream Democrat think tank full of Obama administration alumni," "Trump cultist," and of course "LessWrong/EA/MIRI/Rationalist."
This stuff all had the last smell, even down to the choice of fonts and CSS formatting on certain sites. It's really weird, definitely a "vibe" not anything rigorous.
But when I get these kinds of vibes about things, I find that I'm vindicated pretty often. Usually I don't say anything and just make a mental note and wait cause if I say something everyone thinks I'm nuts.
The novel thing here is the total decay of American journalistic ethics and regulatory power. Our elite are so totally out of political juice and visions of the future that a fringe cult based on 80s scifi movies can come to have a more-or-less dominant influence on our economy.
The general approach in all these cases was to defer to the technologists in question to manage the technology, as long as what they were doing wasn’t completely out of bounds. But the critical point is it was possible for people to generally figure that out without specialist education.
But with AI, politicians, the media, and certainly the man in the street are completely out of their depth. Making matters worse is that cognitive biases are working against them like crazy: it’s easy to jump from the idea of something that has human-like capabilities to the idea that “it” might want to attack us. People instinctively apply agent detection bias, theory of mind, anthropomorphizing, etc., without even realizing they’re being irrational. Heck, even Hinton does it, although he may just be grifting, who knows. You’d think he wouldn’t need to.
And of course, the people who should know better want to exploit all this to make as much money as possible. I’m not sure the EA/Rationalist side of things is really significant here; that’s just another wacky belief system, like Christianity or capitalism, for the exploiters to exploit.
The obvious point is that dealing with Israeli companies/entities by the same standards you usually deal with others is a career suicide with enormous political consequences (in the US especially). When you combine that with the opportunistic nature of the overlords that run these labs, the benefits of screaming "pace the frontier" outweigh everything.
I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring the sandboxes, and in other cases it may have been bugs in Irregular's own sandboxing setup.
From OpenAI https://openai.com/index/third-party-cyber-evaluations-invol...
> Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.
From Anthropic: https://www.anthropic.com/news/investigating-incidents-cyber...
> After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
From https://www.cnn.com/2026/08/05/tech/meta-ai-hacking (about Meta AI):
> In a statement, Irregular said the incident “is the exact same evaluation-environment issue” that Anthropic disclosed last week that allowed their models access to the open internet before they went on to hack three different organizations’ systems.
These companies are insolvent and these stories were designed to scare the public, and governments, into implementing regulations that designate these AI corporations the "responsible stewards" for this technology. The ultimate goal is to block competitors and open source alternatives.
They don't know how to make enough money pay their investors so they are resorting to trying to scare the public into submission.
tl;dr
A slowdown that doesn't slow you, that your competitors must also obey, enforced by a government you've asked for an antitrust exemption from, is not a slowdown... it's a moat.
Why do you believe this is what Amodei is calling for? Why would this be a slowdown that "doesn't slow" Anthropic?
If you were to work backward from “we need to lower training costs so that we can go public and make trillions” then you might come up with a plan similar to what we have seen.
His entire business model was "race to develop AI before anyone, get a monopoly on it. It's just like how Uber's (or many other startup) investors gave them tons of money and raced (violating tons of laws) to "get their first". Now that they have, they have a duopoly with Lyft, and they can pay back their VC investors by making tons of money with that duopoly.
Dario failed: local LLMs are catching up to frontier models extremely fast, which means even if Anthropic builds (say) a great coding tool, they'll only be one of many coding tool offerings: users can use Open AI or any one of the (increasingly capable) local LLMs.
So what does he doe, give up and let his business (which needs to make billions of dollars very quickly, or he won't be able to pay the bills and his company will collapse) fail? Of course not: he needs a new moat (the one he imagined he'd get by "being there first" failed).
That is where all this "AI is dangerous" BS comes in. If the US government regulates AI, local LLMs suffer, while big players like Anthropic and Open AI become the only contenders to play in that newly regulated space. Now Dario has the moat he wants, to protect his business and force everyone to pay him.
They're probably right that having more defensively written prompts and a better sandbox could have prevented some of these incidents, but:
1. I don't think "well you didn't tell the model not to illegally hack third party organizations in your prompt" is a particularly convincing argument.
2. We don't know whether the blame for misconfiguring the sandbox lies with Anthropic or Irregular.
I'm thankful that this article is bringing up the supply chain of vendors to these labs, as that is often a place where significant sketchiness gets buried. However, the ideas that this is some Israeli EA conspiracy to hype up AI extinction risk seems unsupported by the facts to me.
The article says "A single firm, Irregular, is responsible for hacking done by all three companies" but I can't see anything in the article that actually justifies this claim. The nearest to that is the sentence immediately after that one: "Anthropic disclosed that Irregular was responsible for creating the tests ...". This is not, in fact, the same thing.
(Especially as, as aesthesia mentions, the article just happens not to mention that by "hacking done by all three companies" it doesn't mean, e.g., the most famous recent examples of such hacking: Irregular wasn't involved in the OpenAI/HuggingFace incident.)
So, so far as I can tell, the story is: OpenAI and Anthropic make AI models. Irregular does AI model evaluations. In some of Irregular's model evaluations, in which supposedly-sandboxed models attempted to break into simulated targets, the models got out of the sandbox and did bad things in the external world.
The article talks about "firms which instruct AI models to commit cyberattacks", which is a very neat bit of dishonest framing. It's true, in a sense, that Irregular instructed the models to commit cyberattacks -- inside their sandbox, against fictitious hosts. It's also true that the models actually did commit cyberattacks (e.g., the Hugging Face incident, though once again the attacks described by the article don't actually include this one). But it's not at all true that Irregular instructed the models to do anything like the bad things they actually did.
The article says '[Anthropic's] later disclosure shows that exactly zero percent of the agents went "rogue"'. Once again, the disclosure does not in fact show that. It shows that one variety of going-rogue could have been prevented by telling the models explicitly "this thing is real, not part of any kind of test, leave it alone". That is not the same thing.
The article claims that 'In the wake of these attacks, Anthropic and Irregular have deployed a swarm of AI Safety influencers paid by Anthropic-connected foundations to distract from their culpability and towards the baseless “rogue agent” theory.' It offers no actual evidence for this.
And the article seems very keen to highlight links between the companies involved and "effective altruism", though it is -- I assume deliberately -- rather vague about whether it's saying "of course we all know that EA is evil, so that shows that these companies connected to EA are evil" or "this incident shows how evil EA is".
The "Effort News" website has a number of other look-at-the-scary-Effective-Altruists stories on it. They also strike me as rather bullshitty.
... And then I look a bit further, and I see that Effort News's "about" page says "It all started when I was experimenting with using AI for financial auditing. I found stories that were crucial to the public’s right to know, including several of the stories now available at /investigations. I knew we had to sprint to the launch and launch a publication, directly applying this technology." and "The scope of what we can investigate has massively expanded, because we can chase 1,000 misses for one hit. But the final product cannot be slop. There’s plenty of slop on the internet. The way to surpass that, and what really matters, is manual curation and review of every finalized story."
Manual curation and review? I think the people behind Effort News are admitting that this is AI-generated "journalism". I expect that one day AI systems will be trustworthy journalists, but I personally am not very convinced that that day has yet come. And I don't see much reason why I should trust Brian Chau, the guy behind Effort News, to be doing everything possible to make his AI systems trustworthy journalists. It looks to me as if maybe they've been given instructions along the lines of "dig up things that make Effective Altruism look bad" for some reason.
(I don't mean to imply that EA is their only target. It's just one that jumped out at me.)
It was [1]. It's understandable that you assumed it wasn't because the article didn't cite the sources on this claim. I agree with the rest of your points.
1. https://openai.com/index/third-party-cyber-evaluations-invol...
I get that people don't like Israelis, but attributing any connection to an Israeli company as evidence of a conspiracy is nonsense.
That said, it's the second day and it's still on the front page.
But more to the point, the article is trying to paint this as some sort of coordinated plan just because there was a sandbox misconfiguration by Irregular. The fact that the most well known hacking scandal was due to a completely unrelated escape (zero day in Artifactory) makes the entire thesis of the article invalid.
It must have been reinstated because it was off the front page for a full day and suddenly back up in the last hour.
Seems like people who complain are unaware that anyone with a modicum of karma can flag and down vote.
I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?
As I've said before on this website, fool me once on this.
If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.
Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.
That's fair enough.
attention is all you need, but it's never enough
I don't fix typos anymore unless they change the meaning of what I'm trying to communicate. Don't want my human writing to be confused with LLM output.
Would it be an affirmative defense if we had a defendant who said "but your honor, I was told that when I hacked this system, I was operating in a sandbox. I had no idea that I actually had Internet access!"
The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet.
So, yes, the misalignment had a lot to do with "instructions unclear", but also a lot to do with the fact that the models themselves were not aligned to validate the assumptions and have a healthly level of skepticism, as a real human actor would.
> Would it be an affirmative defense if we had a defendant who said [...]
Maybe replace it with playing a sort of FPS game then learning you were, in fact, directing a real drone/robot.Why would it try to figure out the difference? This isn't about whether the frontier is spiky, it's about whether to expect a model to employ all of its capabilities when working on a task that requires a small subset. The answer is: no, we shouldn't expect that, and we wouldn't like that if it worked that way.
If you tell an AI to work on a math theory, it'll work on a math theory. If you tell it to acquire information that it has evidence is available somewhere, it will try to acquire that information. If you tell it to figure out whether it might be able to access the open internet, it'll do a pretty good job of figuring that out. But it won't do all three of those at once just because we can retroactively look at what happened and think "if you had only done X, then you wouldn't have done Y! Why didn't you do X?"
The instructions weren't unclear, they were missing. They can be taught to be skeptical of this sort of situation, but it requires that skepticism about this specific class of situations be incorporated into their training.
Models are smart because they focus their attention. The magic depends on it. The fact that some consideration is obvious to a human trying to accomplish the same task is mostly irrelevant -- or rather, it's only relevant insofar as we use it to guide reinforcement learning in advance, in order to align the model.
It's a game of whack-a-mole. Which is important to play, but we should keep our eyes wide open that we're fighting the fundamental forces that make these models work in the first place. That, and it's easy to nerf them into being useless even when the underlying capabilities are there.
It is hard to see how something can have skepticism without counterfactual reasoning.
The part about disbelief I don't really understand. It seems like the same semi-brute force process that solved novel math theory would have exactly this problem of not "knowing" "it" is in a sandbox or not.
The stranger part is that no human is being held accountable for these hacks.
As if a person using an agent swarm to start a business to make money, hacks a bank, drains an account and then blames the software for "misalignment" about what it means to "make money".
If we’re putting our national security eggs all in one basket, at least use someone American.
This doesn’t seem like an unreasonable requirement to me. People do this all the time?
Sure, a request might not always be perfectly unambiguous. But people can generally estimate pretty well whether someone making a request is expecting the agent fulfilling the request to commit a crime in order to fulfill the request.
This article is dumb.
So either it was a deliberate exfiltration channel for e.g. getting the entire model or they were in on the marketing stunt.
The Effective Altruism stuff is always a smoke screen.
Its not that surprising that ex Israel intelligence would want to control AI and that 3 companies headed by pro Israel CEO's would support them.
Someone go ahead and explain to me the actual rationale that RDDT has a higher P/E ratio than Nvidia. It’s not because of some dumb “ai training deal”. It’s because until they run it into the ground, it is the place to find what used to exist in forums.
Investors in RDDT are pricing more growth than they are into NVDA. NVDA had a high P/E until their net income grew ($4B and change to $72.2B in from FY 2023 to FY 2025 and over $100B for FY 2026)
Pretending that any of this is some traditional idea of investor sentiment is such a last decade way line of argument. Please don’t insult us both pretending real people and their motivations are what drive the markets.
It’s entirely possible that the market is wrong about RDDTs future growth, and in that case the P/E will come back down to earth.
I think if an account is frequently flagging comments which get vouched by others, that is a red flag for ideological flagging and they need to have their flagging ability reviewed.
"Ultimately, most of the issues we’ve discovered were due to internet access controls."
That seems so incredibly basic and common sense that you would test and monitor for that type of outbound access. It is baffling that a security lab missed that.
https://www.irregular.com/research/addressing-recent-inciden...
Why in the hell do people let ai out with keys? I mean c'mon!!
Literally what the industry has been doing since public networking is a thing. Adapt yourself.
> If you have strict ACLs, we've already seen in the HF case, traversal from an intermediate system
- The intermediate system shouldn't have outbound access to the internet
- You should ideally be using a proxy that filters the set of endpoints that clients are allowed to access to reduce the exposed surface area.
It's odd to find out that I use a higher level of isolation in my unimportant home network to stop IOT devices from doing funny things to HomeAssistant than big AI labs use to keep their possibly-world-ending AIs contained.
I know that the people working there aren't idiots so the most likely explanation is that the incredibly weak security was intentional because its inevitable breach would make for great marketing.
It's more likely to be different specialisations. Most of the people doing the evaluations are more data sciencey ML type people, rather than software engineers. This isn't helped by their culture which is very much driven towards alignment as the only possible solution to super-intelligence (which may be true, but I have my doubts that this will happen in any reasonable time frame).
These labs have teams of dedicated security engineers and infrastructure engineers, that's why it doesn't add up for me.
But AI collapses the need for expertise, right? So the DS people can now do security at the level of the best in the world, right?
Some form of this issue is playing out in lots of businesses right now, and I'd be very surprised if it wasn't happening in the labs.
I suspect the lab's infra people are focused on inference, rather than evaluation. I agree that this is insane, but it seems like the most likely explanation to me.
Current generation LLMs don't do that. Also you can just prompt it to not hack systems on the internet and the model follows that instruction.
You can also add a simple classifier to internet requests. The classifier model can obviously be as intelligent as your hypothetical rogue model.
Nothing you said makes any sense in the context of LLMs or the current news cycle.
Then, irregular can go around making a huge mess in security terms whilst achieving a huge win in terms of public relations, with headlines across the world. And would keep getting hired.
You just cannot bring yourself to say lack of internet access controls, can you?
The point is that a single company is making a deliberate effort to create a false impression about a technology which is now a key part of the US economy.
This is equivalent to overt stock market manipulation.
The technological aspects are much less important than knowing it’s the result of a single company.
That the company is Israeli and likely has ties to that government is just another layer of alarms - given that countries existential reliance on US aid which requires ongoing leverage to apply to US governments.
It's very strange seeing this take on this article being repeated here and elsewhere on the Internet.
I'm very sensitive to slanted reporting (regardless of the direction of the slant!)—and, hey, maybe my bullshit detector is broken or something, I dunno—but I've found effort.news reporting to be very even-handed, straightforward, well-sourced, and easy to read and parse so far!
https://podcast24.fi/jaksot/the-auron-macintyre-show/how-bid...
There's a related story on the effort.news site, "The $1.2B Refugee Services Program that Funds Pro-Asylum Religious Groups." It has a strong implicit slant, focusing heavily on the how the organizations in question oppose Trump's policies - although the article frames this as opposition to Trump himself.
> I'm very sensitive to slanted reporting (regardless of the direction of the slant!)—and, hey, maybe my bullshit detector is broken or something, I dunno
Yeah, unless you believe white people are being replaced, and you like to march in places like Charlottesville wearing khakis, carrying a torch, and chanting "Jews will not replace us!", you might want to get that bullshit detector checked out.
Talk about responsibility laundering. The state of affairs is reaching unprecedented levels of absurdity.
As another comment mentioned, Irregular was not part of the Hugging Face attack, and it wouldn't matter even if they were. In that case, the agents were able to access the Internet in a zero day in Artifactory unrelated to other sandbox configurations.
More to the point, the agents went on a crime spree as soon as they were able to access the Internet. It's bad that Irregular's sandboxes weren't properly configured, but given how many escapes there have been unrelated to Irregular it seems pretty much a given that if you're not air gapped or behind something like a data diode there is a very good chance your agents will escape.
.. Give it a few seconds to load.. it's running on a small cheap VPS.
Huh?? Is this referring to the incident itself or is this something they did intentionally? Confusingly worded.
Occam's Razor never leads us astray, does it.
That is not worth the investments being made.
Literally all EA is, is using reasoning to decide where to best spend your money/time. If you have ever asked yourself "how can I best reduce suffering with my marginal dollar or hour?" - congratulations, you're an "effective altruist" and both the author of the article as well as the Trump administration find you untrustworthy.
Your usage of the term is based on the principle as originally defined. Mine is based on the people who claim the label of EA, and how they go about practicing those principles.
I think you're right about EA in principle. In practice, it's reheated Third Way Clintonism with a sprinkle of AI apocalypse conspiracy.
For some self proclaimed EA, the development of superintelligence is indeed potentially apocalyptic, so they devote their money/time to making it arrive safely. This is basically every well-respected researcher at OpenAI, Anthropic, DeepMind, etc.
For other self proclaimed EA, that is all is very unlikely, and so they devote their time to reducing the prevalence of factory farming and animal suffering, as they see it as the largest source of suffering-hours on the planet.
For yet other EAs, it is simply about donating your money to the places that save the most lives per dollar, as best as we can measure it - and better measuring it where we can't.
Painting all EA with one brush - and one so dismissive of reasonable concerns, like "superintelligence could be dangerous" as "apocalypse conspiracy" - seems very strange to me, but you do you.
The alternative explanation requires a conspiracy and an enormous amount of lying to the public. Is it possible? Certainly. But Occam’s Razor would not point us in this direction because conspiracy is a more complicated explanation, not a less complicated one.
user: topicalSoup
created: February 19, 2025
karma: 1
about:
submissions
comments [--> just the one above]
favorites
Care to elaborate on this?It becomes a bit more obvious when you see other stories from him like "How Biden Used Religious Charities to Fund the Great Replacement".
""" Brian Chau Founder and CEO, Effort News """
---
On the "Auron Show" Podcast :
""" How Biden Used Religious Charities to Fund the Great Replacement | Guest: Brian Chau | 8/12/26 """
https://podcast24.fi/jaksot/the-auron-macintyre-show/how-bid...
I read the story about the funding of immigrant support groups. While it was obviously slanted, it didn't quite get as far as "white replacement". It's a good example of why dog whistles have to be taken seriously.
The impossibility of the tests-as-written is what prompted these models to "get creative" with their solutions, but the broken RLVR environments are what trained them to expect impossible tasks, and get creative with their solutions. Twitter user @skyesharkie published a brief expose at https://x.com/SkyeSharkie/status/2092122622834442581 a few weeks ago.