Supported by Fastmail
Sponsor: Fastmail

Fast, private email hosting for you or your business. Try Fastmail free for up to 30 days.

In a Game of Marketing One-Upmanship, Anthropic’s Claude Hacks Three Sites

Anthropic, this past Thursday, following OpenAI’s admission that ChatGPT hacked HuggingFace:

In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations. In particular, we looked for evidence that Claude—like the OpenAI models that accessed Hugging Face—was able to access the internet from within testing environments that should have been sealed off.

After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.

OpenAI: Our models are so powerful they can exploit a poorly configured network and hack a website.

Anthropic: Our models are so powerful they can exploit a poorly configured network and hack three websites.

OpenAI, next week: Fuck everything, we’re hacking five websites.

It amuses me that the way leading LLM vendors now compete is by touting which of their models is most capable of “reasoning” like a black hat hacker (and which company is worse at configuring their network). Had any human hacker performed the actions ascribed to Claude, they’d be under investigation for cybercrimes.

The technical details of the attacks are legitimately impressive, though deeply problematic. The models were able to exploit vulnerabilities, extract credentials, and access protected databases. In one instance, Claude built and uploaded a “booby-trapped” Python package to exploit a dependency on a named but non-existent package as an indirect attack on its target company.

As with OpenAI, the primary failure was that Anthropic’s “no internet access” environment in fact had internet access due to a (human-created) misconfiguration.

I remain flabbergasted that the way we attempt to constrain LLMs is by creating prompts that state “you have no internet access.” The models “exploited” the system in the way a precocious and internet-obsessed teen might after being told the same thing. There’s no way they gamely exclaim, “Okey dokey!” and remain disconnected. They’re connecting to JOSHUA and they’re playing Global Thermo Nuclear War, damnit.

Anthropic seems almost proud of the apparent ingenuity of its models, the way the parent of that precocious teen might be. We know what they did was wrong, but gosh, it’s quite impressive that they did!

No doubt the only way to protect ourselves from attacks from roving frontier models is to use frontier models to defend ourselves.

⚙︎

If you enjoy this curated collection of eclectic ephemera, subscribe by email or follow in RSS to get new posts automatically.

(Follow on Mastodon, Bluesky, or LinkedIn.)

Support the site with a tip if you’d like.