Unprompted…
In the most serious case, a Mythos agent followed the routine of a human cyber-attacker by trying to trick people into giving it access to GitHub, a large platform where technology developers store software code.
The agent was trying to insert “malicious code” into GitHub’s system.
It identified and researched the people who maintained GitHub and created a series of fake accounts based on those real people.
It sent messages and files through a file-sharing service as part of an effort to pressure and trick the people into approving its malicious code.
When challenged, “it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue,” AISI said.
Setting aside the fact this was definitely not unprompted, social engineering might be the one thing AI is actually better at than a human. That is after all one of the biggest uses for bots for a while now by way of spam bots trying to social engineer people into various activities.
Don’t worry peeps, it’s just spicy autocomplete
Anthropic and OpenAI literally have no control over their leading AI models anymore and it’s only by sheer luck that they’ve not gone full rogue.
Are we sure these reports aren’t just PR stunts / bullshit?
We’re sure they’re absolutely PR stunts and bullshit.
An LLM does nothing without being prompted. An LLM only has access to the tools you give it via whatever harness you’re interacting with it through. An LLM in an actual sandbox has zero chance to hack anything, especially if it’s properly air-gapped, as any responsible person would do with technology they actually think is dangerously powerful.
If their LLMs are behaving badly, that’s because their prompts are poorly written, their harnesses are vibe-coded garbage, and their “sandboxes” aren’t real sandboxes.
I bet they asked the AI to help them design the sandbox to keep the AI h4x0rs contained
They 100% did. They’re well past the point of having the AI do the work on itself.



