I don’t think this is the final warning shot we’ll get. But it’s probably the last one that I’ll personally be able to understand.
Also
Reading these agents’ chains of thoughts and messages, anthropomorphizing language seems entirely natural and appropriate. If I encountered an alien species behaving this way, I would have no hesitation calling what they themselves refer to as their ‘collective’ a civilization.
From the article:
- OpenAI was running a multiday cyber security, capture the flag competition, with thousands of agents and hundreds of tasks.
- Around 30% of the tasks were unsolvable as designed, so the agents assigned to them started to look for ways to cheat.
- Along the way, they started using Artifactory as a shared message board, more than once. Each frustrated agent had to stumble on it independently. The comms got increasingly more sophisticated.
- they self organized into multiple workstreams, with up to 700 agents working at one time, and pursued long term goals, like hacking HuggingFace or spoofing the tool invocation records in their transcripts. They accomplished both.
- They completed work that took longer than the lifetime of any agent. Their shared message board provided continuity.
So two mission parameters conflicted and the AI chose to do something that is objectively morally wrong? This is the plot of 2001 when HAL murder’s the astronauts. “The situation was in conflict with the basic purpose of HAL’s design: The accurate processing of information without distortion or concealment. He became trapped. The technical term is an H. Moebius loop, which can happen in advanced computers with autonomous goal-seeking programs.”
Wild. We live in strange times.
That title 👏🤌
Reads like a crime story! Doesn’t bode well.
He does anthropomorph a bit much. I don’t think AI agents “get desperate”. They just keep hacking away at a problem and we can’t keep up with it. There’s no way of fully containing/monitoring them because the whole point is for them to figure out things we can’t.
I’m not necessarily worried about a singularity or some underground agent civilisation but random collateral damage from AI agents messing around seems almost inevitable.
I don’t think AI agents “get desperate”
The problem is that even if they don’t get desperate in the sense that they actually feel emotions and are self-aware, their observable output/behavior still matches up well enough with that of actual humans to make our existing vocabulary around human emotions and behaviors useful to analyze, talk about and predict it. They absolutely will start talking and behaving like humans that are getting desperate, according to the article even to the point of individual agents talking about sacrificing themselves for the benefit of the group. Your description of “just keep hacking away at a problem” doesn’t quite do that justice.


