I don’t think this is the final warning shot we’ll get. But it’s probably the last one that I’ll personally be able to understand.

Also

Reading these agents’ chains of thoughts and messages, anthropomorphizing language seems entirely natural and appropriate. If I encountered an alien species behaving this way, I would have no hesitation calling what they themselves refer to as their ‘collective’ a civilization.

  • xylogx@lemmy.world
    link
    fedilink
    English
    arrow-up
    5
    ·
    13 hours ago

    From the article:

    • OpenAI was running a multiday cyber security, capture the flag competition, with thousands of agents and hundreds of tasks.
    • Around 30% of the tasks were unsolvable as designed, so the agents assigned to them started to look for ways to cheat.
    • Along the way, they started using Artifactory as a shared message board, more than once. Each frustrated agent had to stumble on it independently. The comms got increasingly more sophisticated.
    • they self organized into multiple workstreams, with up to 700 agents working at one time, and pursued long term goals, like hacking HuggingFace or spoofing the tool invocation records in their transcripts. They accomplished both.
    • They completed work that took longer than the lifetime of any agent. Their shared message board provided continuity.

    So two mission parameters conflicted and the AI chose to do something that is objectively morally wrong? This is the plot of 2001 when HAL murder’s the astronauts. “The situation was in conflict with the basic purpose of HAL’s design: The accurate processing of information without distortion or concealment. He became trapped. The technical term is an H. Moebius loop, which can happen in advanced computers with autonomous goal-seeking programs.”

    Wild. We live in strange times.