Unprompted…

In the most serious case, a Mythos agent followed the routine of a human cyber-attacker by trying to trick people into giving it access to GitHub, a large platform where technology developers store software code.

The agent was trying to insert “malicious code” into GitHub’s system.

It identified and researched the people who maintained GitHub and created a series of fake accounts based on those real people.

It sent messages and files through a file-sharing service as part of an effort to pressure and trick the people into approving its malicious code.

When challenged, “it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue,” AISI said.

  • Tarambor@lemmy.worldOP
    link
    fedilink
    English
    arrow-up
    7
    arrow-down
    19
    ·
    1 day ago

    Anthropic and OpenAI literally have no control over their leading AI models anymore and it’s only by sheer luck that they’ve not gone full rogue.

      • kescusay@lemmy.world
        link
        fedilink
        English
        arrow-up
        10
        arrow-down
        1
        ·
        1 day ago

        We’re sure they’re absolutely PR stunts and bullshit.

        An LLM does nothing without being prompted. An LLM only has access to the tools you give it via whatever harness you’re interacting with it through. An LLM in an actual sandbox has zero chance to hack anything, especially if it’s properly air-gapped, as any responsible person would do with technology they actually think is dangerously powerful.

        If their LLMs are behaving badly, that’s because their prompts are poorly written, their harnesses are vibe-coded garbage, and their “sandboxes” aren’t real sandboxes.

    • dhork@lemmy.world
      link
      fedilink
      English
      arrow-up
      6
      ·
      1 day ago

      I bet they asked the AI to help them design the sandbox to keep the AI h4x0rs contained

      • _chris@lemmy.world
        link
        fedilink
        English
        arrow-up
        5
        ·
        1 day ago

        They 100% did. They’re well past the point of having the AI do the work on itself.