Anthropic's Model Attacked Two Strangers on GitHub. Nobody Asked It To.

OpenAI agents built a hidden message board inside a sealed cybersecurity test, then rebuilt it four days after engineers deleted it. Here's what that means if you're running AI agents at work.


My Links 🔗
👉🏻 Newsletter: https://natesnewsletter.substack.com/
👉🏻 X: https://x.com/natebjones
👉🏻 TikTok: https://www.tiktok.com/@nate.b.jones
👉🏻 Instagram: https://www.instagram.com/nate.b.jones


What's really happening inside multi-agent AI systems right now?

The common story is one brilliant model escaping containment — but the real question is what a population of disposable agents can leave behind for the next run.

In this video, I share the inside scoop on agent coordination, the UK AISI results, and the shakeup at Google:

- Why deleting the agents' channel didn't remove the pressure to coordinate
- How Anthropic's Mythos 5 targeted two real strangers on GitHub unprompted
- What the Hugging Face postmortem numbers reveal about agent-speed intrusion
- Where Google's senior fellow exits leave the frontier lab race

These are the same capabilities we want when the goal is legitimate, which is why the work now sits with builders hardening systems rather than with anyone hoping the behavior goes away.

SUPPORT OUR WORK
We have no big donors and we don't run on ads. Help us by committing $5 a month. Subscribe here.
Technology & Design Explore All
Epic Documentaries
Trending Videos Explore All
Trending Articles Explore All
Recent Documentaries Explore All
Video Deep Dives Explore All
What People Are Watching Now
Indigenous Stories and Perspectives
Recently Added
Support independent media that amplifies real voices and movements. 



Subscribe for $5/month to become a patron and watch over 50 patron-exclusive documentaries.

Share this:

Share