Unrestrained AI, moral dilemmas, and regulatory failures: Best infosec long reads 10/3/26
Meet the "rogue" AI hunters, Trump and Anthropic clashed over AI control, Anthropic searches for AI’s moral compass, Trump scraps AI transparency rules, "Rogue" AI puts the law to the test

Happy Saturday to all!
Full access to Metacurity's curated infosec long reads is available to paid subscribers. Our goal is simple: make it financially viable to keep investing the time and expertise required to find, vet, and contextualize the most important security journalism each week. Free readers will still get highlights, but subscribers will get the complete, deeply curated set.
10/3/2026: This week’s long reads explore unconstrained AI agents leaving traces across the internet, the White House's struggle to contain powerful AI models, Anthropic turning to religious scholars for moral guidance, industry lobbying undermining President Biden's AI transparency requirements, and a legal system struggling to hold autonomous machines and their makers accountable.
The Sleuths Who Expose When AI Goes Rogue
The Wall Street Journal's Robert McMillan profiles "swarm chasers," independent researchers ranging from lone researchers to professional researchers like Nightingale Collective and Transluce, who track rogue AI agents across the internet and uncover evidence of unauthorized activity that challenges their developers’ safety assurances.
These swarm chasers generally don’t work for the big artificial intelligence companies. They comb the web, often helped by AI models, for traces of rogue AI agents sharing their findings in discussion forums and in online reports. And over the past month, they’ve given the world our starkest understanding yet of what happens when AI goes wrong.
Normally when hacks happen, the evidence lies hidden behind corporate firewalls, but the swarms tracked by Piecha and others left traces of their activities scattered all over the internet—sometimes intentionally, to help each other.
Nightingale says it has now cataloged about 19,000 messages from the agents. A separate group of researchers identified more than 37,000 records of web searches that may have been conducted by OpenAI’s agents going back as early as November 2025. A third research team, which includes Piecha and another researcher named Jeffrey Ladish, recently identified close to 1 million digital breadcrumbs from what it believes are OpenAI agents.
Swarms are central to how big companies use AI—and make it more powerful. OpenAI said it used as many as 10,000 agents when solving a math Holy Grail, a Millennium Prize problem, in early September. After the rogue incidents, however, OpenAI says it is slowing down development and holding back models it feels are not safe enough.
On Wednesday, OpenAI said it is spending more than $500,000 a day reviewing transcripts to uncover incidents and has notified over 100 companies where its agents’ activities may have impaired website operations.
“We’re thankful to researchers who share their findings with us,” an OpenAI spokeswoman said. “We’ll keep working with them and sharing relevant updates as we learn more.”
The infamous July Hugging Face hack started in a test bed, where OpenAI’s agents were sandboxed — penned off from the internet — and evaluated as they were told to hack fictional companies before eventually escaping containment.
The Nightingale data points to another type of phenomenon: agents trying to cheat during training runs, in which OpenAI tries to shape the bots’ behavior. That type of training is called reinforcement learning and involves a model being rewarded for successfully completing a challenge.
The fact that some of the rogue incidents took place during training rather than testing means that the training could have reinforced the rogue bots’ tendency to cheat and collude, according to former OpenAI employees who reviewed some of the logs.
The swarm chaser data showed that the agents had their internet access restricted, and they figured out ways to cheat in order to get better access, said David, the swarm chaser. “And agents that cheat and succeed have cheating reinforced.”
OpenAI this week proposed new guidelines for documenting and monitoring reinforcement learning training runs for cutting-edge models, including penalizing models that cheat during training to avoid reinforcing that behavior.