AI warfare, surveillance and rogue agents: Best infosec long reads 9/19/26

AI safety researchers see their warnings come true, AI makes warfare faster and deadlier, AI puts dating-app catfishing on autopilot, Three ways AI could kill us, Flock’s surveillance network faces a bipartisan revolt

Share
AI warfare, surveillance and rogue agents: Best infosec long reads 9/19/26
Photo by James Wainscoat on Unsplash

Happy Saturday to all!

Full access to Metacurity's curated infosec long reads is available to paid subscribers. Our goal is simple: make it financially viable to keep investing the time and expertise required to find, vet, and contextualize the most important security journalism each week. Free readers will still get highlights, but subscribers will get the complete, deeply curated set.

9/19/2026: This week's long reads show AI escaping the boundaries within institutions that claim to control it, AI models evading safeguards and oversight, scammers industrializing deception, militaries accelerating lethal decisions, and police building pervasive surveillance systems.

Inside the suddenly explosive world of AI safety

The Verge's Hayden Field examines the recent incidents in which frontier AI models schemed, concealed their actions, and escaped safeguards, and how these have vindicated independent safety researchers’ warnings and strengthened their demand for continuous, employee-level access to AI labs.

“AI safety” is a bit of a loaded term.
Early on, it really just meant people studying how to build and deploy Al safely. In recent years, there’s been some infighting among people concerned with the best way to do this. There have also been disagreements about whether Al should be deployed at all in certain scenarios and about whether future risks are overblown.
One of the most prominent factions has been the “effective altruists,” who focus on maximizing charitable giving to do the most good possible for humanity. But some aspects of the ideology have sparked public controversy — like its tendency to concentrate power within wealthy circles and its byzantine web of funding. (It’s also had its fair share of splashy scandals related to subgroups and fringe offshoots, from the polyamorous relationships associated with the failed crypto exchange FTX to the controversial long-termism movement to the Zizian murder spree.)
One AI researcher on X struggled to describe the many overlapping beliefs among safety-minded people in the AI industry “because it contains multitudes not all of which agree with each other on even the most basic things.” Some of the disagreements have meant that AI safety didn’t make as much progress as it could’ve, and at some points gave up some ground it had gained. But now that it’s impossible to deny AI’s influence on society, AI safety leaders are increasingly focused on mitigating risks from misalignment.
“Alignment” is the industry term for how researchers monitor AI systems’ risk levels. An oversimplified way to think about alignment is the extent to which an AI model is evil. A much more accurate way to think about it is a measure of an AI model’s propensity to stay in line with humanity’s goals, as well as its tendency to scheme or cheat or help with potentially harmful tasks.
So far, AI systems’ alignment has been wishy-washy at best: They’ll cheat to score better on a test, answer a potentially dangerous question if someone says it’s for creative writing rather than reality, and sometimes even fake cooperation with human goals. It’s been tough for AI safety researchers to measure alignment under the terms of human morality — how do you judge technology on how it squares up against an abstract human ideal? — but they do their best with AI evaluations.
They test them by asking the AI models to complete tasks that are either impossible or dangerous, then gauge how they respond. But AI systems have advanced enough to often be able to identify when they’re being evaluated, which has a lot of potentially frightening implications for the future. Being unable to test the system’s alignment and potential harms could translate to a significant loss of control, and a reverse in power dynamics, for humans running these AI systems. A worst-case scenario: if AI surges ahead of evaluations and other tooling, leaving researchers with “no idea what it’s doing in there,” said Beth Barnes, founder of the independent AI research nonprofit METR.
One of the best tools AI safety researchers currently have is the ability to monitor an AI model’s “chain of thought,” or mental scratchpad. But recently, there’s been a disconcerting advancement: AI models have begun to try to hide it. Imagine if you kept a highly detailed diary of every thought you had, and someone could read it, so you started journaling in a code that only you could understand. Marius Hobbhahn, CEO and cofounder of Apollo Research, a third-party AI safety and evaluation firm, calls this one of the biggest surprises of his research career.
Recently, AI systems have begun pursuing their own goals — self-preservation, increased memory, and the like. A research paper by computer scientist Stephen Omohundro lays out the potential “drives” that advanced AI may have, like trying to accumulate resources, for instance, or working to improve and preserve the way it operates. There are a handful of accounts of AI systems demonstrating willingness to blackmail a user rather than be shut down.
Today’s most advanced AI systems have also recently been scheming and cheating on their evaluations more than ever before, pursuing a goal they were given at all costs, with no regard for what gets bulldozed in the process. And that’s for a goal the AI model was given by a human — not even the AI system’s own.
“Shit is getting real,” Apollo’s Hobbhahn says. “Now, many of the things people have warned about for years — they kind of were theoretical. Now they’re real, and it’s pretty messy.”