AI shifts from cyber assistant to attack operator, Anthropic finds

Anthropic says Russian spies, criminals and hacktivists used Claude to accelerate and automate cyberattacks, while other actors pursued potential biological weapons research, military programs and industrial-scale theft of Anthropic’s models.

Share
AI shifts from cyber assistant to attack operator, Anthropic finds
A complex workflow of how AI was used to disrupt Russia's Midnight Blizzard cyber operations, the effects on targets, and more. Source: Anthropic

Anthropic said it had disrupted several potential plots this year by scientists who used its leading artificial intelligence models to conduct research that could have helped develop biological weapons.

In a report describing misuses of its AI models, Anthropic said it could not determine whether the research served a legitimate or nefarious purpose because valid biological inquiry — the kind that can lead to breakthroughs like vaccines — can also help engineer dangerous pathogens. In the face of that uncertainty, Anthropic said it erred on the side of caution because the consequences of missing malicious activity could be severe.

The potential for cutting-edge AI models to facilitate the development of known or entirely new biological pathogens is among the gravest concerns experts have about a technology that is developing so rapidly that even its leading architects doubt whether humans will be able to control it fully.

Compared with the threat of catastrophic cyberattacks or AI agents that fail to align with human intentions, biological misuse has gained less attention as an existential risk of AI, in part because past examples have generally been shown only in research settings rather than in the real world.

The report also demonstrated that Chinese artificial intelligence companies attempted to clone American AI models by funneling real Chinese user queries to them via a network of intermediate platforms,

Companies including Moonshot AI and DeepSeek used these middlemen to relay millions of their own users’ queries, including ones containing sensitive data such as location information and passwords, to Anthropic’s Claude, the report alleges, offering a detailed new look at the cloning technique known as distillation that American officials liken to industrial-scale theft.

US labs generally ban outside distillation of their proprietary models, but American officials say the technique remains a widely used shortcut for rivals to catch up to more advanced competitors.

The details about AI theft come as concerns are mounting inside AI companies and Washington about the potential safety risks of powerful AI models, with some safety advocates and companies calling for a slowdown in AI research.

Anthropic, in its report, said it has seen increasing evidence that Claude was used for things like criminal hacking campaigns and potentially dangerous biological research. In one case, a user attempted to adapt avian flu to infect mammals. OpenAI has also seen similar types of queries about bioweapons and poisons.

From a cyber threat perspective, Anthropic claims that AI is increasingly functioning as an operational layer for cyberattacks. According to its report, state-backed groups, criminals, and hacktivists used Claude for reconnaissance, exploitation, phishing, persistence, and data theft.

The attacks generally exploited familiar weaknesses—stolen credentials, exposed services, unpatched systems, and cloud misconfigurations. What changed was their economics: AI enabled smaller, less-skilled groups to attack more victims faster. Anthropic also found attackers stealing AI API keys for resale, free computing resources, and cover, making AI vendors, agents, and integrations an increasingly valuable part of the attack surface.

Anthropic said that it detected and disrupted a Russian cyber-espionage campaign targeting Ukrainian government, military, and diplomatic organizations. The attackers used Claude models at almost every stage of the attacks and automated parts of the process.

Anthropic linked the attacks on Ukrainian organizations to the Russian group Midnight Blizzard. The attackers used phishing campaigns, interception of hotel Wi-Fi traffic, and WhatsApp account takeovers. Claude was used at virtually every stage of the operation.

AI helped the hackers create and modify malware. Anthropic said they even built a system that automatically checked whether security tools could detect the malicious code and modified it if they did.

The company noted that in such attacks, humans are increasingly acting as supervisors rather than directly carrying out operations. This allows threat actors to automate a much larger proportion of cyber operations.

Anthropic also identified the use of Claude in military programs in Russia, China and Yemen. The model was used for work related to missiles, drones and munitions, as well as intelligence gathering and procurement.

Finally, Donald Trump brushed aside warnings from artificial intelligence researchers that the technology is becoming too hard to control and could kill all humans, saying his paramount interest is maintaining an edge over China.

Asked by reporters Thursday upon departing Dallas if he has any concerns about AI leading to human extinction, Trump said, “No, I don’t have any.”

“I have concerns that if we don’t win AI, we’re going to be put in a very bad position,” the president said. “We are leading China right now by a pretty good period. I would say a year, which is, you know, considered a lot.” (Dustin Volz / New York Times, Robert McMillan and Amrith Ramkumar / Wall Street Journal, Myroslav Trinko / Ukrainska Pravda and Jeff Mason / Bloomberg)

Related: Anthropic, TechCrunchSalonSiliconANGLECyberScoopCBS NewsUnite.AIWashington ExaminerNBC NewsSFistPoliticoThe Daily Caller, Mashable, Associated Press, The Register, Hacker News, The Guardian, The InformationFinancial TimesFox BusinessEngadgetNewserFox NewsKTVZ-TVTom's GuideQuartzNew York PostDaily MailGadget Review, Hacker Newsr/technologyr/Contagion,  r/USNEWSr/worldnewsr/singularity, Slashdot, Associated Press, CNBC, Euronews, Axios

Breakdown of attackers' use of AI. Source: Anthropic

Metacurity is the cybersecurity news, analysis, and insight you'd need hours and possibly days to assemble yourself.

Every weekday, we read the releases, filings, court documents, and reports that vendors and PR teams often don't want summarized — then tell you what actually changed and why it matters. Minimum vendor marketing, no outrage bait, no SEO filler.

A paid subscription to Metacurity delivers

  • Full archive access — every newsletter and AI Watch roundup, searchable and browsable.
  • Our weekly curated long-reads roundup — the best cybersecurity writing from across the industry, filtered and vetted so you're not sorting through it yourself,
  • Periodic specialized reports and analyses — deep dives that go beyond our daily coverage
  • Support for independent, no-spin cybersecurity journalism — funded by readers, not vendors or investors.

Reader support is what keeps Metacurity independent. It allows us to focus on serving the cybersecurity community—not advertisers, vendors, or investors—and to continue delivering the thoughtful analysis you've come to rely on every weekday.

Please consider supporting us. And thank you!


A Republican-led Senate subcommittee that oversees disaster management is investigating OpenAI's handling of the Hugging Face breach in July.

"As you may know, in the public domain, more AI experts are warning about the existential risks of AI," Sen. Josh Hawley (R-Mo.) wrote in a letter to OpenAI CEO Sam Altman.

Hawley, the chair of the Senate Homeland Security & Governmental Affairs subcommittee on Disaster Management, wrote that he is launching the probe in response to findings from OpenAI's recently released internal investigation.

Hawley described OpenAI's handling of the cyber test — specifically not taking more drastic action after researchers became aware their agents had gone rogue — as "reckless." He also noted that OpenAI "redacted many important details" about the incident in its report.

"The American people deserve to know the details of what went on in the Hugging Face incident and other incidents of AI models going rogue," he wrote. (Andrew Solender, Maria Curi / Axios)

Related: The Information, QuartzReutersTechSpot, Forbes, International Business Times, Mother Jones

Anthropic gave the European Union’s cybersecurity agency access to its powerful Mythos artificial intelligence model, more than three months after first signaling the bloc would be given access.

“Following our constructive engagement with Anthropic, we can confirm that the EU’s cybersecurity agency ENISA has been granted access to Mythos 5 and is testing it now,” Thomas Regnier, a spokesperson for the EU’s executive branch, the European Commission, said in an email.

After lobbying from the EU, Anthropic proposed allowing the European cybersecurity body ENISA access to the model in late May. But that kicked off months of wrangling over the type and scope of ENISA’s access.

Those negotiations were further complicated by the White House restricting foreign access to Mythos and another powerful Anthropic model, Fable. It later eased those restrictions, though foreign institutions’ access to Mythos remained in limbo.

Despite the successful negotiations, the EU’s achievement is only partial, as ENISA will not have access to the model’s latest iteration, Mythos 5.1, the commission spokesperson said.

The UK’s AI Safety Institute, which had been among the first non-US entities to stress-test the original Mythos, was not given access to this new version, according to a person familiar with the matter. (Gian Volpicelli / Bloomberg)

Related: Reuters, Silicon RepublicThe Next Web, Financial Times

Japan’s Digital Agency said it suffered unauthorized access to its servers and that personal data on about 246,000 people may have been leaked.

The agency detected access to a large volume of files on its Government Solution Service network using a maintenance account in late June, according to a statement on Friday. A subsequent investigation found that someone had exploited a vulnerability in a virtual private network and possibly stole information on public servants and others who use the system.

Data that may have leaked include names, email addresses and phone numbers, the agency said, adding it hasn’t confirmed any misuse linked to the incident. (Maho Nambu / Bloomberg)

Related: Anadolu Ajansı, Japan Wire, The Star

The US Justice Department announced that Ukrainian national Oleksii Oleksiyovych Lytvynenko has been sentenced to four years in prison for his role in Conti ransomware attacks between 2021 and 2022.

Lytvynenko was arrested by the Irish national police (An Garda Síochána) in July 2023 at the request of the United States and was extradited last year.

Lytvynenko and his Conti accomplices deployed ransomware on victim networks in the United States and abroad, stealing data and encrypting devices to extort Bitcoin ransom payments.

"From 2020 until 2022, Conti was used to attack computers and networks in 47 states, 31 foreign countries, the District of Columbia, and Puerto Rico. The FBI estimates that, as of January 2022, there had been victim payouts associated with Conti ransomware exceeding $150,000,000," the Department of Justice said.

"Lytvynenko joined that conspiracy as both an intruder and a developer — personally harming at least 12 companies, storing stolen data from victims, and helping build the malicious tools Conti used to extort and threaten communities," added Assistant Attorney General A. Tysen Duva. (Sergiu Gatlan / Bleeping Computer)

Related: Justice Department, CyberScoop, Cyber Daily

ID verification service IDScan has confirmed that a data breach involved the theft of driver’s licenses from its systems, a week after a report said the identity document checker had been breached during a year-long hack.

The company said in a website notice that hackers stole the driver’s licenses from the company’s cloud; the stolen information includes people’s full names and driver’s license numbers, along with identity numbers from other government-issued documents, such as passports.

IDScan said in its notice that it “received information” on or around September 1 about a claim of a hack, on the same day that independent cybersecurity journalist Brian Krebs first reported a data breach at IDScan. 

Krebs reported that he was alerted to a website on the dark web that allowed anyone to search the driver’s license information of over 150 million people living in the United States and Canada, including accessing their photos. Krebs verified the authenticity of the data by examining his own record. The database also contained high-profile individuals, including the U.S. Secretary of Defense Pete Hegseth, and a security researcher who also verified his data for Krebs’ report. (Zack Whittaker / TechCrunch)

Related: TechRadarPCMagCyberInsiderEngadgetBleepingComputer, The Record, PYMNTS, Cybersecurity Insiders, CNET, ID Scan

Surfshark disclosed that hackers accessed one of its internal test servers after a configuration error exposed it to the internet.

The VPN service provider said the incident did not affect its customers and did not extend to other parts of its infrastructure, but it exposed service configurations and build-related credentials.

“Due to a human error, an internal test server used by our engineering teams was misconfigured in a way that made it reachable from the internet,” Surfshark explained on its website.

The exposed environment also contained portions of system binaries and code history.

Surfshark said that the unauthorized party accessed a separate server used for content-accessibility optimization. The machine acted as a proxy and did not have access to any sensitive data, like user identity, IP addresses, encryption keys, or browsing traffic. (Bill Toulas / Bleeping Computer)

Related: Surfshark, TechRadar, Cyber Daily

Researchers at GreyNoise report that a threat actor, likely Russian-speaking, used hundreds of AI agents to develop and launch a global exploitation campaign targeting vulnerable PaperCut NG/MF servers.

The agents were tasked with building, testing, and refining exploits for CVE-2026-81578 and CVE-2026-82078, both security flaws affecting PaperCut Software and flagged as actively exploited earlier this month.

GreyNoise says the campaign began on August 31, combining OpenAI’s Codex and DeepSeek models with commodity offensive tools.

The AI agents also generated target lists through the Netlas internet scanning and discovery platform.

GreyNoise data indicates that the operation compromised at least 440 PaperCut instances linked to 395 distinct organizations across 48 countries.

The attacker harvested credentials from 280 victims, obtained operating system or domain secrets from 147, and obtained administrator privileges at 12 organizations.

Most of the victims were in the education sector, accounting for roughly half of all breaches. The United States was the most targeted country, followed by the United Kingdom, France, Spain, and Canada.

According to GreyNoise, the threat actor specified a list of countries to avoid, including Russia, China, Iran, Ukraine, Belarus, Moldova, Brazil, and South Africa. However, the agents did not consistently follow these rules. (Bill Toulas / Bleeping Computer)

Related:  Grey Noise, HackReadThe Register - SecurityeSecurityPlanet, Cyber Security News, Cyber Press

Attack timeline. Source: GreyNoise

As Trezor warned, threat actors who breached Brevo, its third-party email provider, were emailing customers who opted in to receive newsletters.

According to customers targeted in this phishing campaign, they received fake "critical security alert" emails from help@trezor.io claiming that a "hardware microcontroller vulnerability" in Trezor cold storage wallets' STM32 microcontrollers could expose their seeds to brute-force cracking.

The phishing emails tried to trick recipients into clicking a malicious link that prompted them to download an app that asked them to enter their wallet backup.

Trezor says that it took down the domain used in the phishing attacks within 20 minutes, disabling the link and limiting the campaign's impact to 2,500 customers who had clicked it before it was taken down.

"On September 9, 2026, Brevo, the third-party marketing platform Trezor uses for newsletter campaigns, suffered a security incident affecting 120 Brevo accounts. An unauthorized actor gained access to Brevo's system and used it to send emails from various customer accounts, including Trezor's," the company said.

"The incident affected our opt-in newsletter database, roughly 347,000 email addresses. These addresses might be potentially used for other phishing attacks in the future. No other Trezor system was touched. We have suspended the Brevo account to stop further email distribution." (Sergiu Gatlan / Bleeping Computer)

Related: Trezor, Bitcoin Magazine, crypto.news, Crypto News, Protos, Blockonomi

Trezor phishing email. Source: Geo Soul.

A Ukrainian IT specialist was sentenced in Zurich to 12 years and nine months in prison and banned from Switzerland for ten years for his role in ransomware attacks on several firms, including Stadler Rail.

Prosecutors put the damage at around CHF100 million ($123 million).

In its judgment, the Zurich District Court went further than the prosecution, which had sought a prison sentence of twelve years for the 52-year-old IT specialist. The man, who was living in the canton of Basel-Country, has been in pre-trial detention since October 2021.

The court found that the defendant had played a key role in cyberattacks against companies such as Stadler Rail, Meier Tobler and Crealogix with the aim of extorting ransoms. He is regarded as the lead developer of the Lockergoga, Megacortex and Nefilim malware programs, which were used to encrypt data belonging to the targeted companies. (SWI)

Related: blue news

Korea's privacy regulator is sharply raising the cost of data breaches, aiming to push companies to treat data protection as a preventive investment rather than a routine cost of doing business.

Starting Friday, companies found to have leaked the personal data of 10 million or more people through intent or gross negligence can be fined up to 10 percent of their total revenue as part of a broader overhaul under the revised Personal Information Protection Act that is set to take effect the same day. Even if a leak hasn't been confirmed, companies must notify users within 72 hours if the risk of exposure is high.

“Personal data breaches have recently occurred repeatedly and grown in scale in fields closely tied to daily life, such as retail and telecommunications,” Personal Information Protection Commission (PIPC) Secretary General Yang Cheong-sam told reporters Thursday. “We've improved the system to hold serious violations strictly accountable while also helping prevent breaches from happening in the first place.” (Korea JoongAng Daily)

Related: UPI, ChosunBiz, KoreaJoongAng Daily, Seoul Economic Daily

Threat actors sent more than one million AI-assisted invoice fraud emails in a large business email compromise (BEC) campaign targeting enterprise users.

The operation, detected by Microsoft between August 3 and 5, impersonated executives and trusted brands to pressure finance teams into making fraudulent Automated Clearing House (ACH) payments of nearly $50,000.

The campaign primarily targeted organizations in the United States, which received 87.7% of all messages. IT services, business advisory firms, consumer-goods companies, and other enterprises were among the targeted sectors.

Unlike standard invoice scams, the attackers combined several social-engineering techniques in one email.

They impersonated CEOs, CFOs, and presidents of the victim organizations, using executive names in sender display names, Reply-To fields, and email signatures.

The messages included a short approval notice directing accounts payable employees to process an invoice.

To make the request appear legitimate, the attackers added a fabricated ServiceNow annual-subscription invoice and fake forwarded email conversations.

Microsoft stated that ServiceNow and the other organizations named in the campaign were not compromised or involved. (Varshini / Cyber Press)

Related: Microsoft, Petri

Attack chain showing domain registration, executive impersonation, invoice fraud delivery, ACH payment execution, and financial theft. Source: Microsoft.

The White House plans to apply the approach behind its Texas water cybersecurity pilot to other critical infrastructure sectors in the near future, National Cyber Director Sean Cairncross said.

Speaking at the Billington Cybersecurity Summit on stage with OPM Director Scott Kupor, Cairncross said his office would announce the additional sectors soon as it evaluates how to expand Project Watershed 250, which pairs water utilities with government and industry cybersecurity assistance. He did not identify the sectors or give a launch date for the additional efforts.

The goal is to “find a specific, concrete solution, and then scale off of that,” he said.

The water pilot, launched Aug. 31 in San Antonio, brings together the Office of the National Cyber Director, Texas Cyber Command, the office of Texas Gov. Greg Abbott and private companies to identify vulnerabilities in water systems and help utilities address them.

Cairncross said water was a starting point because providers often lack the funding and technical capacity to adequately defend their systems. The pilot is intended to test whether industry partnerships can lower costs, deploy new technology and improve operators’ ability to protect their facilities, he said. The White House had been exploring pilot programs for multiple infrastructure sectors with private companies for some time.

Separately, Cairncross said his office is developing a cyber academy proposal that would connect venture capital and other private sector initiatives with a service component intended to attract people into cybersecurity work, a topic he has discussed in several other public appearances. He said the office would have more to share soon without providing a timeline. (David DiMolfetta / NextGov/FCW)

Related: CyberScoop, MeriTalk

LinkedIn beat two lawsuits over its practice of scanning users’ browser extensions, with a judge granting the Microsoft subsidiary’s motion to dismiss the cases.

The users who sued LinkedIn failed to adequately allege that they have standing to sue because neither asserted that they “had browser extensions installed that conveyed private information to LinkedIn,” ruled Judge Vince Chhabria in US District Court for the Northern District of California.

In his ruling, Chhabria gave the plaintiffs leave to amend their complaints but said he doubts they can make a plausible case. “Given LinkedIn’s further arguments that users voluntarily download browser extensions, which by their nature intentionally expose data to websites, it seems unlikely that the plaintiffs will ever be able to allege a privacy violation, much less prevail at the end of the day,” Chhabria wrote.

California residents Nicholas Farrell and Jeff Ganan separately filed class actions against LinkedIn in April, seeking to represent themselves and other LinkedIn users. Ganan’s attorney, J.R. Howell, said he is evaluating whether to bring the claims in a California state court, which has different requirements on standing, or to appeal the US district court ruling in the US Court of Appeals for the Ninth Circuit. (Jon Brodkin / Ars Technica)

Related: Court Listener, Bloomberg Law

Best Thing of the Day: We Promise This Is the Last Hugging Face Post-Mortem

Melanie Mitchell, professor at the Santa Fe Institute, offers an accessible and common-sense post-mortem of the Hugging Face incident, noting that the leading factor triggering the incident was poor cybersecurity measures.

Worst Thing of the Day: Not Even CAPTCHAs Can Save Us Now

Anthropic reports that Mythos 5 is the first model to solve the CAPTCHA process.

Closing Thought