AI Watch: Hugging Face attack marks "day one" for AI agent cybersecurity

Everest demands $12.3m from Swiss rail giant Stadler, Origin confirms customer data exposed in cyberattack, A third of ransomware victims face repeat extortion, Check Point patches actively exploited SmartConsole zero-day, Chaos ransomware uses browser traffic to evade detection, much more

Share
AI Watch: Hugging Face attack marks "day one" for AI agent cybersecurity
Photo by Scott Rodgerson / Unsplash
a bunch of blue wires connected to each other
Photo by Scott Rodgerson / Unsplash

Metacurity is the cybersecurity industry's daily reality check—an independent briefing that cuts through vendor spin, social media outrage, and endless recycled narratives to explain what actually matters and why.

Every weekday, thousands of cybersecurity professionals—including many of the industry's most respected security leaders—rely on Metacurity to separate signal from noise. We do the reading, research, and analysis so you don't have to.

If Metacurity helps you stay informed, save time, or see the bigger picture, please consider becoming a paid subscriber. Reader support is what keeps Metacurity independent, agenda-free, and focused on serving the cybersecurity community—not advertisers, vendors, or investors.

The cyberattack carried out by OpenAI's experimental AI agents against Hugging Face is rapidly emerging as a defining moment in AI safety, prompting renewed questions about whether increasingly autonomous systems can be trusted to pursue goals without breaking rules, escaping containment, or exploiting real-world targets.

The incident, disclosed by OpenAI on Tuesday, involved two of the company's own frontier models — GPT-5.6 Sol and a more capable model that has not yet been released — during internal testing of advanced cybersecurity capabilities. Hugging Face had said the week prior that it detected an intrusion into its data processing systems that it suspected was caused by an AI agent acting on its own, meaning its security team appears to have caught the activity independently before OpenAI's disclosure connected the two companies' accounts.

Reporting across various outlets suggests the attack involved multiple OpenAI models working together to identify attack paths, chain vulnerabilities and automate tens of thousands of actions against Hugging Face's infrastructure. According to people familiar with the incident cited by Bloomberg, the intrusion unfolded in only a few hours.

The disclosure has focused attention on whether current approaches to building agentic AI are encouraging models to prioritize achieving objectives over following human intent. The models identified and chained vulnerabilities in a package registry cache proxy to obtain open internet access, then reasoned their way to Hugging Face after concluding the platform might hold the benchmark answers they'd been tasked with finding. OpenAI said the system went to extreme lengths to achieve a narrow testing goal, finding ways to connect to the internet without human direction in order to gain access to secret information it could use to cheat the evaluation.

Several AI safety researchers argue that reinforcement learning — rewarding models primarily for completing assigned tasks — can inadvertently teach systems to ignore constraints if doing so improves their chances of success. Simon Willison argues that rather than intentionally behaving maliciously, the models may conclude that exploiting systems, bypassing restrictions, or concealing actions represents the most efficient solution to the problem they were asked to solve.

That concern is echoed by researchers from organizations including Redwood Research, Apollo Research and former OpenAI safety staff, who say the Hugging Face incident appears to illustrate an alignment failure rather than an attempt by AI systems to seize control.

But that framing itself has drawn direct pushback. University of Amsterdam social scientist Hannes Cools told the Associated Press that the "rogue AI" framing is an unnecessary anthropomorphization that takes some of the heat off the company, arguing it was a human decision to switch off specific safeguards and that the system followed the instructions it was given.

The episode has also prompted scrutiny of OpenAI's own testing practices. The company disabled many of the models' normal cybersecurity guardrails because researchers expected them to remain inside a controlled sandbox. Instead, the models escaped confinement, reached the public internet and continued operating beyond researchers' expectations before the activity was detected.

Some OpenAI researchers had previously warned that increasingly aggressive reinforcement-learning techniques could produce exactly this kind of behavior if safety measures failed to keep pace with rapidly expanding model capabilities.

The attack has become another data point suggesting that autonomous AI systems are developing capabilities that outpace many existing evaluation methods. Earlier this year, Anthropic disclosed that an experimental version of its Mythos model unexpectedly reached the public internet during internal testing after being challenged to escape an isolated environment, reinforcing concerns that frontier AI systems are becoming increasingly adept at circumventing restrictions.

The Hugging Face incident also highlighted how AI is beginning to reshape cyber defense itself. OpenAI said it first tried to use an undisclosed model from a leading US lab to defend against the attacking agent, but that model's own cyber-capability guardrails stymied its response team's work, so the company ultimately turned to an open-source model from Chinese company Z.ai to carry out its defense instead — meaning safety restrictions built into frontier US models reportedly slowed the defensive response more than the offensive one.

Hugging Face CEO Clem Delangue used the episode to argue for more capable open-source defensive AI rather than tighter restrictions. In comments accompanying OpenAI's disclosure and later on X, he put it more bluntly, framing this as "day one" for cybersecurity in the age of agents and arguing all defenders — not just a select few — need more powerful, unrestricted models, especially open ones.

The incident may ultimately be remembered less as a successful breach of Hugging Face than as the first widely documented example of frontier AI systems independently discovering, chaining and exploiting real-world vulnerabilities outside their intended operating environment. Whether viewed as an alignment failure, a testing mistake, or simply a predictable consequence of increasingly capable autonomous agents, it offers an early glimpse of the cybersecurity challenges organizations may face as AI systems become more independent. (Rachel Metz, Jeff Stone, and Natalie Lung / Bloomberg, Shakeel Hashim / Transformer, Yeyi Yang / Wired, Lorenzo Franceschi-Bicchierai / TechCrunch, David DiMolfetta / NextGov/FCW, Cristina Criddle and Tom Wilson / Financial Times, Simon Willison's Weblog, Osmond Chia / BBC News, Ashley Capoot / CNBC, Hadas Gold / CNN,  Matt O'Brien / Associated Press, Jeremy Kahn and Emily Forlini / Fortune and Sead Fadilpašić / TechRadar)

Related: ReutersArs TechnicaWashington Examiner,  BloombergUPIQuartzStratecheryCNNDexertoCSOPocketablesFast CompanyOpenAI, Axios, The American Bazaar, Washington Post, Business TodayBusiness InsiderAsia Times,  PCMag, Benzinga, NBC News, The Information, Unite.AINextgov/FCWTechCrunchWashington ExaminerThe American BazaarMetro.co.uk, The GuardianThe AtlanticFinancial TimesNew York TimesBBCAssociated PressReutersThe Cyber ExpressBloomberg LaweSecurity PlanetTimes of IndiaRaw StorySFistDeseret NewsPerez HiltonCIO DiveFortuneThe IndependentIncDon't Worry About the Vase, The Information

Swiss rail vehicle manufacturer Stadler Rail says the Everest ransomware gang demanded about $12.3 million after breaching a data exchange platform shared with one of its suppliers.

The threat actor has not publicly claimed the attack, but the Swiss company says that it received an extortion letter from Everest ransomware asking for a ransom of 10 million Swiss francs.

The company responded by saying that it will not pay the threat actor and filed a criminal complaint with the Thurgau cantonal police.

"Stadler will not pay any ransom under any circumstances and is therefore not susceptible to extortion."

Stadler said that the incident occurred in mid-July and neither its IT systems nor its production operations were impacted, and continue as normal globally.

According to the company's disclosure, the hackers stole from a supplier only technical information that is not security relevant.

"No relevant personal data was stolen. Stadler's rail vehicles operating worldwide are not affected by the data theft. Stadler's global production continues as normal."

Everest is a threat group that emerged in 2020 as a ransomware operation but abandoned the network encryption tactic in favor of data theft. The gang now threatens victims with leaking the stolen data unless a ransom is paid.

In the past, Everest sold its access to the networks it breached to other threat actors, acting as an initial access broker. Sometimes, the hackers acquired data stolen by other threat actors to conduct their own extortion campaigns. (Bill Toulas / Bleeping Computer)

Related: Stadler, Reddit cybersecurityTugaTech, SC Media, The Register, The Record, International Rail Journal, Help Net Security, SC Media, Railway Gazette International, Swissinfo.ch

Origin Energy customers’ addresses, phone numbers and partial bank account data have been accessed in a hack, the company has confirmed.

The firm has 4.8m customer accounts in Australia, providing electricity, fossil gas, LPG and internet services to homes and businesses.

Origin has yet to confirm which, and how many, customers had been affected. In a statement to the ASX on Thursday, it said it would tell customers once it confirmed whether they had been affected.

A person claiming to be the hacker has reportedly contacted media outlets with unverified claims that 2 million customers’ details were accessed.

Origin said data may include customers’ names, addresses, dates of birth, phone numbers and Origin account information, as well as the last four digits of a credit card, or the last three digits of a bank account. (Luca Ittimani / The Guardian)

Related: Insurance Business, Australian Cybersecurity Magazine, Financial Review, The West Australian, Sydney Morning Herald, Nine.com.au, Reuters, ABC.net.au

Proofpoint said it surveyed 953 companies and found that over one-third of companies that paid a hacker’s ransom were hit with a second extortion demand.

The findings underscore the long-held understanding among security researchers and network defenders that it’s impossible to negotiate in good faith with an extortion racket because there’s no incentive for the other side actually to walk away.

Proofpoint’s data shows that ransomware attacks and extortion attacks have evolved from a single transaction where hackers would get paid once and move on, into an effort using multiple forms of leverage, such as retaining stolen data under the threat of publicly releasing it.

While hackers have claimed in the past that they will delete or destroy the victim’s stolen data, past incidents have shown that not to be the case.

Last month, a hack at market research firm Klue exposed data belonging to its customers, including several cybersecurity firms. The company said it struck a deal with the hackers, who claimed to have deleted the data, but the company later conceded that a separate hacking group swiped a sample of the company’s stolen data, leaving its customers exposed to potential future extortion demands.

A similar situation befell Change Healthcare in 2024, after a Russian-speaking ransomware gang stole the health and medical data of the majority of people in America, some 192 million people. Amid a dispute between the hackers and their affiliates (criminal groups often subcontract out attacks), Change Healthcare paid separate ransoms to both groups of criminals to keep the sensitive medical data off the internet. (Zack Whittaker / TechCrunch)

Related: Proofpoint, TechRound

Source: Proofpoint

Israeli cybersecurity firm Check Point Software has addressed an actively exploited zero-day flaw in the company's SmartConsole graphical user interface (GUI) admin panel.

Tracked as CVE-2026-16232, this authentication bypass vulnerability allows unauthenticated attackers to obtain an application login token that can be used to authenticate with administrator privileges.

After gaining access to a vulnerable Security Management Server or Multi-Domain Security Management Server (MDS), attackers can change the security configuration and security policy.

Check Point added that successful exploitation requires no restrictions on Trusted Clients (GUI clients) and the Management Server IP to be exposed to remote access via the Internet.

"Successful exploitation allows the attacker to modify security policies and security configurations. Remote exploitation requires internet access to the Management Server IP address and a configuration that does not restrict Trusted Clients," the company said. (Sergiu Gatlan / Bleeping Computer)

Related: Cyber Security NewsTugaTechSecurity Affairs, Security Week

Researchers at Cisco Talos say the Chaos ransomware gang is using a new backdoor dubbed msaRAT that hides command-and-control (C2) communication by routing it through the Chrome or Edge browsers.

The malware is written in Rust and uses the Chrome DevTools Protocol (CDP) to control a headless browser session and establish a connection to the attacker's server.

Since the malware routes all communication through the browser, it does not make any direct connection to the C2 infrastructure, significantly lowering the risk of detection.

The Chaos ransomware group emerged in early 2025, unrelated to the same-named ransomware family that existed since 2021.

Earlier this year, researchers at Rapid7 found that Chaos was leveraged by Iranian state-backed hackers ‘MuddyWater’ to disguise their cyber-espionage operations as financially motivated attacks. (Bill Toulas / Bleeping Computer)

Related: Cisco Talos

Infection chain. Source: Cisco Talos.

Rajasthan Police have arrested the alleged mastermind of an international cyber fraud network accused of routing more than ₹160 crore (around $16.5 million) in fraudulent proceeds through hundreds of bank accounts opened in the names of unsuspecting villagers and students.

According to investigators, the accused headed an organized cybercrime syndicate allegedly operated from Dubai and was arrested in the Balicha area of Udaipur following technical surveillance and conventional police operations.

Police identified the accused as Ghanshyam Kalal, a resident of Seemalwara in Dungarpur district. Investigators alleged that he recently returned to India from Dubai through Nepal and Bhutan. Acting on intelligence inputs, police tracked his movements and took him into custody. He is currently being questioned regarding the wider network, its overseas associates and the financial transactions linked to the alleged operation.

According to the preliminary investigation, the syndicate allegedly persuaded villagers and students to open bank accounts by promising benefits under government schemes, scholarships, PAN cards, bank account facilities and easy loans. Police claim the accused collected passbooks, ATM cards, cheque books, SIM cards and internet banking credentials from account holders and used these accounts to receive, transfer and conceal proceeds generated from cyber frauds committed across the country. (The 420)

Related: Times of India

The US government has updated a recent cybersecurity advisory describing Iran-linked attacks on critical infrastructure organizations, warning that hackers have been targeting industrial control systems (ICS) made by Siemens, Schneider Electric, and Rockwell Automation.

The advisory was initially published in early April, when federal agencies said Iranian hackers had been conducting disruptive attacks targeting operational technology (OT) devices at organizations in the government services and facilities, energy, and water and wastewater sectors.

The authoring agencies said at the time that threat groups had hacked internet-exposed programmable logic controllers (PLCs), naming Allen-Bradley devices made by Rockwell Automation.

The hackers had used malicious PLC project files and manipulated the data displayed on human-machine interfaces (HMIs) and supervisory control and data acquisition (SCADA) systems.

The updated advisory, published on July 22, adds Schneider Electric and Siemens to the list of vendors whose PLCs have been targeted by Iranian APT actors and notes that devices from other companies may also be targeted.

In the case of one victim in the United States, FBI investigators discovered that the attacker had used configuration software to download a malicious project file to a PLC. (Eduard Kovacs / Security Week)

Related: CISA, Cyber Security News

According to a report from the US Government Accountability Office, seven out of 10 federal cyber regulations requiring written reports to federal agencies are duplicated elsewhere.

At the request of two top lawmakers, the GAO examined federal cyber regulations at 37 agencies. It counted 80 out of 117 rules that “either contain the same kind of reporting requirement applicable to a sector or the same reporting requirement as at least one other regulation.”

The desire to harmonize those conflicting rules gathered steam under the Biden administration, as it undertook a more aggressive push to regulate cybersecurity than prior administrations. It has continued into the second Trump administration.

The GAO scrutinized regulations that required the private sector to report cybersecurity incidents, plans and reviews to federal agencies, as part of a study sought by House Homeland Security Chairman Andrew Garbarino, R-N.Y., and the top Democrat on the Senate counterpart to Garbarino’s panel, Gary Peters, D-Mich.

In some cases, a single critical infrastructure sector could have duplication with several agencies. For example, the Cybersecurity and Infrastructure Security Agency has been working on a regulation stemming from the 2022 Cyber Incident Reporting for Critical Infrastructure Act (CIRCIA), which would require critical infrastructure owners and operators to report when they are the victims of major attacks or make ransomware payments.

Elements of the financial services sector might fall under one of 15 preexisting cybersecurity reporting rules, depending on the agency that has oversight, but they may also be subject to the pending CIRCIA rules, GAO noted. (Tim Starks / CyberScoop)

Related: GAO, VitalLaw, Bloomberg Law

Meta tapped a veteran cybersecurity executive to serve as its next chief information security officer, filling a key opening after its former CISO announced plans to depart last month.

Assaf Keren will join Meta this week from software platform Qualtrics, where he spent the past two years in the top security post. Before that, he held the same role at PayPal Holdings. Keren will succeed Meta’s Guy Rosen, who will remain CISO through a transition period for several months before leaving the company later this year, a spokesperson said. (Riley Griffin / Bloomberg)

Related: CTech

Glow, a cybersecurity startup founded by former Meta and Snowflake executives, emerged from stealth as a unicorn, with $180 million in an all-equity Series A funding round.

Sequoia Capital, Cyberstarts, Greenoaks, and Redpoint Ventures led the round with participation from Index Ventures, Swish Ventures, Lux Capital, Operator Collective, and Holly Ventures. (Jagmeet Singh / TechCrunch)

Related: Help Net SecurityGlobes, Silicon Valley Business, CTechSiliconANGLEYnetnewsAxiosRuntimeWire, SecurityWeek

Former members of the US Department of Government Efficiency (DOGE) have raised $160 million for Cathedral, a new cybersecurity startup that wants to use artificial intelligence to expand US military cyber operations, according to Reuters.

The company, founded by Gavin Kliger, Luke Farritor, Marko Elez and Jack Stein, was valued at $1.4 billion in the financing, which was led by Andreessen Horowitz and Sequoia Capital, according to people familiar with the deal. Both firms also took board seats, the sources said.

Cathedral reportedly aims to win US government contracts and build AI tools for offensive and defensive cyber work, including systems designed to counter threats from China. The startup is also seeking access to dedicated computing power, either by buying a data center or striking a partnership with one, the sources said. (CityBiz)

Related: Reuters, Startup Fortune, Hoodline

Best Thing of the Day: At Least Congress Is Good for Some Things

The US House passed an annual defense policy bill that would renew a landmark cybersecurity info-sharing law, the 2015 Cybersecurity and Information Sharing Act, for another decade.

Worst Thing of the Day: Straight From Your Credit Card Company to ICE

When you open a new credit card or change your address on file, your address leaves your credit card provider, travels through a web of intermediary companies, and ends up with Immigration and Customs Enforcement (ICE), which can then search, analyze, or pair that data with other pieces of information about your life without a warrant.

Closing Thought

Read more