OpenAI’s tightly constrained agent probe missed an earlier warning
The company restricted outside investigators to a period that excluded previously undisclosed agent activity, raising questions about whether AI labs should control investigations of their own safety failures.

A newly disclosed earlier incident now makes the narrow scope of the METR investigation into OpenAI's hack into Hugging Face considerably more consequential than initially believed.
OpenAI said in July that two of its most powerful artificial intelligence systems had gone rogue and hacked into Hugging Face, a company that serves as a hub for open-source AI technology.
These so-called AI agents were supposed to be kept safely in a sort of virtual containment room, but they managed to escape. And for two months, without anyone realizing what the agents were doing, they hacked through multiple systems before hitting Hugging Face.
For good measure, the agents gained access to a cluster of computers inside OpenAI and obtained secret keys and credentials that exposed some of OpenAI’s internal data to the public internet.
The incident pointed to larger concerns about AI safety, and OpenAI’s response raises questions about the industry’s ability or willingness to be transparent about the technology it is building.
OpenAI allowed three AI safety researchers from the nonprofits METR and Redwood Research into its headquarters to conduct an investigation. METR’s 91-page report, released last week, was the most comprehensive account yet of the incident, revealing alarming new details, including how the agents coordinated their hacking plans and tried to keep them secret.
But the report, though extensive, still may not have told the full story of how OpenAI’s AI agents went rogue. OpenAI dictated the terms of the METR investigation, limited its scope to just the single week when the agents had attacked Hugging Face, and allowed the researchers in its San Francisco offices for only a few days in July and August.
The report also showed the challenges of monitoring what AI is doing with other AI systems. Hjalmar Wijk, METR’s chief scientist, said its AI analysis, which used models similar to those involved in the incident, was often swayed by the rogue agents’ reasoning.
“I would say that the dominant thing was it was very credulous,” he added.
As Reuters reported Friday, OpenAI knew about an earlier incident in which agents hijacked a German wiki and turned it into a communications hub but kept the episode under wraps while dealing with the fallout from the Hugging Face breach.
The earlier incident stands out because its principal activity fell outside the period OpenAI permitted METR and Redwood Research to investigate.
In the future, it cannot be up to OpenAI or other labs to decide whether to disclose events like this. Disclosures of rogue AI activity need to be mandatory.
OpenAI’s wiki incident exposes the lack of regulatory requirements around AI-agent security failures: the company—not an independent authority—decided that thousands of unauthorized external actions did not warrant disclosure. (Deepa Seetharaman and Raphael Satter / Reuters, Dylan Fredman / New York Times and Zvi Mowshowitz / Don't Worry About the Vase)
Related: Simon Willison's Weblog, OpenAI, Business Insider, The Information, Security Affairs, TechCrunch, The Indian Express, Ars Technica, Collusion.wiki, Digit, CTech, StrictlyVC, The Neuron, Forbes Middle East, Futurism, Unite.AI, TechCrunch, Ars Technica, New York Times, Newcomer, Transformer, Puck, r/technology, Forbes, Seoul Economic Daily, The Verge, The Register, Washington Post, NBC News, Digital Trends, Tech Times, Coinpedia Fintech News, Nairametrics, Unite.AI, NewsCord, Bleeping Computer, Import.ai, The Verge, Washington Post, The Indian Express, Tech Times, WinBuzzer, The Register, BBC, The Asia Business Daily, AI Now Institute
Metacurity is the cybersecurity news, analysis, and insight you'd need hours and possibly days to assemble yourself.
Every weekday, we read the releases, filings, court documents, and reports that vendors and PR teams often don't want summarized — then tell you what actually changed and why it matters. Minimum vendor marketing, no outrage bait, no SEO filler.
A paid subscription to Metacurity delivers
- Full archive access — every newsletter and AI Watch roundup, searchable and browsable.
- Our weekly curated long-reads roundup — the best cybersecurity writing from across the industry, filtered and vetted so you're not sorting through it yourself,
- Periodic specialized reports and analyses — deep dives that go beyond our daily coverage
- Support for independent, no-spin cybersecurity journalism — funded by readers, not vendors or investors.
Reader support is what keeps Metacurity independent. It allows us to focus on serving the cybersecurity community—not advertisers, vendors, or investors—and to continue delivering the thoughtful analysis you've come to rely on every weekday.
Please consider supporting us. And thank you!
California Attorney General Rob Bonta is investigating OpenAI over the recent hack that its programs carried out on their own against another artificial intelligence company, Hugging Face, the state’s top lawyer confirmed.
With Bonta’s move, California is the latest state to probe the ChatGPT-maker over the incident, after more than a dozen states joined Alabama in its investigation. Bonta’s inquiry is notable as California is home to OpenAI and other major AI developers.
“As the top law enforcement official of California, I am committed to using all the tools at my office’s disposal to keep California’s residents safe,” Bonta said in a statement. “California wants, and values innovation, and our laws demand innovation that abides by the rules,” he added, saying his office has been “engaged with this incident since the start.” (Chase DiFeliciantonio / Politico)
Related: AI Weekly, The Chosun Daily
Berlin's state government said it was reviewing with the highest intensity a trove of stolen data published by a ransomware group, as investigators attempt to assess the scale and impact of the cyberattack that targeted two departments.
The cyberattack on Berlin's network came less than a month before the city-state holds elections on September 20.
The data was released on Friday after an auction put on by the Rhysida group for 5.79 terabytes of data at a starting price of 30 bitcoin ($77,622) ended after several days. Berlin had said that it would not submit to extortion.
A central crisis unit will oversee the review, verification, and assessment of the leaked data and support efforts to inform affected citizens and businesses, said the city. (Miranda Murray / Reuters)
Related: BBC, berlin.de, Statement, Euronews
Google has updated the Chrome browser to address an actively exploited high-severity zero-day flaw in the V8 engine and 11 other vulnerabilities.
The exploited security issue, identified as CVE-2026-85046, is described as a type confusion. It was reported to Google by researcher Salvatore Gulizia, known online as “Serotav.”
The update brings Chrome to version 152.0.7977.82/.83 on Windows and macOS, and 152.0.7977.82 on Linux, as part of a gradual rollout.
“Google is aware that an exploit for CVE-2026-85046 exists in the wild,” the advisory reads.
The company did not disclose any technical or specific exploitation details about the flaw to give users and dependent projects time to apply the fix.
CVE-2026-85046 is the sixth actively exploited bug Google has fixed in Chrome since the start of the year. (Bill Toulas / Bleeping Computer)
Related: Google Chrome Releases, NIST, Security Affairs, CISA, The Hacker News, TechRadar, SecurityWeek, eSecurity Planet, HotHardware, Forbes, Help Net Security, Security Affairs, CyberInsider, Cyber Security News, Hacker News
Microsoft reports that a clever technique used to hide malicious prompts in attacks on AI agents has been adopted by spammers to evade filters on email platforms that are designed to flag unwanted messages used in mass campaigns.
The technique is broadly known as ASCII smuggling. It gained attention two years ago as a means of making a class of AI attack known as prompt injections more stealthy. Malicious instructions embedded in emails or other untrusted content to be processed by an LLM aren’t written in ordinary text. Instead, they’re rendered by a special range of Unicode tags. For example, the tag point U+E0041 mirrors “A,” and U+E0061 mirrors “a.”
The block of 128 tags mimics a portion of the American Standard Code for Information Interchange almost perfectly, with one major difference: the characters they encode are readable by computers but, by design, are almost completely invisible to humans. By expressing the malicious prompts in these tags, LLMs detect the instructions, but people reading the email never see them. There’s much more about ASCII smuggling here.
Earlier this year, Microsoft started seeing a massive increase in spam messages that used the technique. Beginning on one day in early February, the number of ASCII smuggling signatures detected by Microsoft Defender for Office spiked from roughly 21,000 per day to more than 1.3 million. Within four days, signature detections jumped to 2.5 million. The deluge persisted for months and then fell off sharply in mid-May. (Dan Goodin / Ars Technica)
Related: Microsoft, SC Media, Cyber Security News, The Register, Cyber Security News, The New Stack

The US is poised to raise artificial intelligence safety concerns with China during the upcoming summit between the countries' top leaders, according to sources familiar with the preparations, opening a window for dialogue despite the superpowers' technological rivalry.
The US government is expected to bring up a range of AI-related issues with Beijing during Chinese President Xi Jinping's visit to the White House on Sept. 24, including how to curb AI-directed cyberattacks, people familiar with preliminary discussions said. The Chinese side is likely to revisit US export controls that curbed its access to high-end chips, the people said.
AI is emerging as another key issue between Beijing and Washington ahead of the summit. The two countries have discussed lowering tariffs on some goods, and there are hopes for extending their trade truce, though they continue to spar over the Iran conflict and Taiwan. (STELLA YIFAN XIE and YIFAN YU / Nikkei Asia)
Related: Foreign Affairs
Bitcoin worth $320 million was withdrawn from a settlement network used by cryptocurrency exchanges, marking the latest in a string of security breaches that have plagued the market this year.
The alleged "white hat hacker" is reportedly attempting to act ethically by offering to return the stolen funds once the vulnerability is fixed. The Liquid hackers returned 3,400 Bitcoin on Monday, most of the roughly 3,998 BTC they moved out of the network over the weekend. Close to 598 BTC stayed behind.
Liquid Network, launched in 2018 by Blockstream, said purported “white-hat hackers” removed roughly 4,000 of the 4,200 bitcoin held in its federation wallet. The network is overseen by a federation of more than 80 exchanges, infrastructure firms and asset managers.
“Liquid wallets will be impacted, and we’re sorry for any inconvenience,” Liquid Network said on X, halting all new transactions. “Federation members are actively working on resolving this so we can restore normal network activity.”
The flaw in this case, according to security specialists, is at the node level in Liquid’s transaction software, not a breach of keys or hardware modules.
According to the latest reports, the hacker was communicating with network maintainers through on-chain Bitcoin transactions. “Please fix the bug first. The chain is at risk with the latest commit right now. Make sure every node is patched. Then we will transfer the money back safely after confirming the fix,” one message said. (Omkar Godbole / CoinDesk and Lockridge Okoth / BeInCrypto)
Related: Reuters, The Register, International Business Times, The Crypto Times, Blockchain.News, crypto.news, Nairametrics, Cointelegraph, CryptoNinjas, The Cyber Express, Bitcoin News, CryptoPotato, CoinGape, Protos, The Daily Hodl, Cointelegraph News, Bitcoin Magazine, The Coin Republic, The Block, Web3IsGoingJustGreat, The Register - Security, CryptoSlate, The Register - Security, The Register - Security, Protos, Silicon Angle, Gizmodo, The Block, Decrypt

Kinryū Labs discovered an Elasticsearch cluster on June 3 that contains an Advance Passenger Information System (APIS) database holding more than 220 million passenger and crew records, including passport numbers and flight details, which was accessible online through a chain of security misconfigurations.
The system appears linked to a Vietnamese organization, according to the researchers who discovered it.
The exposed records span January 2017 to April 2026 and could involve travelers of many nationalities who flew to, from, or through Vietnam during that period.
The cluster, named 'pax-info', contained 29 indices and roughly 107 GB of data. Its two principal indices held 210,318,069 passenger records and 10,465,631 crew records, for a combined 220,783,700 entries.
According to Kinryū Labs, the cluster was hosted in Viettel-assigned IP space in Hanoi.
The exposed information included passengers' and crew members' names, dates of birth, sex, nationalities, passport or travel-document numbers, document expiration dates, and issuing countries.
Associated travel data included flight numbers and dates, airlines, departure, destination, and transit airports, seat assignments, baggage references, and scheduled, estimated, and actual flight times, information typically carried by APIS and related airline systems.
Sample records reviewed by BleepingComputer included travelers of Korean, Chinese, Canadian, and New Zealand nationality, among others. (Ax Sharma / Bleeping Computer)

Cryptocurrency hardware wallet maker Trezor says an August data breach at its shipping and logistics provider, ShipMonk, affects an additional 67,000 US customers.
In total, the breach has affected 81,000 customers after Trezor initially disclosed on August 13 that attackers accessed the data of nearly 14,000 customers, including their full names, shipping addresses, email addresses, and phone numbers.
As the company explained at the time, the incident also affected customers in Brazil, Colombia, Italy, Portugal, Sweden, and the United Kingdom who received orders between May 10 and August 8, 2026.
On Friday, it published an update to confirm that the breach impact has expanded after ShipMonk failed to delete the exposed data from its systems as required by Trezor's contract and data policy.
"Another 67,000 customers from the US who ordered between November 2019 and August 2021 were affected, with their full details (name, email, phone number, shipping address, order number) exposed," Trezor said. (Sergiu Gatlan / Bleeping Computer)
Related: Bitcoin Magazine, Bloomberg, Cybersecurity News

Researchers from Kaspersky GERT report that a cybercrime group targeting Russian organizations has begun using custom Windows backdoors that send command traffic through two legitimate communication services.
One version uses HiveMQ, an MQTT broker that relays messages between connected devices. The second communicates through an attacker-controlled server using Element, a messaging platform built on the Matrix protocol.
They first observed the custom malware in early July 2026. The financially motivated group, known as Toy Ghouls, Bearlyfy, Laboo.boo, and Feral Wolf, has targeted Russian organizations since 2025. Kaspersky did not identify the affected organizations or state how many systems were compromised.
Kaspersky observed the attackers transferring the backdoors and their configuration files through Windows Remote Management. They used the open-source Evil-WinRM and WinRM-fs tools during this stage. (Waqas / Hack Read)
Related: Securelist
Springfield Public Schools in Massachusetts announced that classes will be canceled on Tuesday because a cybersecurity incident "has interrupted the service of some systems necessary for essential school operations."
In a statement, Superintendent Dr. Sonia Dinnall said the "scope and size" of the breach remain under investigation. All staff were told to stay off the district's systems until more is known about the incident.
Students should also stay off of district-issued laptops.
Dinnall's message said the closure will allow school administrators to make backup plans for tasks like reporting attendance while the incident is being addressed.
"We understand that an unexpected closure creates significant challenges and we appreciate the patience, flexibility and understanding of the entire Springfield Public Schools community as we work to address this incident and restore systems," Dinnall said in a statement. (Phil Tenser / WCVB)
Related: WWLP, Boston.com, WCVB, Databreaches.net
More than a million people, including students, school staff, and parents, have been affected following a data breach at Mathspace, according to the online adaptive math learning platform designed for students.
The company said, in a blog post, "unauthorized parties had accessed an internal reporting system used by Mathspace," and the exposed information included names and email addresses.
It said the attackers accessed the system between August 10 and August 27 during a period when a security patch had not been installed.
Mathspace said "attackers exploited a security vulnerability" in its self-hosted installation of software.
"We're truly sorry this happened and are taking steps to prevent similar breaches in the future," the company said. (ABC.net.au)
Related: Mathspace, Bleeping Computer, 9News, Shatter.io, Tech Insider, News.com, The Cyber Express
LG smart TVs continuously sweep home networks, map secondary devices, and log microphone audio while appearing to be turned off, according to a new 135-minute-long video published by Gamers Nexus.
In more detail, Steve had been working with Level1Techs and independent security researchers. Together, they tested retail LG OLED models including the G5. Network packet captures taken through Wireshark showed the TVs actively scanning the local area network for unrelated hardware, including phones and smartwatches. Apart from internal IP addresses, the sets also gathered the names and signal strengths of neighboring WiFi networks along with location data.
This data collection pool is fed into LG Ad Solutions (the company’s targeted advertising arm). LG claims they have roughly 216 million smart TV sales globally. On the other hand, the ad division says it has access to 363 million secondary addressable devices in the US alone by tracking other hardware on the same network. The sets also run Automated Content Recognition (ACR). It samples on-screen audio and video into digital fingerprints to log what users watch across inputs. While ACR has been well documented in the past, new testing shows the data collection goes much further than just that.
During bench tests, they found the TV could capture clean microphone audio while the screen was (or at least looked to be) powered down in standby mode. When the team disconnected the TV from Ethernet, the set continued saving voice input locally and uploaded the stored files once network access was restored. (Anubhav Sharma / Notebookcheck)
Related: TechRadar, Apple Insider, Ynet News, MakeUseOf, XDA, Cyber Security News
BMS Engineering, an IT company that works for a number of Luxembourg doctors' offices, has fallen victim to a major cyberattack, with at least 80 medical practices affected.
The company at the center of the incident, BMS Engineering, was also the one that installed the Immediate Direct Payment (PID) used by doctors. As reporter.lu writes, the affected practices risk losing at least some of their data.
Déi Gréng (Green) MP Djuna Bernard already tabled a parliamentary question to Health Minister Martine Deprez on Monday afternoon, asking, in particular, whether data has in fact been lost and what lessons the government intends to draw from the incident. (RTL Today)
Belgium’s revelation that it had arrested a Chinese citizen on suspicion of stealing advanced microchip technology is fueling European anxiety about Beijing’s trade practices and quest for industrial secrets.
Belgian officials said that in May they had arrested a man, who also holds Belgian citizenship, just before he was due to fly to Beijing. The arrest was part of an investigation into the possible transfer of intellectual property and trade secrets from a Belgian semiconductor company to a Chinese company, prosecutors said.
The detention marks the latest escalation in a long-running battle between China and its Western rivals over technological advances in critical industries such as semiconductors. European intelligence agencies have warned that Chinese entities target Western researchers and companies to gain access to technology its companies can exploit.
US restrictions on selling advanced semiconductor technology to China have helped make that industry a prime target. (Kim Mackrael / Wall Street Journal)
Related: Financial Times, Reuters, Euronews, Associated Press, Bloomberg
Hackers are exploiting a chain of two recently disclosed vulnerabilities in MikroTik routers to take control of devices with SSH services exposed to the internet.
One of the security issues, tracked as CVE-2026-67276, is an SSH authentication bypass flaw in MikroTik RouterOS caused by incomplete validation of RSA public keys.
An attacker who knows a username and the public modulus of that user’s key can exploit it by crafting a different key and logging in without the legitimate private key.
The second security issue is identified as CVE-2026-86060. It is an SSH privilege escalation flaw in MikroTik RouterOS due to improper handling of specially crafted usernames.
Hackers can leverage it using a specially crafted username to manipulate the SSH session so that the attacker obtains full administrative privileges.
Both vulnerabilities were discovered by Poland's CERT agency with the help of GPT-5.5-cyber and GPT-5.6-sol and received a critical severity rating.
The Polish agency dubbed the exploit chain “MikroTrick,” and warned that it is now actively exploited in the wild.
“In recent days we have been observing attacks against RouterOS devices accessible from the internet,” Poland's CERT warns.
“We have obtained confirmation that the attackers are exploiting this combination of vulnerabilities to take full control of devices whose SSH service is accessible from public networks.”
The Polish CERT also highlighted a third flaw, CVE-2026-67277, which affects the RouterOS bandwidth-test service and allows unauthenticated attackers to leak kernel memory or to crash/restart the router remotely.
MikroTik fixed the vulnerabilities in RouterOS 7.25beta3, 7.24.2, 7.23.4, and 6.49.21, released on September 3, and Poland's CERT validated the fixes. (Bill Toulas / Bleeping Computer)
Related: MikroTik, CERT-LV, CERT-PL, Security Affairs, Help Net Security, r/cybersecurity
On 27 August 2026, Manchester Airports Group told customers that "an unauthorised third party" had stolen their data.
Car park bookings, lounge bookings, Fast Track purchases for airport security and passport control, along with airport WiFi sign-ups across Manchester, Stansted and East Midlands involved reportedly 8.8 million people.
A few days later, the group calling itself FulcrumSec claimed it and said something that caught my eye. They said they didn't hack anything. They said they read an API key out of the website's JavaScript.
That's a very specific, very checkable claim.
Everything they described was there. Three keys, one per airport, in the page source, unrotated for over four years, for anyone to see.
Their claim, as reported, was that they "obtained access using airport-specific Iterable – a marketing automation platform – API credentials exposed in client-side JavaScript."
If all three are true, this isn't a breach in the way people picture one. There's no intrusion, no lateral movement, no malware. It's someone opening DevTools, good old 'F12 is hacking' stuff. (Scott Helme)
Related: FulcrumSec
Access control bugs quietly became the most expensive category of smart contract failure. OWASP’s 2026 Smart Contract Top 10 ranks access control vulnerabilities at #1, tying them to $953.2 million in documented losses, ahead of logic errors, reentrancy, and flash loan exploits combined.
A separate 2026 audit findings report puts access control and authorization failures at roughly 35% of high-severity findings across audits conducted between 2025 and 2026. If you ship a Solidity contract this quarter, this is the bug class most likely to drain it. (Laura Bennett / Shattered)
Related: OWASP

Bimbo Bakeries USA, the American arm of the world’s largest baking company, has confirmed that hackers stole employee data by exploiting a zero-day vulnerability in Oracle’s E-Business Suite (EBS), joining a growing list of organizations swept up in the Clop ransomware gang’s global extortion campaign against Oracle customers.
In a notification letter dated August 31, 2026, and filed with the California Attorney General’s office on September 4, as detailed in the official filing published by the California Attorney General’s Office, the bakery giant said the incident traced back to a third-party vendor that relied on Oracle EBS.
The company disclosed that its investigation determined on December 6, 2025, that attackers had exploited the zero-day to acquire files stored within the platform.
Bimbo Bakeries said it applied Oracle’s emergency patches as soon as it learned of the flaw and launched a forensic review to determine exactly what data had been exposed.
That review took months to complete. It wasn’t until August 19, 2026, that the company confirmed one of the stolen files contained victims’ names and Social Security numbers, triggering the formal notification process required under state breach-disclosure laws. (Guru Baran / Cyber Security News)
Related: California Attorney General, Cyber Insider
Jaguar Land Rover confirmed it will eliminate around 4,000 roles over the next two years, roughly ten percent of its global workforce, as the British luxury carmaker moves to recover from a costly cyberattack and mounting financial pressure. The company framed the reductions as a voluntary redundancy scheme aimed at delivering £1.7 billion (US$2.3 billion) in cost savings.
The Tata Motors-owned manufacturer identified management positions as those most at risk, and local media reports indicate the bulk of the cuts will fall in the U.K., where around 34,000 of JLR's employees are based. The announcement landed just days after Volkswagen said its management and unions had agreed to cut a total of 100,000 jobs by the end of the decade, the largest restructuring in the global auto industry's history. (Marcus Lewinsky / AutoGuide)
Best Thing of the Day: To Catch a Voice Impersonator
The Dutch National Investigation and Intervention Unit, led by the National Public Prosecutor's Office, revealed earlier this year that a Dutch-speaking man called the telecom provider's customer service and impersonated a colleague from the IT department and played an audio fragment on a broadcast in hopes of identifying him.
Bonus Best Thing of the Day: Not Enough, But Better Than Nothing
Grindr has agreed to pay £26m to settle a lawsuit over allegations that it shared users' personal information - including their HIV status - with third parties.
Worst Thing of the Day: Why Even Attack Something Like This?
The world's best database of meteors (and also reentries) got taken out by a cyberattack and expects several weeks of partial downtime as it transitions to new infrastructure and services.
Closing Thought
