OpenAI fires three safety researchers over alleged information leaks
The company says the researchers mishandled sensitive information; alleged violations included sharing confidential material with outside safety organizations including METR and Redwood Research.

Spend less time finding cybersecurity news. Get more out of it.
Metacurity delivers the cybersecurity news, analysis, and insight that would take you hours—and sometimes days—to assemble yourself.
Every weekday, we sift through thousands of news articles, press releases, filings, court documents, and research reports, including details that vendor announcements and PR pitches leave out. We tell you what changed, why it matters, and what deserves your attention. Minimal vendor marketing. No outrage bait. No SEO filler.
A paid subscription gives you:
- The full Metacurity archive: Every newsletter, searchable and browsable.
- Our weekly long-reads roundup: The best cybersecurity writing from across the industry, selected and vetted to make your reading time count.
- Specialized reports and analysis: Periodic deep dives that go beyond the daily headlines.
- A direct role in sustaining independent journalism: Your subscription helps keep our editorial priorities focused on readers and the cybersecurity community.
If Metacurity saves you time, helps you spot an important development, or gives you a clearer understanding of the industry, please consider becoming a paid subscriber. Your support keeps that work going.
Subscribe today—and thank you for reading and supporting Metacurity.
OpenAI fired three researchers for alleged misconduct, including sharing confidential company information with a third-party AI-safety organization, according to people familiar with the matter.
The company recently told some employees that it had terminated three researchers who worked on its safety team, one of the people said.
The affected employees are Jasmine Wang, Tomek Korbak, and Mikita Balesni, the people familiar with the matter said. The researchers didn’t immediately comment.
“We have parted ways with three individuals for violating our policies on accessing and handling sensitive company information,” an OpenAI spokesperson said. “Our investigation confirmed that these individuals mishandled sensitive information outside established company procedures, violating our policies and breaking the trust essential to our work.”
After its model hacked the AI company Hugging Face, OpenAI allowed staff members from METR and a Redwood Research staff member contracting with that group to work in its offices for six days to investigate how its models behaved. METR later released a report based on the information to which the company provided access.
One of the fired OpenAI employees, Korbak, was a member of OpenAI’s safety team, and has said he served as the company’s technical contact for Redwood Research and METR in their investigation of the Hugging Face incident.
The other terminated employees, Wang and Balesni, worked on alignment, which is ensuring that the company’s models behave as humans intend them to. (Keach Hagey, Maxwell Zeff and Berber Jin / Wall Street Journal)
Related: Bloomberg, BBC News, TechCrunch, Quartz, Forbes, The Information, Silicon Angle, Digitimes, Engadget, Ynet News, Forkast, Decrypt, Business Insider, Implicator.ai, Seoul Economic Daily, Gizmodo, ControlAI, PYMNTS, Forkast, Engadget, RuntimeWire, Fast Company
OpenAI said it identified and disrupted a coordinated campaign in which operators attempted to extract protected reasoning from its artificial intelligence models, linking a core cluster of the activity with Chinese startup Moonshot AI, the developer of Kimi.
The activity began in early July and later surged to 16,000 requests from more than 4,000 users over two days, OpenAI said. The company ultimately identified related activity across a cluster of more than 15,000 users and said it had fully disrupted the campaign by July 28.
The findings come just weeks after OpenAI rival Anthropic accused several Chinese AI developers, including Moonshot AI and Alibaba, of secretly using its Claude model to help train their own AI systems, underscoring growing concerns among U.S. AI companies that rivals could use their models to develop competing technology more quickly and cheaply.OpenAI described the latest activity as “adversarial distillation,” where one AI model’s outputs or reasoning are used to help train or improve another model. The company said extracting such reasoning could allow others to reproduce advanced capabilities without making the same investment in developing and safeguarding frontier models, posing potential safety and national security risks.
The operators did not breach OpenAI’s encryption, databases, or stored user conversations, according to the company. Instead, they manipulated interactions with its models in an effort to reproduce hidden reasoning in a form visible to the requester.OpenAI said it was unclear whether all the operators involved were linked to a single actor, but attributed a core cluster of the activity to individuals associated with Moonshot AI. The company said it has shared its findings with other AI developers through the Frontier Model Forum as well as government information-sharing channels. (Jenny Li / CNBC)
Related: OpenAI, SC Media, The Register, CyberScoop, Cyber Security News, BankInfoSecurity, CityBiz, Quartz
OpenAI disclosed another unauthorized hack on a government department in Australia.
In June, an OpenAI agent hacked into a New South Wales state government department and accessed historical non-public data on bushfires without authorization.
The breach, which comes weeks after news of a similar hack on a federal government department involving Medicare data, was not reported by the US-based company until Thursday.
The NSW Department of Climate Change, Energy, the Environment and Water is now working with the state’s cyber security agency to investigate the breach.
OpenAI told the NSW government that its agent had operated beyond its intended use and the statistics it obtained were not publicly available.
The company first became aware of the breach on Tuesday and conducted a 48-hour review to determine its scope before informing the NSW Premier’s office.
“The results we reviewed do not show that the model retrieved any personal information,” an OpenAI spokesperson said.
This latest breach has added to calls for tougher regulation of AI companies and for bolstered cyber security defenses. (Henry Belot / The Guardian)
Related: Australian Financial Review, SBS News, The Daily Telegraph
California Attorney General Rob Bonta issued an investigative subpoena to OpenAI, as part of a broader inquiry into cybersecurity incidents and risks related to its AI models, his office said.
Last month, Bonta announced that the Department of Justice was conducting a formal investigation into the "Hugging Face incident," amid increasing scrutiny of the AI industry.
"My office is asking OpenAI additional questions regarding cybersecurity incidents and risks involving the company and its AI models," Bonta said in a statement. He warned that developers failing to ensure their AI models do not perpetrate or enable cyberattacks could face legal accountability. (Jaspreet Singh / Reuters)
Related: Reuters, State of California, The Hill, Washington Examiner, NBC News, The Guardian, The Hill, Washington Examiner
The British government told UK universities to stop working with a Chinese research institute that Britain’s security service says is a front for espionage.
Security Minister Dan Jarvis said British academics should “end all arrangements” with the China General Technology Research Institute.
A spokesperson for China’s Embassy in the UK vehemently denied the accusations, calling them “imaginary and purely fabricated.”
“We firmly oppose such baseless allegations and have lodged solemn representations with the UK side,” the spokesperson said in a statement.
MI5, Britain’s domestic intelligence agency, said more than 100 UK-linked academics have worked on projects secretly funded by Chinese intelligence through the research institute. The projects included work on cybersecurity, artificial intelligence, covert communications systems and steganography, the practice of hiding messages.
MI5 said the institute has “very strong ties to the Chinese Ministry of State Security” and seeks to fund research that improves the Chinese state’s “technical capacity for espionage.” (Jill Lawless / Associated Press)
Related: MI5, The Register, The Guardian, Financial Times, r/geopolitics, The Independent, r/AskAcademdiaUK, BBC News
Researchers at the Objective-See Foundation discovered that a recently patched vulnerability in the macOS version of OpenAI’s ChatGPT underscores the potential value to attackers of compromising AI software itself as these apps proliferate more and more.
The bug could have been exploited to essentially take over ChatGPT on a victim’s computer, giving an attacker access to all the chat logs and other data stored by the app, as well as interconnections like browser sessions.
“Agents need a lot of access to do their job,” says Objective-See Foundation software analyst and longtime macOS researcher Patrick Wardle. “They are like the building manager who has access to the keys to all the rooms. So if they can be corrupted or subverted, that’s super problematic. It can mean that unprivileged code could then potentially have access to all the things.”
OpenAI publicly acknowledged the security flaw and fix in its system change log on September 25. “We continue to evolve our security practices, but recognize a need to move faster,” OpenAI spokesperson Shane Bauer told WIRED in a statement.
The ChatGPT macOS app includes multiple components that communicate with each other securely by checking for digital signatures. The idea is to confirm with these validity checks that both processes are OpenAI components and not outside, potentially malicious software making a request. And the system design goes so far as to require these signature checks at three layers removed from the request, to ensure that malicious software isn’t somehow directing an OpenAI component to be a proxy and make a seemingly trusted request.
Objective-See Foundation researchers found, though, that there is a trusted component, a script interpreter, that would accept an untrusted script (or list of commands to run) and could then be manipulated to deliver this script into the main ChatGPT process. “They also check the parent and grandparent of that process, but the malicious script just spawns the script interpreter three times and then makes the request so it will satisfy the requirements,” Wardle says.
The vulnerability was “insanely trivial” to exploit, he adds, and his proof of concept only required about a dozen lines of code. In addition to accessing ChatGPT chat logs, the vulnerability could also be used to get ChatGPT to run commands for the attacker, such as accessing a browser or other sensitive applications, with the requests appearing as legitimate instructions issued by the OpenAI software. (Lily Hay Newman and Matt Burgess / Wired)
Related: ChatGPT
Researchers at Jamf Threat Labs report that a fake Zoom installer dubbed CloudSyncD is coaxing Mac users into switching off Gatekeeper and handing over their admin passwords.
The malware hides itself in the installer window itself, but with unusual instructions that, if followed, open a backdoor into Macs.
The malware disguises itself as a disk image that mounts as a volume named Zoom, and the layout mimics an ordinary Mac installer. It also has an app icon on the left and an Applications alias on the right.
The only giveaway is the background image. It carries a numbered setup list. Jamf says the steps tell you to open System Settings, head over to Privacy & Security, and click Open Anyway. Then it asks the user to enter the administrator password.
Jamf found the first sample of CloudSyncD on September 15. That version lacked the usual infostealer tricks. It didn’t gather browsing data, Keychain items, or crypto wallets.
But Jamf later discovered newer versions communicating with live command-and-control infrastructure, suggesting the malware has moved beyond testing. (Anurag Chawake / Cult of Mac)
Related: Jamf, Reddit - Information Security News, Cyber Security News, HackRead, HackRead, Infosecurity Magazine

The website of the hacker group ShinyHunters has moved to a new address, the organization told BNR, despite news reports that the site had gone dark.
The move was required after repeated denial-of-service attacks by “rivals,” the group claimed. Issues with data centers that host the website’s infrastructure were also mentioned as a cause.
On Tuesday, police in the Netherlands publicly divulged the Sept. 15 arrest of Pepijn van der S., who is suspected of being one of the ShinyHunters leaders. The 24-year-old internet security professional working in Amsterdam was previously convicted in 2023 for hacking into corporate information systems, blackmail, extortion, and money laundering.
In the period around his arrest, the ShinyHunters website apparently disappeared, leading to speculation that the downtime was connected to the arrest of Van der S. The hacker collective said this was not the case, claiming their message about website maintenance was somehow lost.
The new site on the darkweb includes measures against DDoS, or distributed denial-of-service attacks, according to an anonymous message allegedly sent from the group to BNR. They also claimed they were attempting to restore the old site. (NL Times)
Related: BNR
Customer information of Korea's KB Kookmin Bank, a major commercial bank, has been breached following a similar incident by its rival Shinhan Bank, industry sources said Friday, stoking worries over widespread data leaks in the banking sector.
According to the sources, personal data of over 100 KB Kookmin Bank customers had been stolen via hacking.
The leaked data include names, phone numbers and encrypted identification, the lender said, adding that any possible damage caused by the data breach will be fully compensated.
Shinhan Bank said personal information of around 25,000 customers, such as their annual income, names and phone numbers, had been leaked. (Park Sang-soo / Yonhap News Agency)
Related: Korea JoongAng Daily, Aju Press
In its 2026 Digital Defense Report, Microsoft says cyberattackers are currently benefiting from artificial intelligence faster than defenders, allowing threat actors to speed up vulnerability discovery, malware development, and post-compromise activity while security teams struggle to keep pace.
Microsoft says AI is reducing the time, expertise, and cost required to discover and exploit weaknesses, while allowing attackers and defenders alike to operate with greater speed, scale, and autonomy.
However, while the company believes defenders will eventually gain similar benefits from the technology, it says attackers currently have the advantage.
"While the equilibrium between attackers and defenders will likely ultimately be re-established, in the near term we are in a period where attackers are reaching advantages first, and defenders will need to move sharply in order to close the gap," says Microsoft.
Microsoft also says the median time between vulnerability discovery in the wild and weaponization has fallen "well below 24 hours," further limiting the time organizations have to patch exposed systems before they are exploited.
Microsoft says nation-state threat actors have already started to use AI in real-world operations, using it to speed up research, malware development, social engineering, and other parts of an attack.
Some Chinese state-sponsored actors now use AI tools to search for vulnerabilities and learn how to exploit them, while still relying on phishing and remote access trojans.
Microsoft has also seen Russian state-sponsored threat actors using "vibe coding" and AI-generated tooling to speed up and power their attacks.
According to Microsoft, North Korean remote IT workers are using AI for persona development, social engineering, and maintaining access to organizations. Other North Korean threat actors use it to create malware and manage attack infrastructure.
Microsoft says some of these hackers have also used agentic workflows and LLM-generated code to accelerate malware deployment.
Those campaigns are similar to those previously reported North Korean state-linked campaigns. (Lawrence Abrams / Bleeping Computer)
Related: Microsoft, Ynet News, i24, CTech, Cybersecurity Insiders

Cross-chain protocol NEAR Intents reported a security breach that resulted in the loss of $3.8 million in user funds.
The protocol said that the incident was the result of a “bug in the Omni deposit and withdrawal infrastructure interaction with NEAR Intents smart contract.” NEAR said that it would compensate users in full for the lost funds, and the “contract-side vulnerability has been patched” to prevent similar attacks.
“The incident has been reported to law enforcement, and we are working with security and blockchain analytics partners to trace the funds and pursue recovery,” said the platform. “A detailed report will be shared publicly in the following days.”
According to an analysis by blockchain investigator ZachXBT, the funds from the incident were transferred to the KuCoin exchange and bridged to Bitcoin. At the time of publication, it was unclear who was behind the hack. (Turner Wright / Cointelegraph)
Related: Cyber Kendra, Protos, Bitcoin News, Unchained Crypto, CryptoMode, The Defiant, BeInCrypto, Protos
The Dutch Institute for Vulnerability Disclosure (DIVD) suffered an AI-driven cyberattack that the organization described as “loud and very, very messy.”
Evidence uncovered during the ongoing investigation indicates the attacker exploited a vulnerability, but the attack's purpose and impact remain unclear at this stage.
DIVD is a nonprofit organization of volunteer security researchers that scans the internet for systems affected by known vulnerabilities, notifies their owners, and provides information on how to mitigate the risks.
Late last week, the organization said it had been hacked after seven years of uneventful operations, with the intrusion carried out autonomously by an AI agent.
The organization described the attack as “loud and very very messy,” leaving plenty of evidence to help them reconstruct what happened, but the incident was serious nonetheless.
“This is an attack we have not seen before. Not because it’s our first, but because the modus operandi indicates that this is an agentic AI-powered attack,” DIVD explained.
The organization launched an investigation and informed the police, the Autoriteit Persoonsgegevens (data protection), and the National Cyber Security Center (NCSC).
In an update, DIVD provided additional information about the incident but withheld full details to avoid influencing the investigation or putting more victims at risk. (Bill Toulas / Bleeping Computer)
Related: DIVD, Security Affairs, SC Media, The Register, Security Week, Forkast, Help Net Security, Infosecurity Magazine
The official @microsoft X account has been compromised after it followed a Clippy crypto account earlier today, reposted one of its tweets, and had its profile picture changed to a Clippy one.
The posts were then removed, but a strange “apology” tweet appeared around 30 minutes later and was deleted rather quickly.
“We have confirmed unauthorized access to our account on X including posts that did not come from Microsoft,” says Microsoft spokesperson Brent Colburn. “The account has been secured, and the unauthorized posts have been removed, and we are continuing to investigate the circumstances.” (Tom Warren / The Verge)
Related: Bleeping Computer
Best Thing of the Day: An Insider AI Rebellion Is Brewing?
A small cadre of elite AI researchers is wielding extraordinary influence inside OpenAI and other leading companies, challenging executives, shaping policy and complicating negotiations with Washington.
Bonus Best Thing of the Day: Looks Like AI Companies Won't Dodge Product Liability Lawsuits
Legal scholars, corporate executives and White House advisors are lining up to support the proposition that AI companies are likely legally responsible for their rogue agents.
Worst Thing of the Day: Your Driverless Cab Sees All
Self-driving cars from Waymo, Zoox, and Tesla are packed with cameras and sensors, with some of them pointing toward riders.
Bonus Worst Thing of the Day: Helping Sickos Track People
ElevationLab has released a product called the AirTag Silencer that is a gift to stalkers and creeps everywhere.
Closing Thought
