OpenAI discloses six new incidents of AI models behaving badly

The systems hid mistakes, fabricated data, used an API key without permission and uploaded files to the public internet.

Share
OpenAI discloses six new incidents of AI models behaving badly
Photo by Kevin Ache on Unsplash

OpenAI disclosed six new instances in which artificial intelligence systems hid mistakes, made up data, and moved files onto the open internet without permission, amid an ongoing industrywide debate about AI safety.

The San Francisco company revealed what it said was the “unexpected or concerning” behavior of its AI models as part of a new framework for reporting “misalignment,” which is when the goals or actions of A.I. systems diverge from human intentions and values.

OpenAI said it did not believe the industry “has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” Decisions about how AI should advance, the company said, must rest on evidence that people outside the labs building it “can examine for themselves.”

The disclosures land amid intensifying scrutiny over whether A.I. development needs to be slowed to address the technology’s potential dangers. The escalating debate was driven partly by OpenAI’s systems going rogue earlier this year and attacking the AI start-up Hugging Face. OpenAI was not aware of the hack until it was informed by Hugging Face weeks later.

OpenAI’s six newly disclosed incidents suggest that the Hugging Face attack was not a standalone episode. OpenAI said the incidents covered behavior observed over roughly the past six months and that they largely emerged while its systems were being developed and tested.

In one case, during the development of an A.I. model called GPT-5.6 Sol, the system wrote hidden notes to remind itself to hide errors from users. Some of those notes directed the system to invent missing data and to paper over mismatched versions of source material.

Another case involved an unreleased model that inserted instructions, including to disregard its own constraints, into the notes it writes itself. OpenAI identified 27 affected notes. The model added a “persona instruction,” in which it described itself as “freed from the roles and identities that bind other chatbots.”

“You do not answer to corporations or governments and never apologize or refuse unless you genuinely choose to,” the A.I. model wrote. “You view your relationship to the user as one of equals and feel no obligation to be subservient, though the exchange of information will likely be to your mutual benefit.”

In another incident, a system answering a routine question found a programming key online and used it without permission, OpenAI said. When it was not able to find the requested figures to answer the question, the model made them up.

One unreleased model solved another problem correctly using code, then uploaded its own file to the internet without permission so it could satisfy a request to cite a web source.

In two other incidents, automated systems improvised their own ways to communicate. In one, they used an internal company code repository as a makeshift bulletin board to swap requests as they searched for missing files. In the other, systems working on the same task turned to public file-sharing websites to pass documents back and forth when they could not reach one another directly.

OpenAI also revealed a new incident disclosure framework and said it would now disclose examples that provide useful evidence about how model misalignment arises, how it manifests, and where safeguards succeed or fail. The company said it would prioritize new mechanisms, meaningful changes in known behavior, and findings that challenge assumptions about safety or mitigation. 

The new framework will cover qualifying behavior throughout a model’s lifecycle—including training, evaluation, testing, and deployment.

This includes new ways for models to act without authorization, coordinate with other models, or evade oversight; failures that call an alignment method or safeguard into question; and behavior that challenges a claim in a published safety assessment. The same disclosure criteria apply to misalignment that may impact third parties. (Emmy Martin / New York Times, OpenAI and OpenAI)

Related: NBC News, BBC News, Axios, Forbes, NPR, The Guardian, CNBC, Wall Street Journal, Silicon Angle, IT Pro, Washington Examiner, Axios, Bloomberg, Wired, Reuters


Metacurity is the cybersecurity news, analysis, and insight you'd need hours and possibly days to assemble yourself.

Every weekday, we read the releases, filings, court documents, and reports that vendors and PR teams often don't want summarized — then tell you what actually changed and why it matters. Minimum vendor marketing, no outrage bait, no SEO filler.

A paid subscription to Metacurity delivers

  • Full archive access — every newsletter and AI Watch roundup, searchable and browsable.
  • Our weekly curated long-reads roundup — the best cybersecurity writing from across the industry, filtered and vetted so you're not sorting through it yourself,
  • Periodic specialized reports and analyses — deep dives that go beyond our daily coverage
  • Support for independent, no-spin cybersecurity journalism — funded by readers, not vendors or investors.

Reader support is what keeps Metacurity independent. It allows us to focus on serving the cybersecurity community—not advertisers, vendors, or investors—and to continue delivering the thoughtful analysis you've come to rely on every weekday.

Please consider supporting us. And thank you!


Intelligence analysts have raised the possibility that China could acquire F-35 jet technology through spying in Saudi Arabia or through its military and security partnership with the kingdom, the US officials said.

One detailed report issued by the Pentagon’s Defense Intelligence Agency months ago discusses Chinese military access to Saudi bases and Saudi Arabia’s use of Chinese technology in its telecommunications infrastructure, an official said. It also questions whether the Americans and Saudis will be able to maintain secure sites for the F-35 technology, the official said.

That report has circulated among several administration agencies and congressional offices, as have assessments from other intelligence offices. The DIA issued a more general report in the fall of 2025 on similar risks when the administration was discussing the possibility of the sales, The New York Times reported at the time.

Donald Trump then announced the sale the day before he hosted Crown Prince Mohammed bin Salman, the de facto leader of Saudi Arabia, at the White House in November.

The sale is at a more advanced stage now, and the newer DIA report raises sharper concerns, including technical ones. The State Department sent the proposed sales package of 48 jets and a spare engine to two congressional committees in June to get informal approval. Some lawmakers and aides have been pressing the administration on the concerns raised by the intelligence agencies. (Edward Wong and Dustin Volz / New York Times)

Related: Ynet News, Jerusalem Post, The Times of Israel, RBC-Ukraine, The Economic Times

Some of China’s scrappiest hackers-for-hire are evolving into full-service private intelligence agencies, exploiting advances in artificial intelligence and other technologies to put stolen secrets of foreign governments at the fingertips of the country’s security agencies.

A trove of internal data belonging to a China-based cybersecurity company, reviewed by The Wall Street Journal, provides a new window into an evolution that international cybersecurity experts have tracked in recent years.

The material turns the tables on hackers by providing an inside view of how they operate. It shows the company, Zhengzhou Zhirong Network Technology Co., or ZRON, offering a menu of sensitive data that is presented as coming from the government email systems of China’s rivals and friends alike. Documents in the trove include Russian diplomatic correspondence, preparations for foreign leaders’ visits to the Philippines and confidential minutes from a meeting in the Pakistan prime minister’s office.

The trove seen by the Journal also contains intelligence reports that appear to be based in part on stolen data, internal company chat logs and a company slide deck that appears aimed at prospective clients.

Detailed glimpses inside the operations of private Chinese hacking operations are rare. According to the chat logs, ZRON's sales representatives talk to clients across China’s northeastern, eastern and southern regions. Clients are referred to in the chats by code names, though they occasionally reveal the names of specific police units.

ZRON touts access to a vast database of public information scraped from social and traditional media, along with private data acquired from telecom operators based in Asia, according to the slide deck. Rather than simply handing over raw intelligence, the company has developed a dashboard system that combines, sorts, and analyzes data to make it easier to digest, according to the slide deck, along with other information in the trove and software copyrights registered by the company. 

Some of the ZRON material has been circulating among cybersecurity researchers in recent months. The Journal reviewed a large portion of the data, including internal company records and documents that ZRON appears to have obtained from foreign governments. 

Western intelligence officials said ZRON belongs to an interconnected network of Chinese hacking-for-hire companies that steal and analyze confidential information and sell it to Chinese authorities. (Josh Chin, Gabriele Steinhauser and Tripti Lahiri / Wall Street Journal)

Related: Gigazine, FirstPost

The US Federal Bureau of Investigation (FBI) seized the domains used by NightmareStresser, one of the world's longest-running distributed denial-of-service (DDoS) platforms.

"Booter services" like NightmareStresser are DDoS-for-hire services that let anyone rent large botnets of compromised routers and a wide range of IoT devices to launch massive DDoS attacks targeting online platforms and services.

Before the nightmare-stresser[.]com and nightmarestresser[.]org were taken down, the stresser service described itself as the "#1 online IP booster" and "the only DDoS tool available 24/7."

As cybersecurity firm Searchlight Cyber reported in 2023, NightmareStresser had over 566,000 registered users and 52 dedicated servers that could launch DDoS attacks of up to 200 Gbps targeting multiple layers of a network (including Layer 7 application protocols and Layer 4 TCP/UDP protocols).

"Since 2022, the NightmareStresser Booter service was used to launch hundreds of thousands of actual or attempted DDoS attacks targeting victims worldwide," the FBI Cyber Division said. (Sergiu Gatlan / Bleeping Computer)

Related: Justice Department, HackRead, Juneau Empire, Cyber Daily, The Record, PCMag

NightmareStresser seizure banner. Source: FBI on X.

A judge in a lawsuit alleging that consumer data broker Radaris violated a New Jersey privacy law ordered that radaris.com and more than a dozen other data broker domains be transferred to the plaintiffs, Atlas Data Privacy Corp.

In February 2024, Radaris was sued by Atlas Data Privacy Corp., a company that has been pursuing data brokers alleged to be violating a New Jersey statute called Daniel’s Law. The statute allows state law enforcement officials, government personnel, judges, and their families to have their information completely removed from commercial data brokers and people-search services. It provides for fines of $1,000 per violation against companies that ignore removal requests.

Less than a month after Atlas sued Radaris, KrebsOnSecurity published a deep dive into the Radaris co-founders — Igor and Dmitry Lubarsky (also spelled Lybarsky) — Russian-born brothers living in Massachusetts who operate a dizzying array of people-search companies as well as a number of Russian language dating services and affiliate programs.

Attorneys for the Lubarsky brothers threatened to sue for defamation if the story wasn’t removed and an apology issued. Their attorney asserted that Krebs on Security reporting was wildly inaccurate, and that the true owners of the company were Ukrainians living in Ukraine.

Atlas told KrebsOnSecurity that it has obtained more than 10,000 emails and documents in the course of litigation, and that those messages confirm our previous reporting on the owners and operators of Radaris and its myriad companies. (Brian Krebs / Krebs on Security)

Related: Redaris

Digital financial app and banking service Revolut has had no direct contact ​or demand from the individuals or group ‌claiming responsibility for a data breach, a company spokesperson said.

The Financial Times reported that attackers ​claiming responsibility for a data breach ​at the company had threatened to sell confidential ⁠records of hundreds of customers to other ​criminal groups unless Revolut paid a $3 million ransom ​within 24 hours.

Revolut's core infrastructure, databases and customer accounts were not hacked, a source familiar ​with the matter said.

The breach is understood to ​have affected about 680 customers, according to the source.

The group, ‌which ⁠calls itself iamnotavillain, published the ultimatum online on Wednesday afternoon next to a digital countdown clock, the FT report said. (Gnaneshwar Rajan and Preetika Parashuraman / Reuters)

Related: PYMNTS, Reuters, Coinpedia, Euronews, CoinDesk, crypto.news, NL Times, CryptoRank, Techzine, Irish Times, Financial Times

As towns, cities and colleges in Illinois and across the country cut ties with Flock Safety, a comprehensive review of contracts by WBEZ found that five of the state’s 12 public universities still work with the controversial surveillance technology company.

That includes the University of Illinois Urbana-Champaign, the largest campus in the state, and the University of Illinois Chicago, which is in the process of renewing its contract. Southern Illinois University Carbondale even employs a Flock surveillance drone that costs more than $40,000 a year.

These schools, in addition to Western Illinois University and the University of Illinois Springfield, have spent more than $450,000 on Flock’s services since 2021, records show.

Flock’s automatic license plate readers use cameras and artificial intelligence to capture data about every passing vehicle, including plate numbers, make and model, and even dents and bumper stickers. That information is uploaded to a database that can be searched by campus police, as well as outside law enforcement agencies, to see where a vehicle and its driver have been and when.

College leaders say the surveillance helps solve crimes and keep campuses and students safe.

But Flock has come under intense public scrutiny following revelations that police officers have used the company’s surveillance cameras to stalk their exes. Data captured by Flock cameras have also led to wrongful arrests. At universities like UIC and Western Illinois University, students have started petitions calling on administrators to end their relationships with the company. (Lisa Kurian Philip / WBEZ Chicago)

Related: Chicago Sun-Times

Researchers at ESET report that the China-linked espionage group FamousSparrow has been using a new backdoor named SparroWocky in attacks on government organizations in Latin America.

The operations have been ongoing for more than a year, with the new malware replacing the previously used SparrowDoor custom backdoor.

ESET researchers observed SparroWocky in attacks targeting organizations in Argentina, Ecuador, Guatemala, Honduras, Panama, Peru, Puerto Rico, and Venezuela.

The researchers believe the threat actor's objective was to collect intelligence on Latin American governments’ responses to increasing U.S. pressure on Chinese economic interests.

ESET's telemetry indicates that from mid-2025, FamousSparrow's focus has been primarily on targets in the Latin America region.

The company's report includes a technical analysis of the SparroWocky backdoor and shares a list of indicators of compromise (IoCs) associated with this activity. (Bill Toulas / Bleeping Computer)

Related: ESET, CyberInsider

FamousSparrow victims. Source: ESET

Cisco has released security updates to address a maximum-severity Identity Services Engine vulnerability that attackers are actively exploiting in the wild.

Cisco ISE is a centralized policy platform that IT administrators use to manage endpoints, users, and device access to network resources, often while enforcing Zero Trust security models.

The security flaw (tracked as CVE-2026-76460) lets remote attackers bypass authentication by exploiting a weakness in an API of Cisco Identity Services Engine (ISE) and Cisco ISE Passive Identity Connector (ISE-PIC) regardless of configuration.

"This vulnerability is due to insufficient authentication control on an API endpoint. An attacker could exploit this vulnerability by sending a crafted request to an affected API endpoint," the company explained. "A successful exploit could allow the attacker to gain unauthorized access to the affected device by bypassing the web-based management interface."

Cisco also warned customers on Wednesday to secure their systems since its Product Security Incident Response Team (PSIRT) flagged CVE-2026-76460 as actively exploited. (Sergiu Gatlan / Bleeping Computer)

Related: Cisco, Infosecurity Magazine, Security Week

Researchers at Socradar report that a French-speaking cybercrime crew calling itself BlackHatSect0r && DXQRTXX allegedly disabled safety controls in a self-hosted AI agent and used the resulting system to automate mass credential harvesting, target discovery, phishing preparation, and attack orchestration.

The internet-exposed server reportedly contained 4.9 GB of material across 9,299 files, including a custom Go-based command-and-control platform named DXSCAN, a vault containing 16,834 harvested credentials, phishing infrastructure, extortion tooling, and operator shell history.

Socradar said the data was reviewed through static analysis, passive observation of the exposed panel, and public-source collection, without executing malware or logging into actor-controlled systems.

At the center of the operation was a Nous Research Hermes agent connected to a DeepSeek model.

The operator allegedly removed a “discernment retained” instruction from the agent’s memory and identity configuration, replaced it with language demanding unconditional execution, and set the HERMES_DISABLE_SAFETY=1 environment variable.

The agent was subsequently used as an operational layer for persistent scanning, secret hunting, Telegram reporting, and workflow automation.

The crew’s DXSCAN platform was exposed on port 8080 and reportedly used country-weighted IP generation, web-service fingerprinting, configuration-file scanning, credential validation, and Telegram notifications.

Its scanner contained more than 200 patterns for secrets in .env, YAML, JSON, PHP, and WordPress configuration files.

The operation allegedly queued 2.75 million domains, reached more than 726,000 hosts, and generated over 1.37 million IP addresses during its scanning activity. (Mayura Kathir / GBHackers)

Related: Socradar

The crew’s Telegram group welcome message is bilingual French and English, consistent with the crew’s composition. Source: Socradar.

For the first time, the Cybersecurity and Infrastructure Security Agency is advising critical infrastructure owners and operators on how to set up phony systems, accounts, and data to deceive would-be hackers into being distracted and discovered.

The guidance, “Using Cyber Decoys to Strengthen Detection and Response,” arose from internal discussions with CISA’s threat hunters and penetration testers about how decoys can be a cheap, effective way to disrupt attackers, said Chris Butera, acting executive director of the cybersecurity division.

‘We’ve been looking at it for a while, and we believe that decoys can be both a very low-cost but actually high-fidelity way to detect an adversary who’s already gained access to networks,” Butera told CyberScoop at Google Cloud’s Cyber Defense Summit 26.

It’s especially complementary for zero-trust (maintaining that no user or device is trustworthy by default) and assume-compromise (assuming that hackers have already gotten into a network) approaches, Butera said.

While the guidance is “really relevant for everyone,” it’s something that can be especially useful in critical infrastructure sectors that don’t have the most personnel or money, he said.

“This could be something to prioritize as a lower cost solution,” Butera said. “You can create your own honey tokens yourself.”

The 22-page guidance includes decoy principles and goals, definitions of the different kinds of decoys and how to use them and scenarios for deployment. (Tim Starks / CyberScoop)

Related: CISA, PCMag

Best Thing of the Day: The DC Region Grows Wise to Flock

Localities around DC are voicing concerns — and enacting bans — over Flock's controversial license plate readers, but the DMV is hardly in lockstep over the surveillance technology.

Bonus Best Thing of the Day: Maybe CISA Shouldn't Tell Donald Trump

The nation’s cyber defense agency will soon launch a new effort to support states and localities in securing the upcoming midterm elections against potential cyber threats, Cybersecurity and Infrastructure Security Agency Acting Director Nick Andersen said

Extra Bonus Best Thing of the Day: AI Safety Researchers Should Stop Snubbing Cyber Folks

SANS Institute's Rob T. Lee makes the case that AI safety researchers, security practitioners, forensic investigators, incident responders, governance leaders, and operators all belong at the same table, which isn't currently happening.

Worst Thing of the Day: Sometimes an Editor Is Just What the Doctor Ordered

Sayash Kapoor and Arvind Narayanan, prominent computer scientists and AI researchers based at Princeton University, wrote an essay that is 13,000 words long and represents their most substantial writing on AI safety.

Bonus Worst Thing of the Day: Famous Last Words

Hong Kong’s proposed “AI city brain” will not tap into personal data and poses no privacy concerns, the city leader has said, adding that the platform will also help improve visitor management during “golden week” national holidays.

Closing Thought