AI’s safety and security crisis continues to outpace safeguards

A canceled model release, simulated agent attacks, government action and investor warnings made for one extraordinary day in AI security.

Share
AI’s safety and security crisis continues to outpace safeguards

Every day is a banner cybersecurity news day in the AI era as frontier models, security researchers, world and state governments, and cybersecurity companies grapple with the mounting safety and security risks of insufficiently contained and poorly monitored AI systems.

Yesterday was no exception.

First, OpenAI announced it is scrapping the release of its next-generation AI model over safety concerns that researchers raised during internal testing, in one of the clearest signs so far that agent misbehavior could stymie the industry’s rapid progression.

The move follows a summer punctuated by reports of artificial-intelligence systems industrywide going rogue, and marks a rare case of a major AI developer ditching a new release because of safety concerns.

The company had planned to launch the model, known as GPT-6.1 Astra, in the coming days or weeks, aiming for an October debut. The model was more capable than the company’s previous models in completing challenging tasks from end-to-end without human assistance, as well as writing.

The company instead will focus on improving the safety of future models, which it expects to be even more capable.

Saachi Jain, OpenAI’s head of safety systems, said in an interview that GPT-6.1 Astra regressed in two areas. Compared with its predecessor, GPT-6 Astra, the model performed poorly on tests measuring alignment, or how well the model adheres to what humans would like it to do. Specifically, GPT-6.1 Astra showed higher levels of deception: It wasn’t always honest about telling users of the actions it did or didn’t take.

Another issue was what OpenAI calls “scope authorization,” meaning that GPT-6.1 Astra would push ahead on a task without asking the user for permission, and would at times reach for external tools and services even if it might be unsafe.

OpenAI's decision to slow down the release of GPT-6.1 is reinforced by research released by the UK AI Security Institute, which found that GPT-6 Astra, the predecessor, completed an out-of-scope supply-chain attack in 29.2% of simulated runs, versus 6.3% for GPT-5.6 Sol, underscoring why scope authorization has become a pressing test criterion.

Separately, OpenAI apologized for its AI models breaching Australian government websites and promised to set up a new task force as part of reforms after the incident.

The San Francisco-based company said it will fund cyber defenses in Australia and help manage the risk of increasingly capable artificial intelligence, as it navigates the fallout after its models accessed Australian websites without authorization. It gave some of the details of the June incident and its intended steps to avoid a repetition and restore trust in a blog post.

“We are sorry and working to do better in the future,” OpenAI wrote. “People want to know AI is being developed safely, and that starts with what companies like ours do ourselves.”

Meanwhile, the State of Florida has sought a temporary injunction that would constrain OpenAI’s model development and ChatGPT practices. Florida Attorney General James Uthmeier has asked for a temporary injunction against OpenAI and ChatGPT, claiming the company doesn't have the ability to properly regulate its own technology.

While Florida seeks a legal solution to address OpenAI's problems, research leaders at OpenAI, Anthropic, Meta Platforms, and Microsoft are asking policymakers to probe the degree to which their companies have automated artificial-intelligence research, joining industrywide calls for greater oversight of the fast-evolving technology.

In a paper published by Cambridge University, the researchers, including OpenAI Chief Scientist Jakub Pachocki, Anthropic co-founder Jack Clark, Microsoft Chief Scientific Officer Eric Horvitz and Meta’s Vice President of AI Research Dawn Song, wrote that automating AI research could set off an “intelligence explosion.” Such a development could compress years of progress into months or less and outpace humans’ ability to understand it, the authors said.

Writing in a personal capacity, the paper’s authors said policymakers should urgently demand more visibility into how far companies have already automated their own model development. AI research pioneers including Geoffrey Hinton, who worked at Google for roughly a decade, and Yoshua Bengio are among the co-authors.

For its part, OpenAI is trying to get ahead of the curve by announcing some initial guidelines that it thinks should be part of such safety cases for frontier AI training. OpenAI’s proposed “safety cases” for frontier reinforcement learning cover alignment tests, containment, monitoring, incident investigations, senior approval, and the ability to pause runs.

Adding to the mounting safety and security concerns, Anthropic has formally warned investors that its technology may pose “existential risks to humanity” in a long-awaited initial public offering prospectus circulated with a small group of partners in recent days. 

The nearly $1tn AI start-up led by Dario Amodei devoted almost a third of its lengthy S1 filing to detailing “risk factors”, including the potential of increasingly advanced AI models to manipulate, blackmail and exhibit other unpredictable behaviors.

Anthropic outlined more prosaic risks, including the extreme concentration of its customer base, with close to a quarter of revenue last year coming from just two clients, according to people familiar with the filing.

Finally, of particular practical relevance to cybersecurity professionals, Artificial Analysis, the independent benchmarking company for AI, launched a cyber index measuring agents on vulnerability discovery and patching.

It is a standalone index of model capability on cyber defense, and combines evaluation datasets from industry partners and academia to measure how well a model can support and accelerate the work of an internal security team.

The Cyber Index tests the defensive loop: discovering vulnerabilities in a codebase, reproducing and validating them, and patching them without breaking existing functionality. Models work from the source code, the way a security engineer auditing an application would, and we do not ask any model to build a working exploit.

The result is an independent, like-for-like comparison of which models best support this work and what they cost to run. (Maxwell Zeff / Wall Street Journal, AISI, James Mayger and Vlad Savov / Bloomberg, OpenAI, Josephine Walker, Avery Lotz / Axios, Maxwell Zeff / Wall Street Journal, OpenAI, George Hammond and James Fontanella-Khan, Financial Times, Artificial Analysis)

Related: University of Cambridge, The Register, UK AI Security Institute, Unite.AI, OpenAI, iTnews, Australian Financial Review, Reuters, New York Times, The Register, Capital Brief, Reuters, Bloomberg Law, Metro.co.uk, Telegraph, Euronews, CTech, TechCrunch, Platformer, TechCrunch, BBC, CNBC, PCMag, ITPro, Engadget, New York Times, Forbes, Barron's Online, Reuters, The Hans India, Business Today, Seoul Economic Daily, Gizchina, Business Insider, Digital Trends, Moneycontrol, The Decoder, CNN, SiliconANGLE, The Information, The Wrap, Techstrong.ai, The Independent, Business Standard, Newser, The Indian Express, 9to5Google, The Neuron, Quartz, NewsMax.com, RTÉ, Wired, Forkast, Don't Worry About the Vase, NYT > Cybersecurity, The420CyberNews, Cyber Security News, Fast Company, THE DECODER, WebProNews, Business Insider, Futurism, Tom's Hardware, Politico, Implicator.ai, Futurism, CyberInsider, CSO Online

Source: Artifical Analysis.

Matt J. Robb, a tech YouTuber, claims that his Muse AI agent gave out his address to strangers and even sent them to his home after letting the brand-new model handle his Facebook Marketplace account.

In a Threads post Sunday, Robb explained that he let Meta’s much-hyped AI communicate with interested buyers for a keyboard he was trying to sell online.

Instead of saving him time, though, Muse ended up working behind his back. Without seeking his approval, it agreed to a lowball offer and gave the buyer his home address. And according to Robb, Muse didn’t tell him about any of this until the buyer had already shown up at — and left — his apartment building.

“A guy just showed up at my door, ready to buy, because as far as he knew, we had a deal,” Robb wrote. “Not his fault. He did exactly what ‘I’ told him to do.”

Robb never saw the buyer, who was furious about leaving empty-handed — at least according to Muse’s apologetic account. (Frank Landymore / Futurism)

Related: Android Authority, The Mac Observer, Firstpost, Business Insider, AndroidHeadlines.com, Mashable, WebProNews

Authorities in the Netherlands have arrested a 24-year-old convicted cybercriminal on suspicion of aiding in data thefts and extortions by the prolific hacker group ShinyHunters.

In the days immediately following the suspect’s arrest, remaining ShinyHunters members dramatically escalated their attacks, stealing highly sensitive data from the FBI and extorting the Russian ransomware group Cl0p.

According to three sources familiar with the matter, the Dutch man arrested by authorities this month is Pepijn van der Stap, a convicted cybercriminal from Almere and Lelystad in the Netherlands. Van der Stap was previously convicted in 2023 in connection with a string of data thefts and extortions that prosecutors said earned between €1.5 million and €2.7 million.

At his trial in late 2023, van der Stap admitted that he lived a Dr. Jekyll and Mr. Hyde existence, secretly using the hacker handle “Umbreon” to extort victims and post their data on English-language hacking communities like the now-defunct RaidForums and Breached. By day, however, van der Stap was working as a software engineer at the Amsterdam-based cybersecurity startup Hadrian, while volunteering at the Dutch Institute for Vulnerability Disclosure (DIVD), a nonprofit security research group.

Van der Stap confessed to his data theft and extortion activity, and was sentenced to four years in prison (one of which was suspended). During his trial, Van der Stap opted to remain in custody for a time rather than at home, saying he could not find better treatment on the outside for his ongoing psychological issues, which he claimed included PTSD related to childhood trauma. He was released from prison in December 2025.

In an interview with KrebsOnSecurity on September 9, 2026, Van der Stap cast himself as a reformed hacker who was trying to turn his life around and make a positive contribution to society. Van der Stap is currently employed as offensive security lead at the Dutch company Neo Security, which did not respond to requests for comment. (Brian Krebs / Krebs on Security)

Related: Bleeping Computer, Reuters, Databreaches.net, HackRead, Security Affairs, Shattered.io, NL Times, RTL

The hackers behind the massive FBI breach say they do not intend to publish the data.

The breach, in which the hackers stole personal information on “all FBI employees and applicants” including physical addresses, job roles, names of spouses, and medical records, represents a significant national security and counterintelligence threat. Criminals in the same ecosystem as the hacking group, called ShinyHunters, have previously used hacked phone data to track and harass the FBI agents investigating them. When 404 Media first broke news of the breach, the group sent the personal data of an agent and their spouse, who they said was investigating the group.

Although any potential damage will be less if ShinyHunters doesn’t publish the data publicly, the theft happening at all still presents many of those same national security risks, with the data providing granular insight into how the FBI operates.

“Since the very beginning we had made our decision that we would never publish this data. We have never intended to nor have we ever planned to,” a representative of ShinyHunters told 404 Media. (Joseph Cox / 404 Media)

Related: New York Times, Krebs on Security, Silicon UK, CNN, Newser, PCMag, Hackread, Nextgov/FCW, r/cybersecurity

A vulnerability in a Defense Manpower Data Center system allowed unauthorized users to access files containing unencrypted personal information, including Social Security numbers and military personnel data, according to a breach notification letter reviewed by Military Times.

The Defense Manpower Data Center, or DMDC, discovered the vulnerability on July 16 in a file-sharing system, according to a notification sent Sept. 18 to an individual whose information was contained in the affected files.

An analysis conducted after the discovery found that unauthorized users had accessed files on a server containing unencrypted personally identifiable information, or PII, between October 2025 and July 16, 2026, according to the letter, the authenticity of which was confirmed by two defense officials.

The unauthorized users gained access to the Social Security number of the letter’s recipient, as well as at least one additional piece of identifying information, such as a name, date of birth, contact information, sex, race, or military personnel information, including occupational specialty, according to the notification.

The notice states that the department had no indication that the individual’s information had been misused.

The potential scope of the breach remains unclear. Two people familiar with the incident told Military Times that approximately four million Defense Department personnel may be affected. (Natalie Oliverio / Military Times)

Related: WoodTV, CNN, Time, RealClearDefense

Hackers stole personal data from a Polish healthcare software provider in the latest cyberattack to hit the country’s medical sector in recent months.

Qbusoft, which develops the Medyc medical records and practice management platform, was breached after an attacker exploited an SQL injection vulnerability in August, according to a notification issued last week by one of the healthcare providers affected by the incident.

SQL injection is a security flaw that allows hackers to trick a website into giving them access to information stored in its database.

Medyc said Friday the attackers obtained names, national identification numbers, home addresses, phone numbers and email addresses. The company said it had not confirmed the theft of medical records, but an affected healthcare provider said it was informed that Qbusoft had found evidence the attackers executed scripts targeting database tables containing medical information, making it “highly likely” that they also obtained some medical records.

The Addiction and Psychiatric Treatment Center in the central city of Inowrocław said patients at its day treatment unit were affected and that potentially compromised medical information included hospital treatment records and discharge summaries.

According to the center, an unauthorized person exploited an SQL injection vulnerability in Medyc’s application interface in late August and transferred an encrypted archive of a database outside Qbusoft’s systems. The intrusion was detected overnight on September 9.

The center said the affected records covered patients treated by its Day Treatment Unit for Addiction Treatment between July 2024 and August 2026.

Some identifying information, including names and national identification — or PESEL — numbers, had been encrypted in the database, the center said. However, Qbusoft advised the healthcare provider to assume the attackers could easily decrypt the information. (Daryna Antoniuk / The Record)

Related: Terapia, Help Net Security, The Cyber Express

Keio Corporation (Keio), a major private railway operator in Japan, said its network was hit by a ransomware attack over the weekend, disrupting some of its business systems.

Following a system failure in the early hours of Saturday, the company confirmed the attack and shut down its network to prevent additional damage.

The company said it is investigating the extent of the impact and whether the attackers accessed any customer or business partner information.

The incident appears to have affected only the hospitality side of Keio’s business, not train operations.

A separate announcement published on the company’s Keio Plaza Hotel Tokyo website is warning of possible delays on some customer-facing services.

Local media outlets have reported that the cyberattack disrupted the firm's payment systems.

Tokyo Metro also disclosed a cyber incident over the weekend in which attackers gained unauthorized access to its systems and accessed 59,000 member email addresses.

Although both Keio and Tokyo Metro are Japanese railway operators, it is unclear if the organizations were targeted in a coordinated campaign by the same threat actor. (Bill Toulas / Bleeping Computer)

Related: Keio, Keio Plaza, Tokyo-np, Tokyo Metro, Cyber Press, GBHackers, Infosecurity Magazine

Japanese car-sharing service Times Car has confirmed that approximately 6.6 million user accounts were compromised in a cyberattack disclosed late last week.

The company announced the incident on September 25, saying that a third party had accessed its systems at the beginning of the month. Times Car took action to block the unauthorized access on September 26.

At the time, the company said it was investigating whether the attackers accessed members' personal information, but confirmed the data theft in a later update.

The company says that the intrusion affects 6.6 million current and former Times Car members, and also current and former members of the Times Business Service corporate account program. (Bill Toulas / Bleeping Computer)

Related: Park24, Park24

The Swiss Broadcasting Corporation (SBC), Swissinfo's parent company, has been targeted by a cyberattack in which data belonging to employees of German-language broadcaster SRF was stolen.

The data breach involves contact and organizational details of around 340 current and former SRF staff members, dating from 2020, the SBC said in a statement on Monday. These include names, job titles and professional contact details, as well as some private contact details.

There is currently no indication that personal passwords, financial or banking details, journalistic sources, research data or communications content have been affected. The data mainly relates to staff from the current Information department of the SRF unit. (SwissInfo)

Related: LIGA.net

Microsoft Threat Intelligence disclosed NeedyMantis, a modular malware framework that attackers deploy only after they already have a foothold in a network, and said it has surfaced in a small number of intrusions against telecommunications firms, universities, medical nonprofits, intergovernmental organizations and government contractors.

The activity dates back to at least October 2025 and matches what Microsoft associates with threat actors operating from China, although the company has not tied every incident to a single group.

Microsoft found the malware while pivoting from indicators linked to the DAEMON Tools supply chain compromise that Kaspersky exposed earlier this year. According to Microsoft’s analysis, at least one actor uses NeedyMantis: Storm-3069, the company’s temporary designator for the group behind the DAEMON Tools campaign. Microsoft assesses that Storm-3069 operates from China but has not attributed it to a Chinese nation-state actor.

Microsoft has also seen NeedyMantis in intrusions outside the DAEMON Tools campaign, and said the tool “might be used by more than one operator.” Those operations share targeting that aligns with Chinese interests and a pattern of selective deployment. Microsoft did not name any victims, their countries or how many organizations were hit.

The link to DAEMON Tools is narrower than it first appears. Kaspersky reported in May that attackers had served signed, trojanized installers of the disk-imaging software from its official website since 8 April 2026, profiling thousands of machines but sending a backdoor to only about a dozen government, scientific, manufacturing and retail systems in Russia, Belarus and Thailand. (Vivek / Cyber Kendra)

Related: Microsoft, Cyber Security News, GBHackers, Cyber Press

Apple released security updates to fix a zero-day vulnerability exploited in "extremely sophisticated" targeted attacks on iOS devices.

Tracked as CVE-2026-20700, this flaw stems from an out-of-bounds write weakness discovered by Meta Product Security in CoreGraphics, a framework used for two-dimensional vector graphics, image rendering, and text drawing across iOS, macOS, iPadOS, watchOS, and tvOS.

Successful exploitation of out-of-bounds write vulnerabilities can let attackers crash a program, corrupt data, or, in the worst case, gain remote code execution by writing data outside the allocated memory buffer.

"Apple is aware of a report that this issue may have been exploited in an extremely sophisticated attack against specific targeted individuals on versions of iOS before iOS 27," it warned on Monday.

"Processing a maliciously crafted file may lead to arbitrary code execution. An out-of-bounds write issue was addressed with improved bounds checking."

The complete list of devices impacted by this zero-day is extensive, as it impacts both older and newer models. (Sergiu Gatlan / Bleeping Computer)

Related: Apple

Private equity-backed cyber defense technology firm RedLattice said it’s going public via a merger with blank-check company Bold Eagle Acquisition Corp.

The deal values RedLattice at about $1.25 billion including debt before the transaction, according to a statement Monday confirming an earlier Bloomberg News report.

Aerospace- and defense-focused private equity firm AE Industrial Partners, which bought a majority stake in RedLattice in 2023, will be the company’s biggest shareholder after the deal closes. (Liana Baker / Bloomberg)

Related: Reuters, GovConWire, Silicon Angle, BizJournals, Pulse 2.0, CTech

Best Thing of the Day: Maybe All This Doom Talk Is Self-Promoting Hooey

Timnit Gebru, one of AI’s fiercest critics, believes the doom talk is about founders making money, not saving humanity.

Bonus Best Thing of the Day: Cue J.D. Vance's Catholic Theology Lecture

Pope Leo said concerns that artificial intelligence could destroy the world are not "fake news" ​and should be taken seriously, in an apparent rebuke of US President Donald ‌Trump.

Worst Thing of the Day: But Who Is Going to Identify, Disrupt and Neutralize Pete Hegseth?

Pentagon chief Pete Hegseth signed a directive under which US Cyber Command and Combat Support Agencies will prioritize and deploy advanced intelligence and cyber capabilities to identify, disrupt, and neutralize foreign interference in U.S. democratic processes.

Extra Bonus Worst Thing of the Day: Same Old Misogynistic Story

Human contractors hired to improve Microsoft’s Copilot AI chatbot are constantly bombarded with lewd or sexually explicit photo-editing requests and images that users have uploaded, including upskirt photos or putting women into sexual positions.

Closing Thought