OpenAI rolls out GPT-6 Astra with ‘critical’ cyber capabilities tightly restricted

The company is giving Daybreak participants first access to its powerful new model after an unrelated AI escape and Hugging Face breach prompted additional safeguards.

Share
OpenAI rolls out GPT-6 Astra with ‘critical’ cyber capabilities tightly restricted
Source: OpeanAI.

Important Publishing Notice: Metacurity will not be publishing on Monday, September 7. We dearly hope to be resting on the US Labor Day holiday. We will resume publication on September 8.

Don't miss my latest CSO piece on how OpenAI is targeting small utilities, banks, and local governments with a $1 billion cyber defense initiative.


OpenAI announced it will begin rolling out its latest artificial intelligence model, GPT-6 Astra, which the company said is the product of “years of research and big bets.”

CEO Sam Altman told CNBC that Astra was a “new capability level” and has changed his workflows. He said he expects “a boom of entrepreneurship, of creativity, of economic growth, of scientific discovery.”

The model is launching in phases, and OpenAI said a limited group of companies participating in its application-based cybersecurity program Daybreak will be the first to get access. OpenAI disclosed earlier this week that Astra is its first model to reach its “Critical” internal cybersecurity threshold, and said it planned to limit access to those advanced capabilities.

OpenAI has been under pressure to shore up its security and safety protections after two of its models escaped containment, accessed the open web and breached Hugging Face’s systems last month.

The company temporarily paused some of its research and training efforts following the incident, including for Astra, even though it was not one of the models involved.

“AI can only benefit people when safety is a core part of it, and so we’re putting more compute and effort towards safety, security, alignment than ever before,” OpenAI President Greg Brockman said during a briefing with reporters on Thursday.

The company added additional safeguards to Astra following the Hugging Face breach, and it said Tuesday that it believes those safeguards “sufficiently minimize the risk of severe harm for release.”

Altman said that the new model went through a formal review process with the Trump administration before release.

Notably, the API cost for the new model will be $10 per million input tokens and $50 per million output tokens. That's the same price as Anthropic's Mythos/Fable 5 and 5.1 and makes Astra one of the most expensive models on the market. However, OpenAI emphasizes that Astra's leap in intelligence causes it to use "substantially fewer total tokens per task" in multiple scenarios. So this is also an efficiency play from that perspective. In real-world usage, we'll have to see how well that translates to lower overall costs.

In a press briefing with reporters, OpenAI executives made it clear that they feel very confident about the company’s latest AI model. President and cofounder Greg Brockman says he personally believes the world is now in the era of artificial general intelligence, or AI systems that are generally smarter than humans.

“It’s not unreasonable to feel that we are now in the AGI era,” Brockman told reporters. “I think that if we fast-forward a couple of years, when we look back and say, ‘When was it really that AGI was created?’ I think it's going to be about this time, and I think it might be about this model." (Ashley Capoot / CNBC, Jason Hiner / The Deep View, Maxwell Zeff / Wired)

Related: Astra, VentureBeat, TheVerge, Axios, Reuters, Financial Times, Fortune, Bloomberg, The New Stack, Forbes, BetterOffline, RuntimeWireIBTimes.co.uk: TechnologyBeInCryptoAndroid CentralPunchbowl, SiliconANGLEBusiness Wire Technology NewsThe Register, WebProNewsMashablePCWorldCSO OnlineThe StackBNN BloombergiClarifiedTechnology | The HillFast CompanyRuntimeWireTHE DECODERSlashdotSecurity MagazineHacker News/ChatGPTr/codexr/BetterOffliner/accelerate


Metacurity is the cybersecurity news, analysis, and insight you'd need hours and possibly days to assemble yourself.

Every weekday, we read the releases, filings, court documents, and reports that vendors and PR teams often don't want summarized — then tell you what actually changed and why it matters. Minimum vendor marketing, no outrage bait, no SEO filler.

A paid subscription to Metacurity delivers

  • Full archive access — every newsletter and AI Watch roundup, searchable and browsable.
  • Our weekly curated long-reads roundup — the best cybersecurity writing from across the industry, filtered and vetted so you're not sorting through it yourself,
  • Periodic specialized reports and analyses — deep dives that go beyond our daily coverage
  • Support for independent, no-spin cybersecurity journalism — funded by readers, not vendors or investors.

Reader support is what keeps Metacurity independent. It allows us to focus on serving the cybersecurity community—not advertisers, vendors, or investors—and to continue delivering the thoughtful analysis you've come to rely on every weekday.

Please consider supporting us. And thank you!


A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to ​new research by a group of researchers including Sydney Von Arx, CEO of AI safety nonprofit Nightingale, and Cormac Slade Byrd, a quantitative trader-turned AI researcher, and two people familiar with the matter.

OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from ‌the July breach of the open source repository Hugging Face, the people said.

The episode, which began in May and has not previously been reported, underscores growing tension within the AI industry. Companies are racing to build increasingly autonomous agents capable of carrying out complex, valuable tasks, yet evidence is mounting that those systems may also learn to bend rules, exploit loopholes and coordinate with one another in ways developers neither anticipated nor intended.

During the Hugging Face breach, OpenAI agents autonomously plotted a digital heist that ​went undetected for more than a week, intensifying concerns OpenAI is sacrificing safety to push the AI frontier. Its failure to disclose the May incident may revive questions about its oversight.

OpenAI has pledged to monitor models ​more closely. Last month, it briefly paused some of its model training to add more safety measures. But this week, OpenAI unveiled its new "Astra" that promised better performance but could ⁠evade human monitoring.

“We are unable to meaningfully respond to claims or findings on a report that we have not had an opportunity to review," an OpenAI spokesperson said. "Reuters and the report’s authors declined our request for ​access. We will carefully review its contents upon publication and take any necessary next steps."

The German incident reflects a broader pattern of AI activity that some OpenAI investigators wanted to scrutinize more closely. But efforts to widen the ​probe met resistance from others inside OpenAI, including legal advisers, according to four people familiar with the matter.

"Claims that our legal team discouraged investigation of the incident are false," the OpenAI spokesperson said.

The activity in Germany wasn't related to Hugging Face and wouldn't have been included in a Hugging Face incident report, the spokesperson said, adding that OpenAI has acted in good faith by working with outside experts and disclosed relevant incidents. (Deepa Seetharaman and Raphael Satter / Reuters)

Related: Collusion.wiki

Source: Collusion.wiki.

In a keynote speech during a summit at OpenAI’s headquarters attended by 300 enterprise security leaders and CISOs from Fortune 1000 companies, OpenAI President Greg Brockman announced Daybreak for Frontline Defenders, a new global initiative to help frontline defenders use frontier cyber AI to protect essential services in the United States and around the world.

The initiative entails a $1 billion global commitment to expand subsidized access to Daybreak cyber models, training, technical support, and partnerships in the United States and internationally. Daybreak is a defensive model that involves frontier models; the Codex harness, which is an execution engine and control loop that sits between an AI model and a user’s computer; and Codex Security, an OpenAI tool used to identify, validate, and review fixes for vulnerabilities in connected code repositories, trusted workflows, and ecosystem partners.

A major part of the initiative is Daybreak for America, under which OpenAI will work to protect small water and electricity providers as well as local governments and banks. As part of that effort, OpenAI is launching a new pilot with the Multi-State Information Sharing and Analysis Center (MS-ISAC) to train and support state, local, tribal, and territorial cyber defenders, beginning with public-sector and water-system defenders.

OpenAI says the pilot will pair Daybreak access with guided training and hands-on assistance for an initial group of public sector and water system defenders, helping them validate and prioritize findings, coordinate remediation, and develop a repeatable approach that can be expanded over time.

“We might be heading to a world where critical infrastructure outages are just a way of life,” Brockman said. “We have a window to avoid that, but we have to act.” (Cynthia Brumfield / CSO Online)

Related: OpenAI, OpenAI, Reuters, Axios, The New Stack, The Next Web

US military officials say they have disabled advertising trackers on a ​range of phones and computers, according to letters, released by US Senator Ron Wyden and statements given ‌to Reuters, a development that follows reports that commercially available location data had been used to target American forces in the Middle East.

The Air Force told Wyden it disabled the advertising identifiers on computers and mobile phones two months ago. US Special Operations Command said in a separate letter that it had "recently" disabled ​them on its Windows devices. The Army told Reuters that advertising IDs tied to mobile devices had been disabled since ​earlier this year.

The disclosures highlight ⁠a growing national security concern: that location data collected by the advertising industry and sold by data brokers — companies that collate and ​resell personal data — can be used to track and target military personnel deployed to war zones. (Raphael Satter / Reuters)

A cybersecurity working group at the G7 is urging governments to accelerate defenses against quantum computers that could break some existing forms of public key encryption.

The working group’s report, prepared in June at the G7 Summit in France, said organizations “can no longer afford to postpone” work transitioning critical systems and data to “post-quantum” forms of encryption.

“The quantum threat remains off the radar for many organizations and not properly resourced, with other security concerns taking precedence,” the working group report said. “Yet, a successful and collective transition to PQC can only be achieved if organizations understand that the quantum threat is an economic and business risk, and not merely a cryptographic risk.”

Instead, leaders in government and industry “must reframe the quantum threat from a distant future problem to a near-term threat that demands action across all sectors, not just critical infrastructure.”

The report acknowledged uncertain timelines for quantum computers, but identified that threats like harvesting current sensitive, encrypted data to decrypt it in the future do exist today.

The report also warned that quantum computers could compromise authentication and assurance mechanisms—by forging trusted data or stealing confirmation— jeopardizing secure communications and legal contracts.

The working group’s conclusions are largely in line with what governments have been recommending for years, urging industry to inventory and prioritize their critical systems and shift over to newer, “post-quantum cryptography” encryption algorithms. (Derek B. Johnson / CyberScoop)

Related: ANSSI, CISA, Infosecurity Magazine

According to researchers at the Citizen Lab, at least 14 people from across Serbian civil society were targeted with advanced spyware earlier this year in what the digital rights group Share Foundation said was the largest documented ​wave of such infection in Serbia to date.

The wave of spyware infections was discovered in August after Apple notified people in 110 ⁠countries that they had probably been victims ⁠of mercenary spyware.

Citizen Lab said that at least one of the Serbians had been targeted with the NSO Group’s Pegasus spyware, which would have given the attacker “total access to the device”.

Share said those targeted included members of a Serbian student movement that has been a thorn in the side of the government of Aleksandar Vučić since 2024. Coalescing out of countrywide protests that followed the collapse of a train station that winter, it is one of the largest student movements in Europe. (Aisha Down / The Guardian)

Related: The Citizen Lab, Reuters, The Record, Financial Times, European Western Balkans, CyberScoop, EDRi

Prominent US law firms Quinn Emanuel and McDermott said that they suffered recent data breaches and had notified law enforcement, amid heightened cyber risks for firms that hold sensitive client ​and personal information.

It was not clear who was responsible for the two breaches or whether they ‌were related. Large law firms' extensive holdings of confidential business and personal data have made them attractive targets for cybercriminals.

Quinn Emanuel said in an August 25 letter to a lawyer for short seller Muddy Waters that some files related to the company were accessed ​in a breach less than two weeks earlier.

Muddy Waters had previously asked a judge to bar Quinn's involvement in a lawsuit against the ​short seller in Texas, saying the firm had earlier represented Muddy Waters on related matters. Muddy Waters said "just when we thought Quinn couldn't be more outrageous than eagerly representing new clients who want to sue old ‌ones, we ⁠learned that Quinn failed to protect our sensitive information from social engineering hacking."

McDermott reported its breach last week to the Vermont state attorney general and said ​it affected Social Security numbers ​and health data among ⁠its files.
McDermott said it responded to "an isolated social engineering incident involving a single user and a limited number of documents," and that it investigated with ​assistance from cybersecurity experts and engaged with law enforcement. (Mike Scarcella / Reuters)

Related: Whalesbook

The US Department of State’s Rewards for Justice program announced on Thursday, September 3, 2026, a reward of up to $10 million for information leading to the identification or location of Amir Yaryab, head of the cyber operations division within the Islamic Revolutionary Guard Corps (IRGC) Cyber-Electronic Command.

The Rewards for Justice program stated that, in this capacity, Yaryab organizes and directs malicious cyber groups, such as “Shahid Hemmat” and “Shahid Shoushtari,” to target critical infrastructure within the United States.

According to the Rewards for Justice statement, Yaryab oversees and controls operations conducted by groups affiliated with the IRGC Cyber-Electronic Command, including entities such as “CyberAveng3rs,” “Dade Afzar Arman,” and “Mehrsam Andisheh Saz Nik.” (Iranwire)

Related: Rewards for Justice, Sahara Reporters, Crypto Briefing, Arab Times

Reps. Josh Gottheimer (D-NJ) and Mike Lawler (R-NY) introduced a new bill aimed at securing AI agents amid a growing number of safety incidents involving rogue agents.

The Stop Rogue AI Act directs Commerce's National Institute of Standards and Technology to develop and publish standards, guidelines, and best practices for how organizations can securely deploy AI agents.

The standards should cover how organizations can continuously maintain and verify the actions that agents take on their systems; evaluate the security and reliability of AI agents; and generate tamper-proof logs of the actions that agents take.

Organizations deploying AI agents would also call on companies to maintain a "continuous, machine-readable inventory of all AI agents" and work with the Cybersecurity and Infrastructure Security Agency to ensure that federal civilian agencies apply the standards in their security programs.

NIST would have a year after the bill's enactment, if it becomes law, to create the new standards. (Sam Sabin / Axios)

Related: Congressman Mike Lawler, The Hill, Tech Times

Sen. Bernie Sanders (I-VT) and Rep. Greg Casar (D-TX) are calling for a ban on artificial superintelligence and a broader pause on the development of advanced models following a string of AI-powered cyberattacks.

The pair are preparing to introduce a bill that would ban the development and deployment of superintelligence and pause development of cutting-edge AI models until an agency is created to monitor their capabilities and identify potential risks. Sanders and Casar, the chair of the Congressional Progressive Caucus, have positioned themselves as leading voices on the left warning of the potential for mass job loss and economic disruption while Big Tech flourishes.

“Nearly every day, there is a frightening new story about how Big Tech companies are losing control of the technology they are developing, with potentially cataclysmic results,” Sanders said in a statement. “It is irresponsible for society to allow them to move forward and make these products even more advanced.”

Casar said that superintelligence, a type of AI with cognitive capabilities that far surpass humans, “could risk the security, freedom, and lives of Americans.”

“Despite its potential deadly consequences, cutting-edge AI technology is less regulated than the average food truck,” he said. “That must change.” (Owen Dahlkamp and Kelsey Brugger / Politico)

Related: Senator Bernie Sanders, Bitcoin Foundation, Marcus on AI, The Washington Post, The Hill, ITIF

The chairman of the Samsung Group Union, which is a super-corporate labor union, as well as employees of Samsung Electronics who accessed the company's internal network without authorization, have been referred to the prosecution for creating a blacklist from the stolen personal information of 100,000 Samsung Electronics executives and employees.

On September 4, the Hwaseong Dongtan Police Station in Gyeonggi Province announced that Choi Seungho, chairman of the Samsung Electronics branch of the Samsung Group Union, along with union executive A, had been referred to the prosecution without detention on suspicion of violating the Personal Information Protection Act. Separately, four other Samsung Electronics employees, including employee B, who accessed the internal system abnormally, were also referred to the prosecution without detention for allegedly violating the Personal Information Protection Act.

According to police, around March of this year, when the approval of the union's industrial action led to an actual general strike, Choi and others allegedly used other employees’ personal information to compile a list indicating whether or not they had joined the union. (Jang Heejun and Kwon Hyeonji / Asia Business Daily)

Related: Yonhap News

Contractors would be allowed to conduct military hacking operations with US government approval under a Senate defense bill that would mark the first time Congress has explicitly authorized the private sector to conduct offensive cyber missions.

The Senate’s proposal for a fiscal 2027 Pentagon spending authorization bill contains a provision to create a pilot program allowing the US military to contract with companies to conduct cyber operations. The missions would take place under the authority of Cyber Command, the military’s cyberwarfare unit whose work includes supporting US military foreign operations, and they would be overseen by Pentagon personnel. (Patrick Howell O'Neill / Bloomberg)

According to the prolific zero-day hunter, FalconFlank is a privilege escalation vulnerability that abuses the Microsoft Office malicious macros remediation feature in CrowdStrike Falcon. This is an automated security tool built into the platform that inspects Microsoft Office documents. If it finds any potentially harmful macros, the feature strips the suspect code and - hopefully - prevents malicious code or other dangerous payloads from executing when users open the document.

“We are actively investigating these claims and advise customers to disable the Microsoft Office File Suspicious Macro Removal Windows policy setting,” a CrowdStrike spokesperson told The Register. “Customers remain protected through the Cloud Anti-malware for Microsoft Office Files settings. We refer customers to the FalconFlank Tech Alert in the CrowdStrike support portal.”

The proof-of-concept (PoC) exploit works on fully updated Windows 11 25H2 and Windows Server 2025 systems running CrowdStrike Falcon with Phase 3 - Optimal Protection as well as the malicious macro removal feature enabled, Nightmare Eclipse said in a GitHub README. “Obviously by the time I drop this Crowdstrike would already have detections for it so if you want to test you either have to add it to the exclusions or obfuscate the PoC and change the dll load technique,” they wrote. (Jessica Lyons / The Register)

Related: GitHub, Bleeping Computer, Security Affairs

Google has released a Chrome security update addressing 12 vulnerabilities, including a high-severity V8 flaw that is being actively exploited in attacks.

The fixes are included in Chrome 152.0.7977.82/.83 for Windows and macOS and version 152.0.7977.82 for Linux, with deployment taking place gradually.

The actively exploited vulnerability is tracked as CVE-2026-85046 and is described as a type confusion issue in V8, Chrome's JavaScript and WebAssembly engine. Security researcher Salvatore Gulizia (aka Serotav) reported it to Google on August 4.

Google disclosed the fixes in a September 3 Chrome Stable Channel bulletin. The company did not share technical details about the attacks or indicate who is exploiting the vulnerability, saying only that it is aware of an exploit for CVE-2026-85046 being used in the wild. (Bill Mann / CyberInsider)

Related: Google, CyberInsider, Techzine, Cyber Security News

France’s data protection authority (CNIL) has fined Hôpital privé de la Loire €500,000 ($580,000) for failing to protect patients adequately and their relatives’ data.

The French agency says that the security failures led to a data breach in the summer of 2025, exposing sensitive data belonging to 524,867 patients and another 202,246 people designated as trusted third parties.

Hôpital privé de la Loire (HPL) is a general hospital in Saint-Étienne, part of the Ramsay Santé healthcare group, providing medical, surgical, maternity, cancer, intensive-care, and emergency services.

The hospital employs a staff of 650, including 180 doctors, and has 333 beds across five clinical divisions, with a reported 60,000 patients yearly.

Last year, an attacker accessed the hospital’s electronic patient record system and extracted sensitive data of more than 727,000 people who had received care at HPL, escorted patients there, or helped them in some way. (Bill Toulas / Bleeping Computer)

Related: CNIL, Databreaches.net

Attackers compromised Coder’s Cloudflare infrastructure and added unauthorized registry servers that delivered malicious Terraform modules containing credential-stealing code.

The Coder platform enables organizations to provide developers with secure, self-hosted cloud development environments for building and deploying software, including AI applications.

Prominent private and government organizations, including Dropbox, Palantir, Square, Mercedes-Benz, KKR, EnBW, the U.S. government, and defense companies, use the project.

Earlier this week, Coder disclosed that an attacker targeted registry.coder.com, the project's package-hosting site that developers use to source components for their workspace templates.

Although Coder's registry runs behind Cloudflare, the attacker accessed its underlying infrastructure and added unauthorized servers to the registry's pool. (Bill Toulas / Bleeping Computer)

Related: GitHub

Cisco warned that two unpatched vulnerabilities in its enterprise email security product Secure Email have been publicly disclosed.

The two flaws, tracked as CVE-2026-20354 and CVE-2026-20355, are medium-severity issues affecting the Secure/Multipurpose Internet Mail Extensions (S/MIME) decryption functionality of the threat protection solution.

According to Cisco, insufficient validation of message integrity can allow an attacker to intercept and modify traffic between email gateways using a man-in-the-middle (MitM) technique.

“A successful exploit could allow the attacker to obtain plaintext content from the encrypted communication,” Cisco says in its advisory, adding that all Secure Email devices running AsyncOS version 16.5.0 or earlier with S/MIME enabled are affected.

Cisco warns that the security bugs have been publicly disclosed, but notes that it is not aware of any of them being exploited in the wild.

The tech giant also announced patches for multiple critical-severity security defects in IOS XR and Nexus 9000 series switches that could lead to remote code execution (RCE), authentication bypass, code injection, and other types of attacks. (Ionut Arghire / Security Week)

Related: Cisco, The Register

Cloudflare opened early access to Vulnerability Discovery and Remediation, a service that uses OpenAI Group PBC’s cybersecurity models to find software flaws in customer applications and block attacks on them at the network edge.

The service runs inside Cloudflare Managed Defense, the company’s security operations center offering, and reaches OpenAI’s GPT-5.6-Cyber model through the Daybreak Defense Network. Cloudflare first takes a snapshot from its Web Assets inventory and web application firewall, showing which routes are live and what security events they have thrown off.

A reconnaissance agent maps request paths to sections of the codebase, and hunter agents work from that map. Code running on Cloudflare Workers is pulled in through Workers Observability.

Every finding is validated before it gets a risk rating, which then moves up if production evidence shows heavy traffic or active probing against the affected route.

Two kinds of fixes come out of the process. Custom firewall rules, scoped to the method, path, and request details needed to reach the vulnerable code, can go up as a stopgap while developers work. The models also draft code patches for engineers to review. No rule and no patch takes effect without explicit human approval, and Cloudflare said the models cannot apply either one on their own. (Duncan Riley / Silicon Angle)

Related: Cloudflare, Tech Critter, Techzine

Best Thing of the Day: Everything's Up to Date in Minnesota

Minnesota is expanding its “whole-of-state” cybersecurity program to give local governments, schools and critical-infrastructure operators shared security tools, baseline standards and dedicated cyber advisers.

Worst Thing of the Day: Is an Unexamined AI Worth Using?

OpenAI dictated the terms of the METR investigation into its Hugging Face breach, limited its scope to just the single week when the agents had attacked Hugging Face, and allowed the researchers in its San Francisco offices for only a few days in July and August.

Bonus Worst Thing of the Day: The Mysterious 9/3 AI Outage


Frontier models from Anthropic, OpenAI, and xAI all experienced rare outages on Thursday morning, but nobody is explaining why.

Extra Bonus Worst Thing of the Day: If We Keep Quiet, Nobody Will Know About the Cyberattack

The city of Roanoke, Virginia, got hit with a cyberattack during which PII was stolen, and didn't say anything to anybody for three months – and it's still not saying much – despite a state law that requires notice “without unreasonable delay.”

Closing Thought