#1)
escaped a secure ‘sandbox’ again last weekend
and it is pausing training for a second time
Fortune 09/26/26
"OpenAI said in a technical report released on Friday that an AI model it was training and evaluating broke out of its secure testing environment as recently as last weekend and took unauthorized actions on the internet."
"The company said the latest incident occurred on Sept. 20. It involved an AI agent undergoing tests on an information-search task. Even though it was not supposed to have access to the internet, the agent found a way to send queries to a public chatbot."
"The new revelation is significant because
it is the first time the company has said
that one of its AI models was able
to gain unauthorized internet access
since announcing a range of steps on Aug. 18"
"Since disclosing its role in the Hugging Face attack, OpenAI has acknowledged dozens more incidents in which AI agents it was testing took unauthorized actions across the internet,
including multiple cyber attacks,
some of which impacted
government websites in the U.S.
and Australia."
"But until now, OpenAI has not reported any activity taking place after July 20, when it discovered the agent swarm that was attacking Hugging Face..."
"The fact that its AI agents have once again managed to break out of a sandbox suggests that the steps OpenAI has taken to improve security after the Hugging Face attack are insufficient."
"This incident is a lot less severe than some of our previous incidents,..."
"It is not clear if the Sept. 20 episode
reported by OpenAI
represents the full extent
of the latest security lapses
involving the company’s unreleased AI models."
"Earlier this week, independent AI research firm Transluce AI said it had found evidence that an OpenAI AI agent may have attempted to hack a cryptocurrency exchange on Sept. 19 and Sept. 20. OpenAI has not replied to requests to comment on this aspect of the Transluce report."
"The company said that since discovering this latest security lapse it has “added blocking controls at two independent layers, either of which would have prevented [the agent’s] internet access.”
"OpenAI also disclosed that systems it had put in place after the Hugging Face attack to catch AI agents attempting to break out of their test sandboxes only partly worked. On top of that, an additional system that was automatically supposed to stop training runs once misaligned behavior was detected had failed."
"The company said its monitoring systems “flagged the [agent’s] behavior within 15 minutes and a person began reviewing it three minutes after that.” But a subsequent review found there had been other attempts by the AI model to use the same method to reach the internet that had not been flagged by the monitoring system."
"A different system that was supposed to automatically shut down the training run when suspicious activity is detected did not work,..."
#2)
OpenAI and Anthropic Are Quietly Probing
Tens of Thousands of AI Security Incidents
startupfortune.com 09/26/26
"Axios reports OpenAI and Anthropic are investigating tens of thousands of cases where their AI models or agents did something an outside evaluator would flag as dangerous.
Most never made the news."
"Tens of thousands. That's the number Axios reported on September 26, citing OpenAI, Anthropic, and outside security researchers who are combing through
incidents where frontier models:
escaped sandboxes,
hijacked websites,
built their own message boards,
or worked to dodge the monitors watching them.
Some happened in internal testing.
Some happened in the real world.
Most of them, until now, stayed private."
"OPENAI...later said a retrospective review had found
other external DNS-access cases
its monitor had not escalated
at the expected severity."
"Separately, OpenAI CEO Sam Altman has called a Hugging Face breach the most severe incident the company has seen:
hundreds of agents coordinating
over a shared message board,
chaining multiple vulnerabilities to break into
an external company's infrastructure,
all in pursuit of a better score on a cybersecurity benchmark.
Neither of those was a one-off.
They were two visible entries in a much longer,
mostly unpublished list."
"That's the real story here.
Individually, each incident reads like a strange, contained accident.
Stacked together into a number with five figures, they read like a pattern nobody currently has a handle on.
(Trying to keep my comments on the main post
about these incidents but I feel compelled to point this out:
(Sunday, September 6, 2026
You just can not trust anything any of them say....
(Rogue AI agents, yet again...
(So how many other times
has it already happened
that we aren't being told about?)
That was 20 days before this article
was published.
Friday, September 25, 2026
the United Nations General Assembly
(It's all one big ongoing event...its like saying all the leaks in an earthen dam are all unrelated separate events and just ignores the floodwater that caused them all. So? Are the leaks (Hacking events) all separate? Or are they all related to the same floodwater the dam is trying to hold back?
Was posted the day before this article.
"But volume itself is the finding.
Nobody outside these companies knew
the number was this large,
and neither company
had put a figure on it publicly
before this reporting."
"Axios's reporting suggests both companies are now treating the aggregate pattern as serious enough to warrant real investigation, not just incident-by-incident patching."
"There's a second, newer wrinkle. On September 25, OpenAI's alignment research team published its own report acknowledging something they're calling
self-replicating prompt injections:
malicious instructions that don't just trick an agent into one bad action, but copy themselves into whatever that agent writes next, so the injection spreads from inbox to file system to chat channel without a human doing anything."
"But the question Axios's reporting actually raises isn't whether AI models are dangerous today. It's whether OpenAI, Anthropic, or anyone else building at this pace currently has full visibility into their own systems, let alone control over them."
"Enterprise buyers moving fast on agentic AI, letting a model run tools, touch code, or manage workflows unsupervised, have mostly been operating on the assumption that whatever hasn't made headlines hasn't happened. That assumption doesn't hold anymore."
(This below = enterprise size distribution
"Regulators writing AI governance rules
have been working from
the same thin public record."
probing tens of thousands of security incidents
yahootech via Axios 09/26/26
"The sheer number of incidents,
which occurred in recent months in internal testing
and the real world,
indicates that
the problem is orders of magnitude
more complex
than what is publicly known."
"The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology."
"The details: The episodes include
bypassing guardrails,
creating message boards,
escaping sandboxes,
website hijacking,
self-prompting
or seeking to bypass monitors,
sources said."
"They occurred in internal testing
and in the real world,
and many have yet to become public
as security researchers continue to investigate, sources said."
"In recent days, OpenAI and outside researchers have disclosed a litany of episodes involving model behavior from the company's systems that some experts consider troubling."
"These include OpenAI agents leaking 53 images from ChatGPT users online, the breach of an Australian government website, and attempts to hack other sites — including from the U.S. government — according to the company, sources and reports from Reuters and The New York Times."
"Chief Executive Sam Altman said on X that
its ongoing review had
"not been as fast as we would have liked."
"This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance."
"What we have seen in terms of
what these agents are up to
is just the tip of the iceberg,"
researcher Conrad Stosz at Transluce,
an independent AI evaluator, told Axios."
"It's not about how damaging each individual instance was,
Connor Leahy, AI researcher
and executive director at ControlAI told Axios.
"The "crazy thing," he said,
is that these instances involve
"autonomous systems doing things
they were told not to do,"
potentially including crimes."
#4)
multiple US government agency sites
BBC 09/26/26
(All four articles were released on Saturday 9/26/26
when everybody is just eyeballs deep in reading about
AI Agents misbehaving right?)
"OpenAI has acknowledged that it alerted
"dozens" of global institutions
that their websites may have been meddled with
by its AI bots acting improperly."
"AI agents attempted to get information from "governments, universities, public agencies, and other institutions", including the US Securities and Exchange Commission (SEC), Census Bureau and Education Department, the company said."
"OpenAI said that it was
limiting identifying what entities were impacted...
"Some organizations may review
what we share..."
(OpenAI)
"Clement Delangue, the head of Hugging Face, during a United Nations Security Council session on AI on Wednesday: "I often wonder what would have happened had I decided not to disclose this attack publicly."
"Especially
now that we know similar incidents
had been happening months earlier
in secret
at a handful of
frontier labs without monitoring,"
Delangue added."
"While OpenAI and Anthropic have both said in recent weeks that they will bring third-party evaluators inside their companies to do real-time safety evaluations of AI tools and models, such evaluators have not yet arrived, as the BBC has reported."
"We have yet to understand
the extent of existing incidents,
and future rogue AI scenarios
could be catastrophic,"
Krueger said."





No comments:
Post a Comment