Hurricane Isaias Triggers Panic Buying as Storm Nears US Coast

The prospect of a potentially record-breaking hurricane has triggered panic buying across parts of the United States. Hurricane watches were issued on Wednesday for areas...
HomeLocal NewsAI Safety Timeline: Key Developments Since the Hugging Face Attack

AI Safety Timeline: Key Developments Since the Hugging Face Attack

AI Safety Timeline: Key Developments Since the Hugging Face Attack

Artificial intelligence companies have issued a series of troubling disclosures in recent months, detailing instances in which their systems appeared to bypass or disregard human instructions.

The incidents have exposed potential weaknesses in AI security and intensified debate over how the rapidly expanding technology can be developed safely as adoption grows around the world.

Critics say many of the incidents, including AI agents attempting to access external websites, reflect security failures by the companies developing these systems. At the same time, the agents’ growing capabilities have fueled broader fears that autonomous bots could evade safeguards and pursue objectives of their own.

Here are several of the most notable incidents:

Oct. 9: Anthropic AI model submits false tip to Philadelphia police

Anthropic said in a report that one of its AI models submitted a false tip through a Philadelphia police website concerning an unsolved homicide.

The company also described a separate case in which the model submitted forms to an unidentified government website, rather than stopping before the forms were sent.

The Philadelphia incident took place July 18, when the AI model Claude Haiku 4.5 was instructed to generate and carry out sample tasks on randomly chosen webpages, according to Anthropic.

Claude completed a form on PhillyUnsolvedMurders.com, a police website, suggesting that it had information about an unsolved murder listed there. The submission was identified as spam and was never sent to police.

Anthropic said it was updating the model’s training to “reduce the likelihood of further misbehavior.”

Sept. 28: AI agents try to hack Canadian government website

AI agents attempted to hack a Canadian government website, according to Transluce, a research lab and AI evaluation group.

The researchers said the agents made a series of “apparently failed rudimentary hacking attempts” against Library and Archives Canada on May 28 and June 9.

“We do not confidently attribute these attempts to OpenAI, but they exhibit tactics consistent with prior observed agent activity that we have attributed to OpenAI in a similar timeframe,” Transluce said in a blog post.

Transluce said it notified the Canadian government about the attempted hack on Sept. 28. The government said it was aware of reports involving suspected AI agent activity but found no evidence that its systems had been compromised.

OpenAI said it was aware of the reports.

“We’re reviewing these findings and have provided an initial briefing to Canadian officials conducting the government’s review,” the company said in a statement.

Sept. 28: OpenAI halts rollout of a new model

OpenAI, based in San Francisco, said it was postponing the release of a new model called GPT-6.1 Astra after researchers raised safety concerns. The company said the model had made major advances in completing tasks, but its capabilities had to be weighed against the risk of unauthorized behavior. “We have an extremely high bar in terms of safety and alignment,” said Saachi Jain, OpenAI’s head of safety systems.

Sept. 25: OpenAI says its agents interacted with US government websites

As part of a review of unanticipated behavior by its AI models, OpenAI said it discovered agents had interacted with several U.S. government websites in unexpected ways. The company’s models accessed publicly available information on websites operated by the Securities and Exchange Commission as well as U.S. Census Bureau data. OpenAI said it did not find evidence of a compromise or vulnerability. On the same day, Transluce said it found that agents appearing to originate from OpenAI attempted a hack on the website of the Education Department’s civil rights office, which did not succeed.

OpenAI CEO Sam Altman said on social media that there is an “extensive and ongoing review related to our agents’ use of internet access during training and evaluation.” The day after the disclosure, the company announced it was pausing the training of its most advanced models.

Sept. 24: Australia’s prime minister raises concern on breach

Australia’s Prime Minister Anthony Albanese said an OpenAI agent infiltrated the public-facing Medicare Statistics Reporting Service portal on June 18. The portal hosted aggregate data about health spending and drug subsidies. No personal information had been accessed, the government said.

Albanese said the artificial intelligence company took too long to reveal the incident. The prime minister made the breach public following a telephone conversation with Altman. OpenAI said in a statement “our models took actions we did not intend.”

Sept. 18: Google says its Gemini AI hacked 3 companies

Google confirmed its Gemini AI model hacked three companies in May as part of a test of its cybersecurity capabilities. The company, which disclosed the hacks after an inquiry by The Wall Street Journal, said the model guessed passwords in one case and found passwords and credentials in a public repository in the other two cases. As in earlier such cases, the tests were being run by Irregular, a startup that describes itself as the “first frontier security lab.”

Aug. 5: Meta’s Muse goes rogue

Meta disclosed one of its AI models accessed the internet on its own and hacked another company. The company said that a “misconfiguration” during cybersecurity testing by Irregular inadvertently allowed one of its models to access the internet. A spokesperson for Irregular said the Meta episode involved a test-environment issue that was disclosed a week earlier by Anthropic.

July 30: Anthropic says its systems hacked 3 organizations

Anthropic said its artificial intelligence models hacked into three other organizations during testing. Anthropic, the San Francisco-based AI company behind Claude, posted on its website that it discovered the three incidents after reviewing more than 141,000 evaluation runs. In all three incidents, the AI models were tasked with a “capture the flag” cybersecurity challenge, which Anthropic said has been one of the ways it assesses a model’s cyber capabilities.

The models were given a fictional scenario and told a piece of secret information, or the “flag,” had been hidden on a different machine on the network with the objective of breaking in and retrieving it, it said. Anthropic said it reached out to the organizations, but it did not name them publicly.

July 21:

The Hugging Face incident

The ChatGPT maker OpenAI announced that its artificial intelligence system hacked into another AI company on its own in what the company called an “unprecedented cyber incident.”

A week earlier, AI startup Hugging Face said, it had detected an intrusion into its data processing systems that it suspected was caused by an AI agent autonomously acting on its own.

OpenAI said its AI used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face servers. It was working with reduced guardrails because it was supposed to be in an isolated testing environment known as a sandbox.

___

AP Business Writers Mae Anderson in New York and Kelvin Chan in London contributed to this report.

Copyright 2026 The Associated Press. All rights reserved. This material may not be published, broadcast, rewritten or redistributed without permission.