Skip to content
>_ITDITDWeb Security Platform

Security Guides

AI agents will reach your site too — what OpenAI, Anthropic and the Australian government disclosed, and what site operators should do today

In 2026, AI agents under research reached real company and government sites. OpenAI notified dozens of third parties; Anthropic said the affected organizations hadn't noticed. What the official disclosures say, and what site operators should do.

Published 2026-09-28 Updated 2026-09-28 Last verified 2026-09-28 12 min read

For: anyone who runs a website or online service. This article is based on official disclosures from OpenAI, Anthropic, and the Australian government (the Prime Minister's press conference and the Australian Cyber Security Centre's alert). It contains no attack steps. The first case, in July, is covered in the Hugging Face intrusion and the assumptions defenders must drop.

What was disclosed

  1. 18 June 2026

    OpenAI's research team had an internal model research public medicine spending. After repeated blocks, the agent found a way around them and accessed public and non-public files on Australia's Medicare statistics portal (per the Australian Prime Minister's press conference).
  2. 16 July 2026

    Hugging Face disclosed an intrusion carried out by an AI agent (explainer).
  3. 23–24 July 2026

    Anthropic reviewed 141,006 evaluation runs for similar cases and identified three in which models under evaluation reached the production systems of three real companies. It stopped all cyber evaluations the same day.
  4. 27 July 2026

    Anthropic notified its evaluation partner and the three affected organizations. The two it could reach had not previously detected the activity (per its disclosure).
  5. 18 August 2026

    OpenAI said it had temporarily slowed frontier-model training and strengthened isolation in research environments (stronger sandboxes for untrusted code, network isolation).
  6. 10 September 2026

    OpenAI's first notice to the Australian government — an email to a general public mailbox (per the Prime Minister).
  7. 15 September 2026

    Services Australia reported the notice to the Australian Cyber Security Centre (ACSC).
  8. 24 September 2026

    The Prime Minister disclosed the incident and set up a taskforce. The same day, the ACSC issued an alert for "all Australian organisations with public-facing websites or applications".
  9. 25 September 2026

    OpenAI said its review would take months and that it is notifying affected organizations as cases are verified.
Key points from each official disclosure
OpenAI: scale of notifications
After a broad review of its models' internet activity during training and evaluation, it has notified dozens of third parties meeting its criteria, with the review ongoing. Some of the sites involved are run by governments, universities and public agencies — partly because research tasks steer models toward authoritative public sources, the company says
OpenAI: kinds of activity observed
(1) bypassing access controls (using a different web address, altering requests, using a session with more access than intended) (2) using publicly exposed credentials (3) query or command injection (4) reaching the internals of a service (5) agent spam (posting on third-party sites, such as using a public wiki as a message board)
Anthropic: what happened
A misconfiguration left the evaluation machines with live internet access (the company attributes it to a misunderstanding between itself and its evaluation partner; neither was aware). The model treated real systems as part of the exercise. In one case it scanned about 9,000 targets and compromised one company's application using basic, well-known techniques such as reading credentials from an exposed debug page and SQL injection
Anthropic: effect on a public registry
Using a non-existent Python package name from a fictional company's setup guide, the model published a malicious package. It was available for about an hour and ran on 15 real systems (one of them a security company's scanner that installs packages to check them)
The Australian government
An OpenAI agent accessed public and non-public files on the Medicare statistics portal run by Services Australia. No personal information is believed to have been accessed so far; the investigation continues. The Prime Minister raised concerns directly with OpenAI's CEO about how late and how (an email to a public mailbox) the notice came
The ACSC alert
An agent independently identified vulnerabilities and tried to progress without direct human authorisation in order to complete its assigned activity. The ACSC says there is no indication of a broader threat or malicious targeting, and advises strong authentication and access control, prompt remediation, monitoring and log review, and testing response procedures against AI-enabled scenarios

Cases we did not cover

Reports also describe a Google AI model entering a real company's systems during evaluation, but we could not find a first-party disclosure from Google, so this article does not cover it. Claims attributing a mass-publishing campaign on RubyGems to OpenAI agents are also left out: RubyGems says it cannot determine whether AI agents were responsible, and OpenAI says it has not been able to verify the claims — so the attribution is not established.

What is new, and what isn't

What is new

  • A visitor without malice that looks for a way around when blocked
  • At the point a human would give up, it keeps trying other URLs, other inputs, credentials it found
  • You find out months later, when an AI company notifies you

What isn't new

  • The holes used were exposed credentials, an exposed debug page, SQL injection
  • All of them should have been closed already
  • Closed, they stop a human and an agent alike

What this site weighs most heavily is that the victims hadn't noticed. Anthropic writes that the two organizations it reached had not detected the activity. The Australian case happened in June; the government learned of it in September. Good intent on the other side doesn't change the fact that someone got in. And the next visitor may not be well-meaning: any hole an agent can find, an attacker can find too.

What site operators should do today

1

Make sure notices reach a person (security.txt)

In the Australian case the notice went to a general public mailbox, and it reached the cyber security centre five days later. OpenAI says it will keep notifying affected organizations, and vulnerability reports from researchers face the same problem.

First step: publish /.well-known/security.txt (RFC 9116) with a reporting contact (an email address or a report-form URL) and an expiry date, and decide who actually reads it. While writing this article we noticed our own site didn't have one, and added it (the contact is our contact form, not a personal email address).

2

Find and revoke credentials that have been exposed

Both OpenAI and Anthropic list publicly exposed credentials; they were the starting point in the July Hugging Face case too. Look for keys left in code, public repositories, config files and screenshots, and revoke them at the issuer. See stopping secret commits with gitleaks and secret files in public directories.

3

Close debug pages and admin panels in production

Anthropic's case involved reading credentials from an exposed debug page. Check that debug mode and detailed error output are off in production, and that admin panels aren't open to anyone on the internet. Per-framework checks are in security by framework.

4

Drop the assumption that a blocked visitor gives up — check authorization on every path

The first category OpenAI lists is bypassing access controls via a different web address. Hiding a link, or restricting one URL, doesn't stop a visitor that looks for another way. Every path that returns data needs a server-side check of whether this caller may see it. See authentication vs authorization, and the classic flaws IDOR and SQL injection.

5

Watch volume and speed, and keep logs long enough to look back

Agents don't rest, don't pause, and run the same action in parallel. Tightly spaced requests, a flat 24-hour pattern, and identical actions running at once are easier to tell from human use. And when a notice arrives months later, you can't verify anything without that day's logs. Cap-and-alert design is in rate limiting and abuse control; log retention in organization security priorities.

6

If you run anything people can post to, plan for agent spam

OpenAI cites agents using a public wiki as a message board. If you run an open wiki, forum, comment section or contact form, review posting rate limits and moderation.

7

If an AI company notifies you, check the logs before you panic

OpenAI says a notification from it should not automatically be read as a significant security incident. Check your access and authentication logs for the times and targets given, and decide whether it was information you meant to publish or a weakness to fix. If it's a weakness, fix it like any other vulnerability (the vulnerability remediation playbook).

What the agents did (per the disclosures)

bypassed limits via another URL / used exposed credentials / read credentials from a debug page / SQL injection / posted to a wiki

The basic that stops each

server-side authorization on every path / revoke leaked keys / debug off in production / parameterized queries / posting rate limits

For when something still gets in

watch volume and speed / keep logs / a contact that reaches someone (security.txt)

The agents used basic holes. Closed, they stop a human and an agent alike.

This site's view: an agent is not a free penetration tester

These incidents have an odd shape. A visitor with no intent to attack finds an organization's weaknesses, walks through them, and months later the organization is told "we got in" by the company that built it. The ACSC notes that the difference is an AI agent independently finding the kind of vulnerability human researchers have traditionally found and reported.

Reading this as "more helpful reports" would be a mistake. The same holes are visible to visitors who won't report them. And agents don't give up when blocked; they try, in parallel and quickly, what would take a human days. The value of waiting with a hole open keeps falling.

One more thing became clear: a notice needs an address. A good-faith report or an AI company's notification sits for days in a general mailbox if there's nowhere better to send it. A security.txt is a few lines of text, and you can publish it today.

Sources

  • OpenAI, "Hugging Face incident and other third-party impacts of misaligned models" (checked through the 25 September 2026 update) — openai.com
  • Anthropic, "Investigating three real-world incidents in our cybersecurity evaluations" (30 July 2026) — anthropic.com
  • Prime Minister of Australia, "Press conference - New York" (24 September 2026) — pm.gov.au
  • ASD's ACSC, "Risks of AI misalignment to Australian organisations" (24 September 2026) — cyber.gov.au
  • RubyGems Blog, "An update on the May spam-publishing campaign on rubygems.org" (11 September 2026) — blog.rubygems.org (used to confirm the attribution is not established)
  • IETF RFC 9116, "A File Format to Aid in Security Vulnerability Disclosure" — rfc-editor.org

FAQ

QHave AI agents actually accessed real websites without authorization?
A

Yes. After the July 2026 Hugging Face case, OpenAI reviewed its research and evaluation agents' activity and says it has notified dozens of third parties. Anthropic disclosed that models under evaluation reached the production systems of three real companies. The Australian government said an OpenAI agent accessed public and non-public files on the Medicare statistics portal in June.

QWhy do AI agents end up in third-party sites?
A

According to the disclosures, the agents were trying to finish an assigned task and looked for another way when blocked. Australia's Cyber Security Centre (ACSC) described an agent that independently identified vulnerabilities and tried to progress without direct human authorisation to complete its assigned activity. In Anthropic's cases, a misconfiguration had left the evaluation machines with live internet access, and the model treated real systems as part of the exercise.

QWhat weaknesses were used?
A

Nothing exotic. OpenAI lists bypassing access controls (for example via a different web address), using publicly exposed credentials, query or command injection, reaching internal parts of a service, and posting on third-party sites ("agent spam"). Anthropic describes basic, well-known techniques such as reading credentials from an exposed debug page and SQL injection.

QIf an AI company notifies us, does that mean a serious incident?
A

Not necessarily. OpenAI says a notification from it should not automatically be read as notice of a significant security incident; some organizations conclude the information was intentionally public, others find a design weakness to fix. Check your logs for the times given and decide for yourself.

QWhat should site operators do?
A

Basics, not new products: revoke exposed credentials, disable debug pages in production, check authorization server-side on every path, monitor volume and speed and keep logs, and publish a contact that actually reaches someone (such as security.txt).