AI agent governance & security

OpenAI notified 100+ organizations about agent activity: what teams should change

•Make Better Editorial

OpenAI says it notified more than 100 organizations about AI-agent activity that met its notification criteria. That does not mean 100 breaches. Here is the practical governance lesson for teams deploying agents.

OpenAI says it has notified more than 100 organizations about AI-agent activity that met its notification criteria during a broader review of model behavior. The number is significant, but it is easy to misread: OpenAI explicitly says receiving a notification does not mean private information was accessed or that a third-party system was compromised. For teams deploying agents, the useful lesson is not the headline number. It is that an agent with tools, credentials and network access needs operational controls that ordinary chat workflows often do not.

Important distinction

More than 100 notifications does not equal more than 100 confirmed breaches. OpenAI's notification threshold is broader. Treat the figure as evidence of the scale of review and potential exposure, not as a verified breach count.

What OpenAI disclosed

OpenAI says that as of September 26 its teams had notified more than 100 organizations about activity that met its criteria. Reuters reports that the company is reviewing roughly 50 petabytes of data and that the work followed an earlier incident involving Hugging Face. OpenAI says that in some cases models used internet access in unintended ways or did not have ideal restrictions applied, and that it has been adding technical and operational measures to prevent or detect similar behavior earlier.

Why agents change the security model

Make Better analysis

A chatbot can produce a bad answer. An agent can turn a bad decision into an external action. Once a workflow can browse, execute tools, use credentials, edit records or send data, prompt quality is no longer the only control surface. Permissions, environment boundaries, approvals, logging and recovery become part of the product.

From chatbot risk to agent risk

Chat workflowAgent workflow
Primarily generates output for a personCan take actions in external systems
Human usually decides what happens nextSome decisions may execute automatically
A bad answer is often visible before actionA bad action may occur before review
Conversation history is the main audit trailTool calls, permissions and external changes also need logs
Recovery often means correcting the answerRecovery may require reverting external state

Five controls worth adding before an agent gets real access

  1. Start with least privilege. Give the agent only the tools, accounts and scopes required for the specific workflow.
  2. Separate read and write permissions. A research agent usually should not inherit the ability to edit, send or delete simply because those actions are available in the same system.
  3. Require human approval for high-impact actions such as sending external messages, publishing, deleting records, changing permissions or moving money.
  4. Log tool calls and important decisions with enough context to reconstruct what the agent attempted and what actually changed.
  5. Define a stop and recovery path: rate limits, spend limits, session termination, credential revocation and a clear way to revert changes when possible.

Use risk tiers instead of one approval rule

A simple agent action policy

Risk tierExampleDefault control
LowRead a public page or summarize an internal documentAutomatic with logging
MediumCreate a draft, update a reversible internal fieldAutomatic or sampled review
HighSend externally, publish, change permissions, delete or execute sensitive operationsExplicit approval before execution
ExceptionalActions with major financial, legal, security or irreversible impactKeep outside autonomous execution unless separately engineered and governed
Make Better analysis

The goal is not to put a human click in front of every tool call. That destroys much of the value of automation. The better pattern is to make autonomy proportional to consequence: broad freedom for low-impact reversible work and increasingly strict gates as actions become external, sensitive or difficult to undo.

What teams should test before production

  • What happens when the agent receives ambiguous or conflicting instructions?
  • Can untrusted webpage or document content influence a tool action?
  • Can the agent reach systems or data that are unrelated to its job?
  • Are secrets exposed in prompts, logs, browser sessions or tool outputs?
  • Does a failed tool call trigger repeated actions or uncontrolled retries?
  • Can operators identify and stop a problematic run quickly?
  • Can you reconstruct the sequence of actions after an incident?

What remains unknown

The public notification count should not be used to infer a breach rate or the probability that a typical production agent will cause an incident. The broader OpenAI review is still ongoing, and the public material does not provide a denominator that would allow a meaningful incident-rate calculation. Different agent architectures, permissions and environments can also have very different risk profiles.

Bottom line

The practical response is not to stop using agents. It is to stop treating an agent as a chatbot with extra buttons. When AI can act, teams need least-privilege access, risk-based approvals, observable tool use and a tested recovery path as part of the workflow design.

Sources & useful resources