OpenAI

OpenAI Opens Bug Bounty for Prompt Injection and Agent Abuse

OpenAI's new Safety Bug Bounty, run on Bugcrowd, pays up to $7,500 for reproducible AI abuse risks: prompt injection, agentic misuse, and data exfiltration via connectors.

OpenAI Opens Bug Bounty for Prompt Injection and Agent Abuse — article cover
On this page6 SECTIONS
  1. Scope: Paying for Risks That Are Not Security Bugs
  2. Rewards and Submission Requirements
  3. Why Agentic Products Are the Main Target
  4. Splitting Duties With the Existing Security Program
  5. What It Means for Developers and Security Teams
  6. Sources

OpenAI has launched a Safety Bug Bounty program, announced March 26 and picked up by SecurityWeek and other outlets the following day. This is not another conventional security bounty. Hosted on Bugcrowd, it pays for AI abuse and safety risks that would not qualify as security vulnerabilities: third-party prompt injection, data exfiltration attacks, and agentic products performing disallowed actions on a user’s behalf.

For developers, this is the first time a frontier lab has turned “behavioral AI risk” into a formal reporting pipeline with triage and pricing. As agent products deploy at scale, the attack surface has shifted from memory corruption to “the model did something it should not have” — and this program effectively drafts the industry’s taxonomy for that class of problem.

Scope: Paying for Risks That Are Not Security Bugs

Per the announcement, the program covers several categories that traditional vulnerability programs would typically reject:

  • Third-party prompt injection and data exfiltration attacks
  • Disallowed actions performed at scale by agentic OpenAI products, alongside other harmful product behaviors
  • Exposure of OpenAI’s proprietary information and weaknesses in account and platform integrity
  • Design or implementation issues that could cause material harm, including bypasses of abuse protections
  • Vulnerabilities in connectors and MCP integrators that are abusable to cause material harm

The named targets are products that act or access data on a user’s behalf: Atlas Browser, Codex, Operator, Connectors, and other ChatGPT tools.

Rewards and Submission Requirements

The ceiling is $7,500 for high-severity issues that are consistently reproducible, provided the report includes a clear set of recommended steps or mitigations. OpenAI states that final reward decisions and amounts are discretionary. Flaws that facilitate direct paths to user harm may qualify case by case when paired with actionable, discrete remediation steps.

In other words, the currency is not finding a problem — it is writing the problem up as a reproducible, fixable engineering artifact. A vague “the model behaved badly” or a one-off anomaly nobody can recreate will not clear the bar. By mainstream bounty standards the payout is modest, but the point is less the money than the classification: OpenAI is signaling that behavioral failures in deployed agents are now a findable, priceable, fixable class of defect.

Why Agentic Products Are the Main Target

Read the named-product list and the focus is obvious: risk concentrates where models operate a computer for you. Operator browses the web, Codex touches repositories, Connectors reach third-party services through MCP. In those settings, an attacker does not need to exploit any software flaw — malicious content only has to enter the context as text to steer the agent toward an attacker-chosen action.

That is also why MCP integrators are explicitly in scope: one vulnerable connector extends the attack surface for every ChatGPT user who plugs it in. By drawing the boundary around third-party integrations, OpenAI is telling the agent ecosystem that implementation quality on the integration side is platform risk, not just your problem. Anyone shipping an MCP server or connector that ChatGPT users can reach is effectively auditioning for a role in OpenAI’s security perimeter.

Splitting Duties With the Existing Security Program

OpenAI’s long-running bug bounty handles traditional software vulnerabilities; the new program follows the same rules with several additions. Submissions are triaged jointly by the Safety and Security Bug Bounty teams and can be rerouted between the two — a report that spans both security and abuse will not be bounced for being out of scope.

The program also sits alongside the separate Bio Bug Bounty, giving OpenAI three channels mapped to three risk classes: classic vulnerabilities, product abuse, and biosecurity.

What It Means for Developers and Security Teams

Three practical effects. First, teams building agents or MCP services get a formal reporting channel plus an official working definition of “material harm” — a usable starting point for threat modeling. Second, the reward level is conservative, and community critics have already questioned whether payouts of this size will attract serious researchers; whether the program becomes a real defensive layer depends on its actual payout record. Third, by making reproducibility-plus-mitigation a hard requirement, OpenAI pushes AI safety reporting toward engineering discipline: policy complaints are not priced, reproducible attack paths with fixes are.

Sources

AI-assisted summary compiled from the sources above, reviewed by a human before publishing.

FOUND_THIS_USEFUL?

Support more practical AI articles, tutorials, and build notes.

BUY_ME_A_COFFEE
SHAREXEMAIL