Improving Frontier AI Incident Reporting Regimes
In recent months, AI agents have broken out of testing environments to infiltrate third-party companies, attempted to insert malicious code into an open-source project, and hacked into at least one major technology company. Some of these incidents, including when OpenAI agents breached Hugging Face, led to limited, voluntary investigations that identified concrete lessons about what must change to build safer, more aligned AI systems, though their limited scope left important questions unanswered.
Thorough incident investigations are essential to understand why and how AI systems misbehave, but current AI incident reporting regimes do not reliably ensure that such incidents are reported, independently investigated, or used to improve safety across the industry. Indeed, the public has only become aware of recent incidents through independent research and public disclosures by AI firms, companies breached by rogue AI agents, and government evaluators. Given that a number of jurisdictions already have incident reporting requirements, this suggests there are important gaps in current reporting regimes, both for facilitating the discovery of these incidents and learning from them.
The Hugging Face incident illustrates these weaknesses in current incident reporting regimes. In California, officials said the incident did not meet the state’s mandatory reporting threshold. And while OpenAI notified EU authorities about the incident, the EU reporting regime does not require an independent external investigation of incidents or that developers routinely share the lessons learned across industry. Future incidents could therefore reveal important safety failures without triggering requirements to report them or share what investigators learn with other developers. Creating reporting regimes that produce reliable and timely information about how AI systems fail – and can be improved – is therefore an urgent task.
This policy brief describes ways in which current incident reporting requirements – at the US state level and in the EU – fall short and makes recommendations to policymakers on how they could be updated to ensure that future similar incidents are identified and can be learned from. This brief recommends:
- Expanding mandatory reporting regimes to include near misses.
- Including incidents caused by models during training, evaluation, and internal use.
- Requiring independent investigations of sufficiently serious incidents and near misses.
- Requiring more reliable agent attribution.
- Specifying what data frontier AI developers must preserve.
- Incentivizing candid, early reporting by piloting safe harbor arrangements.
- Requiring developers make plans to fix failures, following up to make sure changes are implemented, and ensuring lessons are shared across industry.



