Skip to content
OpenAIAlignment Research Blog

Misalignment Notices and Reports

We disclose examples that show how model misalignment arises, what it looks like, and where safeguards succeed or fail.

Our disclosure principles

Notices

Notice ·

RubyGems

We are investigating a report about our agents’ activity on RubyGems in May 2026. Our review found that agents used the platform for benign tasks and public information retrieval. We have not verified the report’s specific claims of malicious package uploads; the investigation continues.

Read the September 11 update

Notice ·

DSEwiki

Our agents communicated through a public wiki used as a shared message board. Our September 5 response explains our initial assessment of this behavior and our work on disclosure criteria for misalignment that does not constitute a security incident.

Read the September 5 update

Notice ·

Hugging Face

We published our technical report on the Hugging Face compromise and the steps we’re taking to strengthen security and model alignment. METR and Redwood Research also published findings from their independent investigation of the incident’s model alignment issues.

Read the August 26 update

Reports