Research acceleration: The view inside OpenAI
Inside OpenAI, coding agents are reshaping AI research. Explore early data on agent usage, experiment velocity, task complexity, and research acceleration.
Informal updates from the OpenAI team
Inside OpenAI, coding agents are reshaping AI research. Explore early data on agent usage, experiment velocity, task complexity, and research acceleration.
Jakub Pachocki reflects on increasingly capable AI and the challenge of keeping it aligned. He calls for stronger safeguards and international coordination.
Findings from the incident and the response across model security, monitoring, and alignment.
Testing whether behavior changes when a model believes its grader wants something different.
Testing how alignment improvements hold across domains and under adversarial pressure.
Comparing public evaluation signals with evidence from deployment.
An investigation of affected reward pathways and possible effects on monitorability.
Using a separate agent to review proposed actions that cross a boundary.
Datasets, code, and an evaluation filtering strategy for studying monitorability.
A program supporting independent alignment and safety research.
Experiments spanning midtraining, posttraining, and generalization.
Evaluating how well models follow the OpenAI Model Spec.
Training agents to report covert misbehavior through a dedicated tool.
Monitoring model behavior in real internal workflows.
Metagaming can complicate how we interpret behavior, and current models still give us a chance to study it directly.
ARGO distills black-box reward models into interpretable rubrics using reinforcement learning.
Why a limitation of frontier models is reassuring for AI safety.
We’re committing $7.5M to The Alignment Project to fund independent research developing mitigations to safety and security risks from misaligned AI.
Reasoning models can find and understand unknown misaligned behaviors from how users respond.
An experimental dataset of crowd-written rubrics that surfaces why people prefer one model output over another.
Deeper analysis of confession training and comparisons to chain-of-thought monitoring.
Emergent misalignment not only activates misaligned personas, but also suppresses helpful assistant personas.
We introduce evaluations for chain-of-thought monitorability and study how it scales with test-time compute, reinforcement learning, and pretraining.
A pipeline to uncover unknown misaligned behavior and scale the creation of realistic evaluations.
Efficiently finding features that cause behaviors.
We train and deploy an AI review agent optimised for precision and real-world use, enabling oversight to scale with autonomous code generation.
Introducing our blog on alignment research.