Our framework for reporting model misalignment
Executive Summary
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
You May Also Like
Next Logical Step
Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecede...