Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?
Executive Summary
Anthropic and OpenAI want to embed independent safety evaluators inside their AI labs. Researchers welcome the unprecedented access, but warn meaningful oversight requires transparency, independence, and eventually regulation.
You May Also Like
Next Logical Step
Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexp...