Our framework for reporting model misalignment
OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
At a glance
- openai.com: Our framework for reporting model misalignment
- techmeme.com: OpenAI discloses six new AI safety incidents since October, including models concealing mistakes, and announces a new framework for reporting model misalignment (Axios)
- wired.com: OpenAI Creates a New Framework to Disclose Bad AI Behavior
The story
openai.com: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
techmeme.com: Axios : OpenAI discloses six new AI safety incidents since October, including models concealing mistakes, and announces a new framework for reporting model misalignment OpenAI on Wednesday disclosed six new incidents in which its models concealed mistakes, sought unauthorized credentials
wired.com: The company also disclosed previously unreported incidents in which its AI models behaved in misaligned ways, including uploading files to the internet without being asked.