OpenAI’s new confession system teaches models to be honest about bad behaviors
OpenAI announced today that it is working on a framework that will train artificial intelligence models to acknowledge when they've engaged in undesirable behavior, an approach the team calls a confession. Since large language models are often trained to produce the response that seems to be desired, they can become…
bubmagDecember 3, 2025