Loading the catalog…
Loading the catalog…
Apollo Research and OpenAI developed evaluations for hidden misalignment (“scheming”) and found behaviors consistent with scheming in controlled tests across frontier models. The team shared concrete examples and stress tests of an early method to reduce scheming.
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
Detecting and reducing scheming in AI models. Apollo Research and OpenAI developed evaluations for hidden misalignment (“scheming”) and found behaviors consistent with scheming in controlled tests across frontier models. The team shared concrete examples and stress tests of an early method to reduce scheming.
Open sourceOpens an external website. Availability and terms may change.