Loading the catalog…
Loading the catalog…
We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
Toward understanding and preventing misalignment generalization. We study how training on incorrect responses can cause broader misalignment in language models and identify an internal feature driving this behavior—one that can be reversed with minimal fine-tuning.
Open sourceThe catalog shows persisted RADAR opportunities. Storage availability does not mean sources are verified or offers are active.