Loading the catalog…
Loading the catalog…
We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
PaperBench: Evaluating AI’s Ability to Replicate AI Research. We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research.
Open sourceThe catalog shows persisted RADAR opportunities. Storage availability does not mean sources are verified or offers are active.