Loading the catalog…
Loading the catalog…
A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
Separating signal from noise in coding evaluations. A new analysis from OpenAI reveals issues in SWE-Bench Pro, a popular coding benchmark, raising concerns about reliability and accuracy in evaluating AI models.
Open sourceThe catalog shows persisted RADAR opportunities. Storage availability does not mean sources are verified or offers are active.