Loading the catalog…
Loading the catalog…
A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
Introducing SimpleQA. A factuality benchmark called SimpleQA that measures the ability for language models to answer short, fact-seeking questions.
Open sourceThe catalog shows persisted RADAR opportunities. Storage availability does not mean sources are verified or offers are active.