Loading the catalog…
Loading the catalog…
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
Open sourceThe catalog shows persisted RADAR opportunities. Storage availability does not mean sources are verified or offers are active.