Loading the catalog…
Loading the catalog…
Together.ai
ReasonIF finds frontier LRMs fail to follow reasoning instructions >75% of the time; introduces a benchmark across languages, formatting, and length.
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
Large Reasoning Models Fail to Follow Instructions During Reasoning: A Benchmark Study. ReasonIF finds frontier LRMs fail to follow reasoning instructions >75% of the time; introduces a benchmark across languages, formatting, and length.
Open sourceOpens an external website. Availability and terms may change.