Loading the catalog…
Loading the catalog…
Together.ai
ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.
What RADAR observed and classified to build this opportunity. It is what the source published, not a verification that the offer is still active.
ParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet). ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.
Open sourceParallelKernelBench: Frontier LLMs can't write fast multi-GPU kernels (yet). ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.