ParallelKernelBench tests whether LLMs can write fast multi-GPU CUDA kernels across 87 real workloads. The best model solves under a third, but a few generated kernels beat any public implementation.
The same model behaves differently across provider endpoints. Infrastructure, quantization, load handling, and routing defaults all change the result. Here's how to measure latency, throughput, uptime, and precision, then turn the measurements into a routing policy.
Generation runs through the dedicated /api/v1/images endpoint and understanding through /chat/completions, with the same key and billing. Here's the full contract for both jobs, plus the fix for the 'no endpoints found' error.
We generated 12 landing pages with Kimi K2.7 Code and Claude Fable 5. Kimi cost 94% less and scored within a few points on every page. Here's what actually moved the needle.
Define a taxonomy and a small model tags every generation in your workspace by department, task type, or agent complexity. Filter your logs and group your Activity analytics by the results.
Health in ChatGPT now lets eligible U.S. users securely connect medical records and Apple Health to get more personalized insights and better understand their health.
OpenAI announces Project Camellia in Effingham County, Georgia, with commitments to responsible energy, community investment, jobs, and access to Codex.