A benchmark released through Alibaba’s Qwen GitHub organisation tests a narrower question than whether a model can write a GPU kernel. D2K-Bench, described in a paper posted to arXiv on 2 October, gives agents 26 tasks and 85 workloads and measures how well they turn expert design guidance into fast, correct kernels.
The guidance comes in three layers: high-level algorithmic insight, dataflow design, and low-level optimisation tricks. Each task runs twice, with and without it, holding the task description, workloads, tools, hardware and a 350-turn budget constant, so the comparison is paired rather than inferred.
On Nvidia B200 GPUs across five models, the guidance lifted correctness over 130 model-task pairs from 93.1% to 98.5%, and the paper’s performance score across all 26 tasks from 1.46 to 1.95. For the three models that submitted correct code on every task in both runs, which the paper names as GPT-6-Astra, Claude Opus 4.8 and GPT-5.6-Sol, geometric mean speedup rose from 1.69 times to 2.49 times.
Read that as a result about where the gap sits. Handed an expert’s plan, agents produce substantially faster code, which suggests the weaker link is working out what to do rather than writing it. The paper also flags design properties that go unimplemented even when spelled out. All figures are the authors’ own.
Source: the D2K-Bench paper.
