MLCommons published results for MLPerf Inference v6.1 on 16 September, adding two benchmarks that reflect how AI systems are actually deployed. An end-to-end retrieval-augmented generation test runs a full pipeline — embedding model, retriever, re-ranker, then language models reasoning over what comes back. An edge agentic inference test covers multi-turn work such as coding, where each step depends on the ones before it.
The round drew submissions from 30 organisations, a record, including six first-time submitters. More than half used a new API-centric harness that will underpin MLPerf Endpoints, the suite due to replace Inference for data centres.
MLCommons reports that the best per-accelerator server result on the DeepSeek-R1 test improved 5.7 times against v5.1 a year ago, and 2.99 times on the vision-language test against v6.0 six months ago. Nvidia’s own write-up of the same round cites up to 3.7 times the throughput of GB300 NVL72 on Qwen3-VL — a rack-to-rack comparison, not a per-accelerator one. Both can be true; they are not the same measurement.
Note also that Nvidia’s Rubin and Vera Rubin NVL72 entries sit in the preview category — hardware not yet generally available. AMD’s Ryzen AI Max+ 395 and Instinct MI350P, and Intel’s Arc Pro B70, are listed as available.
Source: MLCommons.
