MLPerf Inference v6.1 adds RAG and agentic tests as Rubin arrives in preview

MLCommons published results for MLPerf Inference v6.1 on 16 September, adding two benchmarks that reflect how AI systems are actually deployed. An end-to-end retrieval-augmented generation test runs a full pipeline — embedding model, retriever, re-ranker, then language models reasoning over what comes back. An edge agentic inference test covers multi-turn work such as coding, where each step depends on the ones before it.

The round drew submissions from 30 organisations, a record, including six first-time submitters. More than half used a new API-centric harness that will underpin MLPerf Endpoints, the suite due to replace Inference for data centres.

MLCommons reports that the best per-accelerator server result on the DeepSeek-R1 test improved 5.7 times against v5.1 a year ago, and 2.99 times on the vision-language test against v6.0 six months ago. Nvidia’s own write-up of the same round cites up to 3.7 times the throughput of GB300 NVL72 on Qwen3-VL — a rack-to-rack comparison, not a per-accelerator one. Both can be true; they are not the same measurement.

Note also that Nvidia’s Rubin and Vera Rubin NVL72 entries sit in the preview category — hardware not yet generally available. AMD’s Ryzen AI Max+ 395 and Instinct MI350P, and Intel’s Arc Pro B70, are listed as available.

Source: MLCommons.


Related