NVIDIA released v0.3.1 of Cosmos-H-Surgical on 18 or 19 September — the repository dates its commits by relative age rather than by calendar date — adding a distilled four-step DMD2 checkpoint alongside the existing base model. The model generates surgical video: given a starting image and a structured text description it predicts the next 92 frames, and it can turn simulated control video into photorealistic footage.
The new checkpoint is about cost rather than capability. NVIDIA reports it running 12.84 times faster than the 50-step base model on an H100. Output is 480p, validated at 832 by 480, at 16 frames per second.
The underlying model is a 15.17-billion-parameter diffusion transformer. Its training set combines roughly 12,600 synthetic laparoscopic cholecystectomy videos with 15,043 real robot-assisted prostatectomy videos from the GraSP dataset. Transfer FVD scores, measured over 1,010 videos per modality, range from 34.7 to 41.1. Weights and source code are under the OpenMDW-1.1 licence.
What makes this worth noting is the scoping. The model card states the models “are not intended for clinical diagnosis or autonomous clinical decision-making”, and the stated purpose is generating training and simulation data for surgical robotics. That is a narrower and more defensible claim than medical AI announcements usually carry, and the numbers above are NVIDIA’s own.
