Microsoft said on 7 October that it is bringing MAI Code 1.1 Flash, the coding model it introduced at Build, onto Windows PCs. The model has 137 billion total parameters and 6.8 billion active, and Microsoft says running it at 3-bit precision cuts its size by nearly 80% while keeping coding quality and a 256K-token context window.
The second piece is routing. GitHub’s HydraFusion already sends each task to a suitable cloud model. On Windows it will also be able to use models running on the device, starting in an experimental preview in the GitHub Copilot app, Copilot CLI and Visual Studio Code later in October. Windows ML is also gaining support for llama.cpp, the popular open-source runtime.
Microsoft’s pitch is cost. It says customers’ needs are outpacing their cloud budgets, and local models make tokens go further. Its Copilot assistant will get the same treatment on Copilot+ PCs in the coming months.
Two things are missing from the announcement. It gives no memory or hardware requirement for the local MAI Code model, and the claim that compression preserves coding quality comes with no published benchmark. Until both appear, it is hard to say which machines can run it or how much it loses against the cloud version.
