AWS made Kimi K3, Moonshot AI’s open-weight flagship, available on Amazon Bedrock on 18 September, through United States and global cross-region inference profiles at standard Bedrock pricing. It is the first open-weight model on the service to support explicit prompt caching, which matters for long-running work that keeps re-reading the same repository or document set.
The model itself is not new. Moonshot published it in July and the weights are on Hugging Face, so the news here is distribution rather than capability: enterprise buyers who will not self-host a model of this size can now reach one under the same procurement, data handling and billing terms as the closed models beside it. AWS states zero data retention and no operator access during inference.
The specifications are Moonshot’s own and AWS repeats rather than tests them: 2.8 trillion total parameters with 16 of 896 experts active per token, a one-million-token context window, and native vision. Moonshot claims roughly 2.5 times better overall scaling efficiency than Kimi K2. It is also candid that K3 trails the leading proprietary models while remaining competitive on published benchmarks.
Sources: AWS and Moonshot AI.
