Moonshot AI has released Kimi K3, a 2.8 trillion parameter model that can be deployed on AWS using Amazon SageMaker HyperPod and Amazon EKS. The model, which was announced on July 27, 2026, is designed for complex tasks such as long-horizon coding and multi-step reasoning. It is available on Hugging Face under the identifier moonshotai/Kimi-K3 and uses the MXFP4 format for efficient inference. The model’s architecture, which includes a Stable LatentMoE framework, allows for efficient scaling and performance improvements over its predecessor, Kimi K2. The deployment process involves setting up a HyperPod cluster with EKS orchestration and procuring reserved GPU capacity through Flexible Training Plans or Capacity Blocks. This ensures that the model can be hosted on AWS infrastructure without contention for resources. The model’s deployment requires a vLLM day-0 inference container and is supported by the Inference Operator, which simplifies the orchestration and management of the model. The deployment process is detailed in the AWS blog and includes steps for creating a HyperPod cluster, setting up the required infrastructure, and applying a configuration manifest to deploy the model. The model’s availability on AWS provides organizations with the ability to self-host one of the most capable models in existence on their own infrastructure. The deployment on EKS involves setting up a standalone cluster and reserving GPU capacity through EC2 Capacity Blocks. The AI on EKS project provides a recipe for automating the end-to-end provisioning of the infrastructure. The model’s deployment is supported by the vLLM framework, which provides native support for MoE architectures and the MXFP4 quantization format. The deployment process ensures that the model can be hosted on AWS infrastructure with the necessary resources and optimization for large-scale inference deployments. The model’s availability on AWS provides organizations with the ability to self-host one of the most capable models in existence on their own infrastructure. The deployment on EKS involves setting up a standalone cluster and reserving GPU capacity through EC2 Capacity Blocks. The AI on EKS project provides a recipe for automating the end-to-end provisioning of the infrastructure. The model’s deployment is supported by the vLLM framework, which provides native support for MoE architectures and the MXFP4 quantization format. The deployment process ensures that the model can be hosted on AWS infrastructure with the necessary resources and optimization for large-scale inference deployments.

Kimi K3 is built on a differentiated architecture featuring Kimi Delta Attention (KDA), Gated Multi Head Latent Attention (MLA), and a Stable LatentMoE framework. The model distributes its 2.8 trillion parameters across 896 specialist experts, activating only 16 per token. This means approximately 104 billion parameters are active during any single forward pass, yielding a 2.5x improvement in scaling efficiency over its predecessor, Kimi K2. The model supports native tool calling, structured output, and an always-on thinking mode for multi-step problem solving. The open weights for Kimi K3 are available on Hugging Face under the model identifier moonshotai/Kimi-K3. The weights are distributed in MXFP4 format, which provides an effective balance between model quality and memory efficiency for large-scale inference deployments. The model requires a vLLM day-0 inference container for Kimi K3, which is available in vllm/vllm-openai:kimi-k3. vLLM provides native support for MoE architectures, tensor parallelism, and the MXFP4 quantization format, making it the recommended serving engine for this model. Deploying Kimi K3 requires a p6-b300 instance (ml.p6-b300.48xlarge), which provides 8 NVIDIA B300 Blackwell Ultra GPUs with high-bandwidth interconnects necessary for efficient tensor-parallel inference across the full expert pool. AWS offers two primary mechanisms to procure this capacity: Flexible Training Plans for SageMaker HyperPod and Capacity Blocks for EC2 GPU instances. The model can be deployed on AWS using either method, with the Inference Operator handling model download, container scheduling, health checks, and endpoint readiness. Once the endpoint transitions to a ready state, it exposes an OpenAI compatible API at the configured invocation path. The deployment process is detailed in the AWS blog and includes steps for creating a HyperPod cluster, setting up the required infrastructure, and applying a configuration manifest to deploy the model. The model’s availability on AWS provides organizations with the ability to self-host one of the most capable models in existence on their own infrastructure. The deployment on EKS involves setting up a standalone cluster and reserving GPU capacity through EC2 Capacity Blocks. The AI on EKS project provides a recipe for automating the end-to-end provisioning of the infrastructure. The model’s deployment is supported by the vLLM framework, which provides native support for MoE architectures and the MXFP4 quantization format. The deployment process ensures that the model can be hosted on AWS infrastructure with the necessary resources and optimization for large-scale inference deployments.

Kimi K3 was released on July 27, 2026, as the first open-weight system to reach the 3 trillion parameter class. The model is designed to deliver frontier-level intelligence while making its weights publicly available, allowing organizations to self-host one of the most capable models in existence on their own infrastructure. It excels at long-horizon coding tasks, agentic workflows, and complex reasoning. The model supports native tool calling, structured output, and an always-on thinking mode for multi-step problem solving. The open weights for Kimi K3 are available on Hugging Face under the model identifier moonshotai/Kimi-K3. The weights are distributed in MXFP4 format, which provides an effective balance between model quality and memory efficiency for large-scale inference deployments. The model requires a vLLM day-0 inference container for Kimi K3, which is available in vllm/vllm-openai:kimi-k3. vLLM provides native support for MoE architectures, tensor parallelism, and the MXFP4 quantization format, making it the recommended serving engine for this model. Deploying Kimi K3 requires a p6-b300 instance (ml.p6-b300.48xlarge), which provides 8 NVIDIA B300 Blackwell Ultra GPUs with high-bandwidth interconnects necessary for efficient tensor-parallel inference across the full expert pool. AWS offers two primary mechanisms to procure this capacity: Flexible Training Plans for SageMaker HyperPod and Capacity Blocks for EC2 GPU instances. The model can be deployed on AWS using either method, with the Inference Operator handling model download, container scheduling, health checks, and endpoint readiness. Once the endpoint transitions to a ready state, it exposes an OpenAI compatible API at the configured invocation path. The deployment process is detailed in the AWS blog and includes steps for creating a HyperPod cluster, setting up the required infrastructure, and applying a configuration manifest to deploy the model. The model’s availability on AWS provides organizations with the ability to self-host one of the most capable models in existence on their own infrastructure. The deployment on EKS involves setting up a standalone cluster and reserving GPU capacity through EC2 Capacity Blocks. The AI on EKS project provides a recipe for automating the end-to-end provisioning of the infrastructure. The model’s deployment is supported by the vLLM framework, which provides native support for MoE architectures and the MXFP4 quantization format. The deployment process ensures that the model can be hosted on AWS infrastructure with the necessary resources and optimization for large-scale inference deployments.

Source: awsml