Kimi K3 Launches as First Open 2.8‑Trillion‑Parameter Model with 1M‑Token Context
Kimi releases K3, a 2.8‑trillion‑parameter open model featuring a 1‑million‑token window and native vision, marking a new scaling milestone.

Kimi announced the launch of K3, its newest open‑source large language model. At 2.8 trillion parameters and a 1‑million‑token context window, K3 pushes the open‑model frontier beyond previous limits. The model also adds native vision capabilities and a sparsely‑gated Mixture‑of‑Experts layer that activates 16 of 896 experts. While its raw performance still trails the leading proprietary models, K3’s scaling efficiency is reported to be about 2.5 × higher than its predecessor K2. For developers, the immediate availability via Kimi.com, Kimi Work, and the API means a powerful tool is now ready for integration.
What happened
On July 16, 2026, Kimi released K3, a 2.8 trillion‑parameter model built on Kimi Delta Attention and Attention Residuals. The architecture introduces a sparsely‑gated Mixture‑of‑Experts system that selects 16 of 896 experts per token, and it scales training efficiency by roughly 2.5 × compared with K2.
K3 ships with a 1‑million‑token context window and native vision, allowing it to ingest screenshots, CAD drawings, or full codebases in a single pass. Early benchmark runs show frontier‑level performance on Kimi’s internal evaluation suite, consistently beating other open models and approaching the scores of Claude Fable 5 and GPT 5.6 Sol on selected tasks.
The full model weights are slated for public release on July 27, 2026, and Kimi plans to roll out low‑ and high‑effort inference modes after launch. Partnerships with inference providers are being finalized to ensure a reliable ecosystem rollout.
Why it matters
K3’s open‑source status lowers the barrier for developers to experiment with trillion‑parameter models without licensing fees, accelerating research and product innovation. The massive context window enables end‑to‑end processing of entire repositories or long documents, which can reshape workflows in code generation, documentation, and multimodal design. However, the model still lags the most powerful proprietary offerings, meaning critical applications may still prefer closed alternatives for peak performance.
- Open weights give unrestricted access for customization.
- 1 million token window supports whole‑codebase reasoning.
- Vision integration opens new multimodal use cases.
- Performance still trails top proprietary models on some benchmarks.
- Current inference defaults to max thinking effort, increasing compute cost.
- Full ecosystem (low/high‑effort modes) not yet available at launch.
How to think about it
Treat K3 as a platform rather than a drop‑in replacement for existing APIs. Start by prototyping tasks that benefit from long context—such as codebase navigation or document summarization—and evaluate cost versus accuracy. When performance ceilings are hit, consider hybrid pipelines that fall back to specialized proprietary models for the most demanding sub‑tasks. Keep an eye on upcoming low‑effort inference modes, which will reduce latency and cost for production workloads.
FAQ
How can developers access Kimi K3?+
What hardware is needed to run the 2.8‑trillion‑parameter model?+
How does K3 compare to GPT‑5 in code generation?+
- news·4 min readKimi K3 Launches with 1 Million‑Token Context Window and 300‑Agent Swarm, Escalating AI Competition
Moonshot AI's Kimi K3 offers a 1 M token context window, 300‑agent swarm and a 2‑3 T parameter MoE model, challenging Anthropic and OpenAI on cost and capability.
- ai·2 min readAlibaba’s Qwen 3.8 Model Launches with Updated Architecture and Open‑Source License
Alibaba released Qwen 3.8, the newest open‑source LLM, featuring architectural tweaks and broader language support for developers.
- ai·3 min readKimi K3 Matches Claude’s Performance at a Fraction of the Cost, Highlighting US AI Policy Gaps
Kimi K3 delivers Claude‑level coding output while costing a fraction, exposing pricing disparities and US policy shortcomings on AI model access.
The week’s highest-signal tech and AI stories, synthesized into a five-minute read. One email a week, no spam, unsubscribe anytime.