Market Review (2026-07-24)
Kimi K3 - A Milestone for China's Open‑Source Large language Models
On July 17, 2026, Moonshot AI released Kimi K3, a next‑generation large model that raises the bar for open‑source performance. With full weights set for release on July 27, Kimi K3 is believed to be narrowing the gap between Chinese open‑source models and the global rivals such as Anthropic and OpenAI.
Kimi K3 is a MoE (Mixture of Experts) model with 2.8tn total parameters, it surpasses the previous record of 1.6tn set by DeepSeek V4, cementing its position as the largest publicly available open model in terms of parameter scale. Featuring a 1mn‑token context window, the model is primarily designed for long‑horizon programming, knowledge work, and complex reasoning tasks. Underpinning this massive scale are three core architectural innovations:
1). KDA (Kimi Delta Attention) hybrid linear attention mechanism. Conventional attention mechanisms suffer from quadratic computational scaling with text length: doubling the input quadruples the cost. KDA addresses this by deploying linear attention in three out of every four network layers (enabling linear scaling of computation with text length) while retaining global attention in the remaining layer (to preserve macro-level contextual understanding). This hybrid design reduces KV cache, the memory resources consumed by temporarily storing historical information during inference, by up to 75% and achieves up to 6.3 times higher decoding throughput in million-token long-context scenarios.
2). Attention Residuals: In deep neural networks, early-layer information tends to degrade or vanish as it passes through successive layers—much like a message distorted after being relayed through many people. Attention Residuals enable the model to selectively retrieve historical information from various depths across layers, rather than relying solely on fixed-path layer-by-layer propagation. This effectively establishes a flexible cross-layer information retrieval mechanism for the giant model, allowing it to access early-stage raw information at any depth and thereby preserving the integrity and accuracy of low-level information in deeper layers.
3). Stable LatentMoE sparse architecture: The core premise of MoE (Mixture of Experts) is to route different tasks to specialized "expert" subnetworks, rather than having a single monolithic network handle all tasks, much like a specialist team that deploys only the most relevant experts per task instead of the whole team. Kimi K3 scales the expert count to 896, yet activates only 16 experts per inference, with an activation ratio of roughly 1.8%. This means that despite its massive total parameter count (2.8 trillion), each computation engages only a fraction of the model's capacity, effectively scaling model capacity while keeping per-inference computational costs within a manageable range.
Kimi K3 ranks among the top performers across multiple authoritative third‑party benchmarks. It scores 57 points on the Artificial Analysis Intelligence Index, placing third globally behind Anthropic's Fable and OpenAI's GPT‑5.6 Sol, while leading all open‑source models with a 13 points improvement over its predecessor K2.6. In programming, it achieves a historic breakthrough by topping the Frontend Code Arena at 1,679 points, surpassing both Claude Fable 5 and GPT‑5.6 Sol, and securing six out of seven sub‑category firsts. On the Text Arena comprehensive leaderboard, it ranks ninth globally with 1,486 points, making it the first open‑source model to enter the top ten. In long‑horizon tasks, it places third on SWE Marathon and first on BrowseComp, while taking the top spot on AutomationBench‑AA in agentic evaluations.
K3 occupies the highest tier among domestic Chinese models: It is charging RMB20/mn tokens for input (cache miss), RMB2/mn tokens for input (cache hit), and RMB100/mn tokens for output, at least three times the cost of K2.6 and significantly above domestic peers such as DeepSeek V4‑Pro, Zhipu GLM‑5.2, and Qwen 3.7‑Max, but still below GPT‑5.6 Sol at US$30/mn tokens for output (50% discount) and Claude Fable 5 at US$50/mn tokens for output (70% discount). Nevertheless, the apparent cost advantage by K3 does not translate into equally attractive savings in real‑world usage due to its lower token efficiency and slower inference speed. According to Artificial Analysis, its average cost per task is approximately US$0.95, only slightly 10% off under GPT‑5.6 Sol's US$1.04 but still offering a reasonable cost‑performance advantage with 65% off over Claude Fable 5 at US$2.75.
Our views: The rapid iteration of Kimi K3 reflects an unprecedented narrowing of the gap between open‑source and closed‑source models. Despite facing shortages of advanced chips, continuous architectural innovation has shortened the gap between Chinese open‑source models and their U.S. counterparts from 6‑9 months to 2‑3 months. As industry‑wide iteration accelerates and competition intensifies, the technological barrier built solely on model parameter scale is eroding, and the model of sustaining short‑term technical leadership through heavy capex is increasingly being questioned by the market. We believe that as LLM capabilities gradually converge, competition will shift toward the infrastructure layer, where China holds certain advantages over the United States in electricity costs and AI R&D talent. However, these cost advantages may face policy headwinds; for instance, Chinese models are no longer accessible in the United States.
Demand for Kimi K3 grew exponentially within 48 hours of its release, quickly approaching the limits of its available computing capacity and prompting Moonshot AI to suspend new consumer subscriptions. We believe the release of Kimi K3 will directly benefit mid‑stream cloud service providers, and we are optimistic about Alibaba Cloud (9988HK, HK$114.90, HK$2.2tn), given its integrated full‑stack capabilities spanning from self‑developed chips to cloud infrastructure and AI applications. Moreover, Alibaba also holds roughly 36% of Moonshot AI's equity. With Moonshot AI currently valued at US$31.5bn, the company is preparing for a Hong Kong IPO, unlocking significant value for Alibaba as a major shareholder. The counter is trading at 18x FY27E P/E. (Research Department)