Kimi K3 Paused New Subscriptions Within 48 Hours of Launch

China's Moonshot AI has released a 2.8-trillion-parameter model that is so good, its own data centers could not keep up.

A readable brief on the AI scaling bottleneck

When a Chinese startup called Moonshot AI unveiled Kimi K3 in mid-July 2026, it was a headline grabber in the making. The model packs 2.8 trillion parameters into a mixture-of-experts design and offers a one-million-token context window. Moonshot says it performs alongside some of the world's best frontier models.

What to take away: Demand for an AI model can hit a wall not because the AI fails, but because the computers running it run out of power.

The trouble began almost immediately. Within roughly 48 hours of launch, the service was handling far more traffic than Moonshot had planned for. In a public statement, the company said its GPUs were operating close to the limits of current capacity. To protect the experience of existing subscribers, it temporarily paused new sign-ups and began reopening slots in batches as capacity came online.

A model too popular for its own infrastructure

Kimi K3 is priced aggressively at about $3 per million input tokens and $15 per million output tokens. It also promises an open-weight release under a modified open license later in the month, which has only added to the buzz. Researchers, developers and curious users all piled in at once.

The pause highlights a recurring reality of the AI era. Building a brilliant model is only half the battle. Serving it to millions of simultaneous users requires enormous compute, and the most powerful GPUs remain a scarce and expensive resource. A startup with no Google- or Meta-scale data center footprint is especially vulnerable to a sudden spike in interest.

Why this matters

For the wider industry, the Kimi K3 moment is a reminder that capability and capacity are different things. A model can rank at the frontier and still be unreachable by new users the moment the queue fills. Whether Moonshot can scale its serving fleet fast enough will be a close watch — not just for its own future, but as a case study for any new lab trying to compete at the top of the AI stack.

Artificial IntelligenceTechnologyChinaGPUs