Moonshot AI halts Kimi K3 subscriptions as compute capacity reaches limit

Moonshot AI has suspended new subscriptions for its Kimi K3 model following a massive surge in user demand within a 48-hour period. The Chinese startup reported that the sudden influx of traffic overwhelmed its existing compute infrastructure. This temporary pause allows the team to stabilise performance for current users while they work on scaling their hardware resources.

This incident highlights the critical importance of elastic scaling and robust infrastructure planning for enterprise-grade AI deployments. Teams must anticipate volatile demand curves and ensure their backend can handle rapid growth to prevent service outages that disrupt production workflows.

  • Kimi K3 subscriptions paused after a 48-hour traffic spike.
  • Infrastructure capacity was unable to maintain performance under the sudden load.
  • Moonshot AI is prioritising service stability for existing enterprise and individual users.
Sovereign AI AI Apps & Platforms Generative AI
All AI news

More AI news

Models

OpenAI disbands preparedness team responsible for assessing catastrophic ai risks

OpenAI has reportedly dissolved its preparedness team, the group tasked with evaluating and mitigating potential catastrophic risks from advanced AI models. This internal restructuring follows several high profile departures from the company safety and alignment divisions.

Models

Deepseek raises V4 API pricing by up to eleven times ahead of reported IPO

DeepSeek has implemented a substantial price increase for its V4 API, with some costs rising by up to eleven times as of 16 August 2026. This shift signals the end of the low cost strategy the company utilised to disrupt the market in 2025. The adjustment coincides with industry reports suggesting the firm is preparing for an initial public offering.

Models

Writer releases enterprise-optimised GLM-5.2 variant with token-saving technology

Writer has launched a new variant of the open-source GLM-5.2 model, specifically post-trained to meet enterprise reliability standards. This release introduces a specialised harness designed to significantly reduce token consumption, directly addressing the high operational costs associated with large-scale model deployment.

Ready to build something
extraordinary?

15 minutes. No pitch deck. Just a conversation about what AI can do for your team.

Talk directly with our AI specialists

15 min, no strings
No sales pressure
Prototype in 7 days