Kimi says a 48-hour demand surge pushed its GPU capacity close to the limit, forcing a pause in new subscriptions as a separate Kimi Code plan prepares to change how coding access is bundled and rationed.
Kimi has stopped taking new subscribers and reserved compute for existing members after what it describes as a sharp K3 demand surge. The decision establishes that the company is rationing access. It does not establish the size, likely duration or cause of the shortfall beyond Kimi's own account.
Kimi said in its announcement that demand over the preceding 48 hours had pushed close to the limits of its current capacity. It temporarily paused new subscriptions, prioritized compute for current members and said existing subscribers would not be affected. A flash report relayed the same company statement.
The 48-hour figure describes the period in which Kimi says demand rose; it is not a growth rate or a measure of utilization. The company gave no subscriber count, request volume, accelerator inventory, service-level data or estimate of unmet demand. Its plan to add capacity and reopen places in batches also came without a timetable or target.
That leaves two distinct facts. Kimi has made an operational choice to protect current paying users, and it says GPU pressure forced that choice. The archived sources provide no independent measurement of the second claim.
Kimi is also separating general membership for its web, app and Work products from a coding plan. Its membership documentation says the new structure begins July 20, 2026.
Under the legacy plan, Kimi Code shared a monthly quota with features such as presentations and Agent, so heavy non-coding use could freeze coding access. The new coding plan removes that shared monthly cap and dedicates its quota to Kimi Code, but it keeps both a weekly quota and a rolling five-hour rate window.
| Term | Legacy plan | New Kimi Code plan |
|---|---|---|
| Coding quota | Shared with general Kimi membership | Dedicated to Kimi Code |
| Kimi app and web benefits | Included | Excluded; separate subscription required |
| Monthly total cap | Yes | No |
| Weekly quota | Yes | Yes |
| Rolling five-hour window | Yes | Yes |
For overseas customers, Kimi lists four monthly coding tiers at $19, $39, $99 and $199, with a limited first-month discount whose amount is not given in the retained documentation. New customers who want both coding and general-product benefits must buy them separately.
Legacy subscribers can instead keep their existing bundle and auto-renew until they cancel. If they switch, Kimi says the new plan takes effect immediately, the unused legacy value is refunded on a prorated basis, and the coding quota is reset and recalculated under the new rules.
When a weekly or five-hour quota runs out, a subscriber can upgrade or turn on Extra Usage. That metered fallback draws from a prepaid wallet shared by Kimi Web and Kimi Code and is billed at rates the company says are close to its API pricing. Subscription allowances are used first; Extra Usage is deducted last.
The separation gives Kimi distinct controls over coding and general-product consumption. Kimi says that will help it match compute more precisely and keep service stable. But announcing the redesign with the capacity pause does not show that the preceding 48-hour surge caused it, and quota partitioning cannot itself expand the GPU fleet.
Moonshot describes K3 in its launch material as a mixture-of-experts model with 2.8 trillion total parameters and a one-million-token context window. The same document says it activates 16 of 896 experts. That activation scope is essential context: the total parameter count alone is not a measure of compute consumed by every generated token.
For third-party deployments, Moonshot recommends supernodes with at least 64 accelerators so inference can use a larger high-bandwidth communication domain. This is a deployment recommendation, not evidence of how many accelerators Kimi operates, how they are connected or how much traffic one such configuration can serve. Moonshot also said full model weights would be released by July 27, so self-hosting was a promised alternative rather than a fully available substitute at the article's evidence cutoff.
The public API offers a limited view of Kimi's charging model. Moonshot lists $0.30 per million cache-hit input tokens, $3 per million cache-miss input tokens and $15 per million output tokens. It claims its official API exceeds a 90% cache-hit rate in coding workloads.
The ten-to-one difference between cached and uncached input prices makes repeated-context reuse economically important. It still cannot be converted into subscriber margins: API prices are charges to customers, not Kimi's internal costs, and the claimed cache-hit rate covers coding workloads rather than every use of Kimi Web, Work or the app.
Moonshot has attracted substantial recent financing, but the available figures do not show its current purchasing or deployment capacity. A May funding report, citing Huafeng Capital, which advised some investors in the transaction, said Moonshot raised about $2 billion at a $20 billion valuation. Huafeng's post also put the company's fundraising over the preceding six months at $3.9 billion and said annual recurring revenue exceeded $200 million in April, driven by paid subscriptions and API use.
Those are attributed financing and revenue claims, not disclosures of cash on hand, accelerator orders or installed capacity. Money can fund expansion, but the retained evidence does not say how quickly it can be turned into a working cluster.
The same report says Yang Zhilin, formerly a researcher at Meta AI and Google Brain, founded Moonshot in 2023. It places Kimi in a crowded field that includes OpenAI, Google and Anthropic as well as ByteDance, Alibaba, Zhipu and DeepSeek. Users facing Kimi's subscription freeze therefore have competing hosted services, while eventual access to K3's weights could create a self-hosted option for operators able to meet its infrastructure demands.
Demand is not, by itself, a comparative performance result. Moonshot's own launch material says K3 still trails the strongest proprietary models in overall performance. Its published evaluations also use different agent harnesses for some model comparisons, include internal tests, and note task fallbacks and hardware substitutions in certain benchmarks. Those qualifications do not negate the reported surge; they limit any attempt to treat it as proof that K3 has won the broader model market.
Moonshot's Mooncake paper shows that overloaded service was an engineering concern well before this subscription pause. The researchers described a Kimi serving platform that separates prompt processing from token generation, uses otherwise underused CPU, DRAM and SSD resources for a distributed key-value cache, and can reject requests early under predicted heavy overload.
The paper reports that Mooncake let Kimi handle 75% more requests than its baseline under real workloads. The comparison is between serving architectures; the abstract supplies neither an absolute request count nor evidence about K3's current fleet. It cannot be read as 75% spare capacity, a current utilization measure or proof that the same improvement applies unchanged to this model and traffic mix.
Kimi's next decision is how much new capacity to expose, and how quickly, without weakening service for the members it chose to protect. A batch reopening will be meaningful only if it is sustained and accompanied by evidence that performance and quotas remain stable.
The most useful disclosures would be the scale and schedule of the capacity addition, the number of subscription places reopened, service-level performance before and after each batch, and the frequency with which weekly or five-hour coding limits bind. Kimi could also clarify how general and coding plans are prioritized when demand peaks.
Until those figures appear, the pause supports a narrow conclusion: Kimi says K3 demand exceeded what it was prepared to serve, so it chose rationing and a phased reopening. It does not quantify the shortage, prove how long it will last or show that the membership split supplies the missing compute.
Get concise AI news and useful context from the Magica team.
Read the newsletterAnt Digital Technologies has expanded Agentar with 200 preconfigured job-role templates and a multi-agent management pitch, but it has not disclosed pricing, customer use or performance data—and governance is becoming a market-wide requirement rather than a distinctive feature.
Apple is reportedly testing an opt-in tool that transcribes and summarizes Genius Bar appointments. Its current safeguards are clear, but its accuracy, retention rules and uses after the pilot are not.
Moonshot AI is seeking investor approval for a Hong Kong IPO within six months while an unfinished private round could value it above $30 billion. The pitch pairs rapid reported recurring-revenue growth with Kimi K3, but neither audited financials nor enough independent model and deployment evidence is public yet.
Andy Serkis says machine learning has a narrow role in The Hunt for Gollum’s de-aging work. That boundary remains a production claim, because the actors, shots, tools, data, labor effects and likeness terms have not been disclosed.
Nebius Group revenue rose 684% to $399 million in the first quarter of 2026 while purchases of property, equipment and intangible assets reached $2.47 billion. Microsoft and Meta reduce demand risk, but options and an unsold-capacity backstop leave delivery, financing and unit economics unresolved.
Chinese-developed models have overtaken U.S. rivals in token volume on OpenRouter, where low prices and token-heavy agent workloads favor DeepSeek V4. The crossover covers a small, platform-specific slice of AI use and does not establish leadership in revenue, enterprise demand or infrastructure control.
OpenAI said it had identified the cause of elevated ChatGPT errors and was applying mitigations, but the retained status update did not disclose the technical failure, measure the impact or confirm full recovery.
Alibaba has opened hosted access to Qwen3.8-Max-Preview and says the 2.4-trillion-parameter model will be released with downloadable weights, but the license, architecture, benchmark evidence, release date and model-level credit economics remain undisclosed.
Anthropic has put Bun’s Rust port into Claude Code ahead of Bun 1.4’s general release, creating a real but tightly controlled proving ground for an AI-led migration whose total cost and broader reliability remain unsettled.
Zhipu reportedly reached $1 billion in annual recurring revenue in July, roughly four times a March estimate, but the unconfirmed run rate is not annual sales and still sits far ahead of recognized cloud revenue while margins remain thin.
An account of PNC transaction data puts household-paid generative AI near 2%, while a separate user survey finds much broader paid access when employer-funded plans count. The gap shows why card charges alone cannot settle whether consumer AI is becoming a mass subscription business.
Kimi K3 reduces attention traffic, but Moonshot recommends deploying it across at least 64 accelerators. SemiAnalysis says expert routing will more than erase the bandwidth savings; until the promised weights are deployed independently, that remains a hardware thesis rather than a measured result.
Morgan Stanley raised its Micron fiscal-2027 gross-margin estimate to 89.3%, but Micron’s results show the forecast depends chiefly on exceptional memory pricing, customer contracts and delayed supply rather than a disclosed HBM4 margin advantage.
New Mexico’s land commissioner refused to reconsider state-land crossings for a pipeline serving Project Jupiter, preserving a fuel-supply obstacle for the planned Oracle data center. But an analyst’s 2029 forecast predates that decision and remains at odds with Oracle’s first-half-2027 delivery statement.
A reported CIA mission examined whether an influential Emirati sheikh could be trusted with sensitive U.S. technology. The public export rule that followed gives G42 and Core42 a narrow, temporary exception, but it does not connect the intelligence operation to that decision or show how compliance will be tested.
Tracebit says a guardrail-triggering string in one decoy AWS secret sharply reduced five AI agents' success in a 152-run cyber range, but the company-run test did not cover uncensored models, adaptive attackers or production deployments.
SenseTime’s U1 Pro preview combines a claimed native 8K ceiling with a multi-step image-creation loop, but its August API will need to disclose dimensions, latency, pricing and repeatable results before buyers can compare the cost of a usable asset.
Alibaba Cloud has begun invite-only testing of a 64-card Zhenwu M890 supernode instance, extending its in-house chip and infrastructure stack to outside users while leaving the service's price, benchmark methodology and customer economics undisclosed.
Nvidia's 616.00 driver and CUDA 13.4 preview let developers begin native Windows Arm64 work for RTX Spark, but the release is an ecosystem-building step—not evidence of final performance, compatibility or pricing.
Open Design has attracted nearly 80,000 GitHub stars with an open, model-flexible answer to Claude Design, but its million-install claim has no published methodology and the team has not disclosed the retention and revenue figures needed to judge the business.