Alibaba Cloud has begun invite-only testing of a 64-card Zhenwu M890 supernode instance, extending its in-house chip and infrastructure stack to outside users while leaving the service's price, benchmark methodology and customer economics undisclosed.
Alibaba Cloud has turned 64 of its in-house Zhenwu accelerators into a public-cloud infrastructure unit, but only for invited testers. The July 18 release moves Alibaba's integrated chip, interconnect, network, storage and software stack closer to customers; it does not yet establish what that integration costs or how it performs on a comparable workload.
The Lingjun Zhenwu M890 instance is available for invited testing in Ulanqab, according to a launch account relaying Alibaba's announcement. Alibaba described this as the first time it had offered AI compute in a supernode form through its public cloud. That is a company-specific milestone, not a claim that Alibaba invented rack-scale cloud compute.
Alibaba says the unit supports FP8 and FP4 computation. Its ICN Switch 1.0 expands the tightly connected, scale-up domain from 16 cards in the previous design to 64, and the company gives an inter-card bandwidth figure of 800 GB/s. The available material does not specify whether that bandwidth is per link, per card or measured on an aggregate basis, so it cannot support a like-for-like performance comparison with rivals.
The company also claims that the M890 instance delivers three times the training performance of the earlier Zhenwu 810E in autonomous-driving and embodied-AI workloads. No model, dataset, framework, convergence target, power envelope or 810E configuration accompanies that multiplier in the retained reports.
For inference, Alibaba says the instance's memory and symmetric high-speed fabric can serve a mixture-of-experts model with 10 trillion total parameters. That is not a training result. Nor does total parameter count show how many parameters are active for each token, the precision used to store weights, throughput, latency or the number of concurrent users.
The cloud instance is also not Alibaba's largest M890 configuration. At the same conference, T-Head displayed an M890 × Panjiu AL128 system containing 128 chips in one cabinet, with claimed petabyte-per-second internal bandwidth and hundred-nanosecond latency, according to the conference report. The 64-card instance is therefore a chosen cloud-service boundary, not the hardware ceiling.
Alibaba presents the instance as one component of its Lingjun infrastructure. The company says its HPN 8.0 network can combine as many as 130,000 heterogeneous cards in one cluster and could later extend to one million. It also advertises proactive fault monitoring, minute-level recovery and 99.7% average availability, according to a report based on company specifications.
Those figures describe the broader platform, not observed performance of one M890 instance. The 99.7% figure has no disclosed measurement window, workload or service-level remedy in the available material. Likewise, the one-million-card figure is an expansion direction rather than a deployed cluster count.
The storage numbers require the same separation. Alibaba says a rebuilt CPFS parallel file system can support roughly 100 PiB in one file system, throughput on the order of 100 TB/s and about 100 million input/output operations per second. They are vendor-stated system limits without a test configuration here; they should not be read as storage delivered to every 64-card customer.
Software is part of the offer. T-Head opened its SAIL stack, spanning operating-system, software-development-kit and interface layers, and said it works with more than 260 training and inference frameworks. That breadth could lower porting work, but a compatibility count does not reveal which operators are optimized, whether model outputs remain stable or how much performance is lost relative to another accelerator stack.
The M890 is not starting from a laboratory-scale installed base. Alibaba said in its fiscal-year update that more than 100,000 Zhenwu processing units were deployed on its public cloud, supporting more than 30 automakers and autonomous-driving companies. It said more than 60% of T-Head's compute capacity served external customers. The disclosure does not say how many of those units are M890s, or whether “compute capacity” is measured in chips, accelerator-hours or another unit.
Alibaba also said AI-related product revenue had grown at a triple-digit annual rate for 11 consecutive quarters and represented 30% of Cloud Intelligence Group's external revenue in the March quarter. It put annualized AI-related product revenue above 35.8 billion yuan. Those are company-defined revenue measures, but they show that the M890 enters an existing commercial cloud business rather than functioning only as a technology demonstration.
The investment burden remains substantial. For the January-to-March quarter, Cloud Intelligence Group revenue rose 38% from a year earlier to 41.6 billion yuan, while Alibaba's total revenue grew 3% to 243 billion yuan. Alibaba as a whole recorded an operating loss of 848 million yuan, compared with a 28.5 billion yuan operating gain a year earlier; higher technology investment was one contributor to rising expenses, according to independent financial reporting. The company has pledged at least 380 billion yuan over three years for cloud and AI infrastructure.
M890 demand could help turn that spending into revenue, but public-cloud delivery does not itself prove low cost. Alibaba Cloud said selected Lingjun AI-compute and CPFS prices would change from April 18, citing higher procurement costs for core hardware in a pricing notice. The notice predates the M890 instance and does not give it a rate, so it is context for input-cost pressure—not evidence that this service became more expensive.

Magica chart of Alibaba’s company-reported Cloud Intelligence Group revenue growth, based on Alibaba earnings disclosures. Source: Associated Press.
Alibaba is early in exposing its own accelerator stack as a cloud instance, but it is not first to put a tightly coupled rack-scale domain in public cloud. CoreWeave made NVIDIA GB200 NVL72 instances generally available in February 2025. The liquid-cooled system connects 72 Blackwell GPUs and 36 Grace CPUs; NVIDIA says the GPUs operate as one NVLink domain with 130 TB/s of aggregate GPU bandwidth, according to the platform announcement.
That 130 TB/s domain total cannot be compared directly with Alibaba's unexplained 800 GB/s inter-card figure. The decision-relevant comparison is maturity: CoreWeave described general availability and provisioning through its managed Kubernetes service, whereas the M890 is still invite-only.
Huawei provides a second check on scale and software differentiation. At WAIC it displayed a 1,024-card Atlas 950 SuperPoD and claimed 1 exaflop at FP8, 2 exaflops at FP4, 256 TB of globally addressed memory and round-trip interconnect latency of 3 microseconds. Huawei also says more than 750 of its earlier 384-card supernodes have been deployed across more than 20 industries, according to its product disclosure.
Those vendor figures are not an independent benchmark against M890, and a system count cannot be compared with Alibaba's chip count. They do show that Alibaba cannot claim differentiation from scale alone. Huawei is also developing an open foundation-software ecosystem: it says its CANN community has 67 projects, more than 12.44 million lines of open-source code and over 3,500 monthly active developers. Muxi and ZTE introduced supernode products at the same conference as well. The competitive contest now spans interconnect, memory, cluster operations, software portability and availability—not merely the accelerator.

Huawei’s official image of the Atlas 950 SuperPoD. The company reports a 1,024-card configuration, 256 TB of unified address space and 3-microsecond round-trip latency. Source: Huawei.
The invited rollout can answer the central commercial question only if Alibaba or customers publish evidence that makes the M890 comparable. At minimum, that means a billing unit and contract terms alongside model identity, total and active MoE parameters, precision, batch size, latency, throughput, utilization and power consumption. Availability claims need an observation period, workload and failure definition; training claims need a dataset, framework, convergence target and fully specified 810E baseline.
Without those disclosures, the defensible conclusion is narrow. Alibaba has made a 64-card slice of its domestic AI supply chain testable as public-cloud infrastructure. The next decision is whether customers can move real workloads onto it at a competitive cost—not whether 64 cards can be made to look like one machine.
Get concise AI news and useful context from the Magica team.
Read the newsletterKimi says a 48-hour demand surge pushed its GPU capacity close to the limit, forcing a pause in new subscriptions as a separate Kimi Code plan prepares to change how coding access is bundled and rationed.
Ant Digital Technologies has expanded Agentar with 200 preconfigured job-role templates and a multi-agent management pitch, but it has not disclosed pricing, customer use or performance data—and governance is becoming a market-wide requirement rather than a distinctive feature.
Apple is reportedly testing an opt-in tool that transcribes and summarizes Genius Bar appointments. Its current safeguards are clear, but its accuracy, retention rules and uses after the pilot are not.
Moonshot AI is seeking investor approval for a Hong Kong IPO within six months while an unfinished private round could value it above $30 billion. The pitch pairs rapid reported recurring-revenue growth with Kimi K3, but neither audited financials nor enough independent model and deployment evidence is public yet.
Andy Serkis says machine learning has a narrow role in The Hunt for Gollum’s de-aging work. That boundary remains a production claim, because the actors, shots, tools, data, labor effects and likeness terms have not been disclosed.
Nebius Group revenue rose 684% to $399 million in the first quarter of 2026 while purchases of property, equipment and intangible assets reached $2.47 billion. Microsoft and Meta reduce demand risk, but options and an unsold-capacity backstop leave delivery, financing and unit economics unresolved.
Chinese-developed models have overtaken U.S. rivals in token volume on OpenRouter, where low prices and token-heavy agent workloads favor DeepSeek V4. The crossover covers a small, platform-specific slice of AI use and does not establish leadership in revenue, enterprise demand or infrastructure control.
OpenAI said it had identified the cause of elevated ChatGPT errors and was applying mitigations, but the retained status update did not disclose the technical failure, measure the impact or confirm full recovery.
Alibaba has opened hosted access to Qwen3.8-Max-Preview and says the 2.4-trillion-parameter model will be released with downloadable weights, but the license, architecture, benchmark evidence, release date and model-level credit economics remain undisclosed.
Anthropic has put Bun’s Rust port into Claude Code ahead of Bun 1.4’s general release, creating a real but tightly controlled proving ground for an AI-led migration whose total cost and broader reliability remain unsettled.
Zhipu reportedly reached $1 billion in annual recurring revenue in July, roughly four times a March estimate, but the unconfirmed run rate is not annual sales and still sits far ahead of recognized cloud revenue while margins remain thin.
An account of PNC transaction data puts household-paid generative AI near 2%, while a separate user survey finds much broader paid access when employer-funded plans count. The gap shows why card charges alone cannot settle whether consumer AI is becoming a mass subscription business.
Kimi K3 reduces attention traffic, but Moonshot recommends deploying it across at least 64 accelerators. SemiAnalysis says expert routing will more than erase the bandwidth savings; until the promised weights are deployed independently, that remains a hardware thesis rather than a measured result.
Morgan Stanley raised its Micron fiscal-2027 gross-margin estimate to 89.3%, but Micron’s results show the forecast depends chiefly on exceptional memory pricing, customer contracts and delayed supply rather than a disclosed HBM4 margin advantage.
New Mexico’s land commissioner refused to reconsider state-land crossings for a pipeline serving Project Jupiter, preserving a fuel-supply obstacle for the planned Oracle data center. But an analyst’s 2029 forecast predates that decision and remains at odds with Oracle’s first-half-2027 delivery statement.
A reported CIA mission examined whether an influential Emirati sheikh could be trusted with sensitive U.S. technology. The public export rule that followed gives G42 and Core42 a narrow, temporary exception, but it does not connect the intelligence operation to that decision or show how compliance will be tested.
Tracebit says a guardrail-triggering string in one decoy AWS secret sharply reduced five AI agents' success in a 152-run cyber range, but the company-run test did not cover uncensored models, adaptive attackers or production deployments.
SenseTime’s U1 Pro preview combines a claimed native 8K ceiling with a multi-step image-creation loop, but its August API will need to disclose dimensions, latency, pricing and repeatable results before buyers can compare the cost of a usable asset.
Nvidia's 616.00 driver and CUDA 13.4 preview let developers begin native Windows Arm64 work for RTX Spark, but the release is an ecosystem-building step—not evidence of final performance, compatibility or pricing.
Open Design has attracted nearly 80,000 GitHub stars with an open, model-flexible answer to Claude Design, but its million-install claim has no published methodology and the team has not disclosed the retention and revenue figures needed to judge the business.