AMD Splits Venice Between a 256-Core EPYC and Helios’ High-Frequency CPU
Editorial Team
••📖6 min read
AMD has detailed the 256-core EPYC 9996, its density-oriented Venice server CPU, alongside a broader 9006 platform with 16 memory channels and PCIe 6.0. But the Helios AI rack uses a separate high-frequency Venice design, and AMD’s largest rack-performance figures remain preliminary company projections rather than independent results.
AMD’s EPYC 9996 is a 256-core, 512-thread, 600 W density-focused server processor with 1,024 MB of L3 cache.
The 16-channel SP7 platform expands memory and I/O, but the high-frequency Venice CPU in AMD’s Helios AI rack is a different design from the 9996.
AMD’s up-to-3.3× rack result versus NVIDIA Vera is a preliminary, modeled 100 kW comparison across six mixed workloads—not an independent benchmark.
AMD has introduced the EPYC 9996 as the top-density model of its 9006 “Venice” server-CPU family: 256 cores, 512 threads, up to 4.1 GHz, 1,024 MB of L3 cache and 600 W of default CPU power, according to AMD’s specification page. The company says it is the highest-core-count publicly disclosed single-socket CPU as of July 2026.
That is a substantial step in socket density, but it is not the CPU AMD is putting in its flagship Helios AI rack. AMD’s keynote described Helios as using a high-frequency Venice design, while the 9996 is the dense version; a contemporaneous product account likewise says Helios uses the standard Zen 6 core at 5 GHz rather than the 256-core Zen 6c configuration. The distinction makes the launch less a single flagship claim than a segmentation of CPU roles around the same platform.
One Venice name, four different product tracks
EPYC is AMD’s server-processor line. Venice succeeds the EPYC 9005 generation, called Turin, which an independent technical assessment says remained broadly competitive nearly two years after launch, including against NVIDIA Vera and Amazon’s Graviton5 in the cited tests. The new family has four tracks with different cores, sockets and timing—not one interchangeable product.
EPYC 9006 SP7: the high-density line, including the 256-core 9996 and a new 16-channel memory platform.
EPYC 9006 SP8: lower-power server parts, capped at 128 Zen 6c cores or 96 standard Zen 6 cores, with eight memory channels.
EPYC 9006X SP7 (Venice-X): standard Zen 6 parts with extra 3D V-Cache; reports place their arrival in 2027.
EPYC 9006 LP (Verano): a low-power line for AI host-node use, with LPDDR5X memory, also planned for 2027.
The 9996’s compact Zen 6c cores trade frequency for core count. By contrast, AMD lists multiple 96-core standard-Zen-6 models at up to 5 GHz; the product tables put the 9996 at 4.1 GHz. That makes the high-core chip an option for throughput-heavy tasks rather than evidence that every Venice deployment will choose the most cores.
AMD has introduced ROCm.AI to guide setup, deployment and optimization on its hardware, while HIP and HIPIFY offer a separate source-level route from CUDA for embedded software. The unresolved question is whether those tools lower enough engineering risk to matter for large planned AI deployments.
Editorial Team
The platform distinction matters, too. SP7 brings 16 memory channels, support for DDR5-8000 and MRDIMM-12800, and PCIe 6.0. A platform analysis puts the configuration ceiling at 8 TB and 1.6 TB/s per SP7 CPU when populated with 512 GB modules. Those are maximum configurations, and the same account notes that such memory would be expensive; they do not describe a typical server purchase.
AMD’s company-reported geometric-mean rack-performance estimates in a modeled 100 kW rack. Source: AMD.
AMD’s company-reported methodology table for its modeled 100 kW rack comparison. Source: AMD.
A broader competitive target, with an uneven comparison
AMD is framing the 9006 generation around agentic-AI systems, where it says CPUs handle planning, retrieval and tool execution alongside GPU inference. Its argument rests on thread density and on the ability to keep accelerators supplied with CPU, memory and I/O resources. Those are AMD’s workload assumptions, not an independently established measure of agent capacity: its “agents per” claims use available CPU threads as a proxy and say actual capacity varies with the model, memory, software, orchestration and system configuration.
The company’s highest-profile rack claim has a similarly narrow scope. AMD’s methodology disclosure projects the EPYC 9996 at up to 3.3 times NVIDIA Vera’s performance within a modeled 100 kW rack constraint. It compares two-socket systems with 512 Venice cores and 176 Vera cores and reports the geometric mean of six workloads: estimated SPECrate 2017 integer performance, Java operations, NGINX/WRK web serving, Redis, Memcached and TPROC-C.
The figure is preliminary engineering projections or measurements as of July 22, AMD says, and manufacturers may use different system configurations. It is also not a clean measure of per-core performance: the configurations have radically different core counts and the result is a mixed-workload geometric mean inside a power budget. It is useful evidence of AMD’s intended rack-level positioning, but not proof of a 3.3× advantage for a particular customer application.
Intel is still a relevant alternative, particularly where its Advanced Matrix Extensions suit select AI work. The independent assessment also notes that Intel’s prior edge in memory-intensive workloads from MRDIMMs is narrowed by Venice’s support for second-generation MRDIMM-12800. AMD’s own comparison set now also includes NVIDIA Vera and Arm server CPUs; the actual choice will depend on software compatibility, workload mix, accelerator design and rack constraints rather than a core-count headline alone.
Helios is the larger commercial bet
Helios is AMD’s rack-scale AI system, combining MI455X accelerators, Venice CPUs, Pensando networking and ROCm software in one rack designed to operate as a single machine. AMD has positioned it against NVIDIA’s NVL72-style racks; keynote coverage says Microsoft was already publicly named as a prospective customer before the launch, and that AMD announced Helios was in full production with shipments starting in the third quarter.
That system framing gives AMD a reason to sell several CPU profiles: the Helios high-frequency CPU for accelerator-heavy nodes, dense parts such as the 9996 for high-throughput work, and lower-power or cache-heavy variants for other deployments. It also narrows the claim that a single 256-core chip is the foundation of the rack. The commercial question is whether customers value this coordinated platform—hardware, network and software—enough to choose it over NVIDIA, Intel or Arm-based alternatives.
Reported processor pricing reinforces that this is infrastructure procurement rather than a conventional upgrade. The published product table lists the 9996 at $14,904, but marks core-type details as not officially confirmed. That price does not establish system cost, power cost or total cost of ownership; the retained material supplies none of those like-for-like deployment comparisons.
Availability and independent testing will decide the claim
The rollout schedule itself is not fully reconciled in the retained reporting. One account says SP7 and SP8 become generally available in the fourth quarter of 2026, while another says SP7 begins then and SP8 follows in the first half of 2027. Both place Venice-X and Verano in 2027. AMD’s official product page marks its model specifications “subject to change.”
The next useful evidence is therefore narrower than another launch slide: independently tested, like-for-like systems that separate core density, frequency, memory bandwidth, power draw and software maturity. That testing would show whether the 9996 earns its place for CPU-heavy throughput, whether the Helios high-frequency design delivers the expected accelerator-host benefits, and how either compares with Xeon, Vera and other Arm options under a buyer’s actual rack and workload limits.
KTransformers v0.6.4 adds LoRA and hybrid fine-tuning to its CPU-GPU workflow for sparse mixture-of-experts models, expands an Ampere inference path and moves a scheduler endpoint to loopback after a report of unauthenticated pickle deserialization. Its published throughput figures are configuration-specific release claims, not a general comparison of fine-tuning methods.
Microsoft has put MAI-Image-2.5-Pro and MAI-Voice-2-Flash into public preview in Foundry. The launch extends the company’s in-house AI options, but its evidence for lower cost and broader deployment is chiefly tied to MAI-Image-2.5 and company-reported workloads—not an independently comparable test of the new image model.
Microsoft has made its MAI-Code-1-Flash coding model available in paid GitHub Copilot plans and says a related model is live in Excel. Its evidence points to a targeted effort to reduce serving costs, while leaving the scale, task coverage and economics of the broader rollout unresolved.
NVIDIA says it will commit $4 million over three years to let CNCF projects test on real GPUs. With Kubernetes’ core Dynamic Resource Allocation framework already generally available, the consequential question is whether shared validation can make diverse GPU implementations dependable without settling operators’ choices over isolation and control.
Microsoft has made MAI-Image-2.5 the end-to-end default for Bing Image Creator and put it into PowerPoint image-to-image features. The company says the PowerPoint deployment can reduce GPU costs by up to 84% versus GPT-Image-2, while its much larger OpenAI partnership remains in place.
Centrica says its roughly 1,300 proposed role reductions reflect lower customer contact and a wider transformation programme. The record shows AI is one tool in that programme, while the headline total reaches beyond the 500 contact-based roles it has specifically identified.
Meta is using a new mass-market film to argue that AI will extend its mission of connection. The company has real distribution, product and infrastructure plans behind the pitch, but the meaningful test is whether it makes access, controls and outcomes legible as those plans reach users.
Google is rolling out an optional selfie-video check for eligible Google Account holders who lose access to their usual devices. The feature adds a biometric reference to recovery, but it excludes several account types and leaves users dependent on Google’s unreported real-world performance and recovery decisions.
Google’s first ATLAS report finds Gemini use across occupations covering 88.4% of U.S. employment, but mostly in a limited share of each job’s tasks. The data offer a detailed view of how Google’s AI is used; they do not show whether it is raising output, eliminating work or changing hiring.
Nokia booked €2.8 billion in AI and cloud order intake in the second quarter, as optical and IP-network sales grew. The orders strengthen its data-centre strategy, but they are not revenue yet, and restructuring, supply constraints and negative cash flow leave execution as the central question.
Super Micro Computer says it received more than $60 billion in fiscal fourth-quarter orders and now expects 15% to 17% gross margins, despite revenue tracking near the low end of its outlook. The preliminary figures point to a much larger AI-server delivery pipeline, but not yet to booked revenue or a durable improvement in profitability.
Nebius says it has brought up its first full NVIDIA Vera Rubin NVL72 rack in Finland. The deployment is an early integration milestone backed by a $2 billion NVIDIA private placement, but the company still has software and production testing to complete before customers can use it.
According to people familiar with the effort, Amazon founder and executive chairman Jeff Bezos has pressed Prime Video to make AI and personalization central to a proposed redesign known internally as Lighthouse. The early test could improve discovery, but Amazon has not said how it would reconcile personalized rankings with the paid placement and subscriptions that make Prime Video a complex storefront.
A report says some Chinese cities are gradually issuing robotaxi permits again after a review prompted by Baidu’s Wuhan outage. Shenzhen’s revised rules show that road testing, demonstrations and unmanned trials remain separately governed—and do not by themselves establish broad commercial service.
Databricks says it will run core operations and analytics on Azure Databricks, deepen Microsoft product integrations and expand use of Azure Cobalt chips. The partnership runs into the 2030s, but the companies have not disclosed its value, workload scale or customer results.
OpenAI is beginning a U.S. rollout of Health in ChatGPT for logged-in adults on web and iOS. The feature can bring optional Apple Health and supported medical-record context into ordinary chats, but connected data may be incomplete and requires user permission before use.
Yelp will supply ChatGPT with reviews, ratings, photos and business details, with branded links back to Yelp and a planned quote-request feature. The non-exclusive deal extends an existing licensing strategy, but neither company has disclosed its financial terms or how ChatGPT referrals will be measured and sold.
U.S. stocks fell as attacks on Saudi tankers added a second threat to Middle East oil flows and earnings from Alphabet and Tesla highlighted the cash demands of the AI buildout. The selloff was concentrated in megacap technology rather than a uniform retreat from equities.
Blackstone’s second-quarter distributable earnings rose 26% as realizations and fee-related earnings increased. Its data-center and other AI-linked investments are becoming a larger part of the firm’s portfolio narrative, but the results do not separately show how much of the gain came from AI while retail private-credit fundraising slowed.