Microsoft has put MAI-Image-2.5-Pro and MAI-Voice-2-Flash into public preview in Foundry. The launch extends the company’s in-house AI options, but its evidence for lower cost and broader deployment is chiefly tied to MAI-Image-2.5 and company-reported workloads—not an independently comparable test of the new image model.
Microsoft, the operator of Azure and a major OpenAI shareholder, has added two in-house generative-AI models to public preview in Microsoft Foundry: MAI-Image-2.5-Pro for image generation and editing, and MAI-Voice-2-Flash for text-to-speech. The release gives developers more choices inside Microsoft’s platform, but the company’s strongest operating evidence belongs to related models already in production—not to an independent test of the new image preview.
That distinction matters because the announcement is also a strategic claim. Microsoft says in its release that it began building purpose-built MAI models in-house a year ago and is now choosing models for individual product surfaces on a quality-speed-cost curve. The available evidence supports a growing internal model portfolio; it does not show that Microsoft has replaced outside models everywhere, or that Image Pro has already proved itself in the same production settings as MAI-Image-2.5.
Microsoft’s announcement calls MAI-Image-2.5-Pro its highest-fidelity image model to date, aimed at detailed generation, editing and in-image text. The Foundry model card identifies it as a preview, diffusion-based text-to-image and image-to-image model. Its stated strengths are controlled edits—such as changing an object or layout, updating text, cleaning artifacts and preserving visual consistency across iterations.
MAI-Voice-2-Flash is a text-to-speech model optimized for low-latency interactions. Microsoft positions it for voice agents, assistants and contact-center or IVR flows; the technical documentation says it supports 15 languages and 18 locales, expressive controls through SSML, and licensed prebuilt voices. A developer can also prompt a voice from a five- to 60-second reference clip, subject to the safeguards below.
| Model | Microsoft’s stated role | Published price | Important qualification |
|---|---|---|---|
| MAI-Image-2.5-Pro | Higher-fidelity image creation and controlled editing | $5 per million text-input tokens; $8 per million image-input tokens; $106 per million image-output tokens | Preview; the pricing separates text input, image input and image output. |
| MAI-Voice-2-Flash | Fast, high-volume text-to-speech | $15 per million characters | Preview; its price unit cannot be compared directly with the image model’s token-based rates. |
Microsoft says in its announcement that Voice Flash is twice as fast and 32% cheaper than MAI-Voice-2 while retaining that model’s natural prosody and acoustic quality. Those are the company’s comparisons between two of its voice models, rather than an independent benchmark or a comparison with a rival service.

Microsoft reports up to 84% GPU-cost reduction for MAI-Image-2.5 in PowerPoint and up to 89% for MAI-Voice-2-Flash in Dynamics 365 Contact Center; company-reported figures for different workloads. Source: Microsoft AI.
Microsoft says MAI-Image-2.5, rather than Pro, now powers Bing Image Creator end-to-end and is used for PowerPoint image-to-image capabilities and key OneDrive image-editing scenarios. In the PowerPoint comparison, Microsoft says MAI-Image-2.5 cut GPU costs by as much as 84% against GPT-Image-2, an image model from OpenAI, the AI company whose models and products Microsoft licenses through 2032 under their amended agreement.
For OneDrive, Microsoft reports in the release a 26% increase in save rates, an approximately 25% reduction in P95 latency and 2.5 times greater efficiency under medium-utilization production workloads after rollout. The measures have different denominators and baselines: save rate concerns a user action, P95 is a tail-latency measure, and efficiency is not defined further in the release. They establish Microsoft’s reported operating outcome for those scenarios, not a portable quality ranking for Image Pro.
The distinction has been blurred in coverage. One report initially described Image Pro as being used in PowerPoint, Bing and OneDrive, but its later wording and Microsoft’s original announcement identify MAI-Image-2.5 for PowerPoint and OneDrive. A separate account of the launch likewise attributes Bing, PowerPoint and OneDrive deployments to MAI-Image-2.5. The retained record therefore does not establish that Pro is serving those product experiences.
Voice Flash does have a named production use: Microsoft says in the release that it powers Dynamics 365 Contact Center, the company’s enterprise platform for building call-center agents, and has reduced GPU costs by up to 89% there. It is also integrated into Azure Voice Live for speech-to-speech agents. Those are again company statements for particular workloads, not an estimate of savings a customer should expect across regions or traffic patterns.
The MAI-Voice family is available through Azure Speech in Foundry Tools. The documentation ties Flash to Azure accounts, Speech resources and regions that support MAI models; it does not supply a general latency figure or a universal deployment-cost calculation.
More importantly, the documentation describes the family as a public preview without a service-level agreement and says it is not recommended for production workloads. Voice prompting requires limited-access approval and consent safeguards. Microsoft says only authorized, licensed voices can be synthesized in production, so a five- to 60-second clip is not a shortcut around the access and consent process.
That creates a meaningful boundary around the release. Microsoft can point to its own Dynamics deployment, while a third-party buyer considering Flash must still assess preview terms, regional availability, consent operations and its own throughput before treating the claimed speed or GPU savings as a budget forecast.
The new previews do broaden Microsoft’s ability to select its own models for workloads it controls. But the company’s April amended agreement with OpenAI preserved a substantial partnership: Microsoft remains OpenAI’s primary cloud partner, OpenAI products ship first on Azure unless Microsoft cannot and chooses not to support the needed capabilities, and Microsoft’s now non-exclusive license to OpenAI intellectual property runs through 2032. Microsoft also says it remains a major shareholder, while OpenAI can serve products across cloud providers.
So the immediate consequence is more model-selection control inside Microsoft’s product and cloud stack, not evidence of a clean break. The financial and infrastructure stakes still run both ways: the companies say they will continue work on new datacenter capacity and next-generation silicon, even as the amended agreement gives each more flexibility.
The next useful evidence is not another product assertion but comparable results for the previews themselves. For Image Pro, that means disclosed evaluation methods and production performance that separate it from MAI-Image-2.5, including the workload and output-token volume behind its $106-per-million-output-token price. For Voice Flash, it means customer deployments that identify region, traffic conditions, latency measurement and the operational effect of its limited-access voice-cloning controls.
Until then, the evidence supports a more precise reading: Microsoft has expanded a credible in-house MAI lineup and is using related versions in its products, but the commercial and technical case for each new preview remains specific to the workload—and, for the headline image model, largely still to be demonstrated.
Get concise AI news and useful context from the Magica team.
Read the newsletterAMD has introduced ROCm.AI to guide setup, deployment and optimization on its hardware, while HIP and HIPIFY offer a separate source-level route from CUDA for embedded software. The unresolved question is whether those tools lower enough engineering risk to matter for large planned AI deployments.
KTransformers v0.6.4 adds LoRA and hybrid fine-tuning to its CPU-GPU workflow for sparse mixture-of-experts models, expands an Ampere inference path and moves a scheduler endpoint to loopback after a report of unauthenticated pickle deserialization. Its published throughput figures are configuration-specific release claims, not a general comparison of fine-tuning methods.
Microsoft has made its MAI-Code-1-Flash coding model available in paid GitHub Copilot plans and says a related model is live in Excel. Its evidence points to a targeted effort to reduce serving costs, while leaving the scale, task coverage and economics of the broader rollout unresolved.
AMD has detailed the 256-core EPYC 9996, its density-oriented Venice server CPU, alongside a broader 9006 platform with 16 memory channels and PCIe 6.0. But the Helios AI rack uses a separate high-frequency Venice design, and AMD’s largest rack-performance figures remain preliminary company projections rather than independent results.
NVIDIA says it will commit $4 million over three years to let CNCF projects test on real GPUs. With Kubernetes’ core Dynamic Resource Allocation framework already generally available, the consequential question is whether shared validation can make diverse GPU implementations dependable without settling operators’ choices over isolation and control.
Microsoft has made MAI-Image-2.5 the end-to-end default for Bing Image Creator and put it into PowerPoint image-to-image features. The company says the PowerPoint deployment can reduce GPU costs by up to 84% versus GPT-Image-2, while its much larger OpenAI partnership remains in place.
Centrica says its roughly 1,300 proposed role reductions reflect lower customer contact and a wider transformation programme. The record shows AI is one tool in that programme, while the headline total reaches beyond the 500 contact-based roles it has specifically identified.
Meta is using a new mass-market film to argue that AI will extend its mission of connection. The company has real distribution, product and infrastructure plans behind the pitch, but the meaningful test is whether it makes access, controls and outcomes legible as those plans reach users.
Google is rolling out an optional selfie-video check for eligible Google Account holders who lose access to their usual devices. The feature adds a biometric reference to recovery, but it excludes several account types and leaves users dependent on Google’s unreported real-world performance and recovery decisions.
Google’s first ATLAS report finds Gemini use across occupations covering 88.4% of U.S. employment, but mostly in a limited share of each job’s tasks. The data offer a detailed view of how Google’s AI is used; they do not show whether it is raising output, eliminating work or changing hiring.
Nokia booked €2.8 billion in AI and cloud order intake in the second quarter, as optical and IP-network sales grew. The orders strengthen its data-centre strategy, but they are not revenue yet, and restructuring, supply constraints and negative cash flow leave execution as the central question.
Super Micro Computer says it received more than $60 billion in fiscal fourth-quarter orders and now expects 15% to 17% gross margins, despite revenue tracking near the low end of its outlook. The preliminary figures point to a much larger AI-server delivery pipeline, but not yet to booked revenue or a durable improvement in profitability.
Nebius says it has brought up its first full NVIDIA Vera Rubin NVL72 rack in Finland. The deployment is an early integration milestone backed by a $2 billion NVIDIA private placement, but the company still has software and production testing to complete before customers can use it.
According to people familiar with the effort, Amazon founder and executive chairman Jeff Bezos has pressed Prime Video to make AI and personalization central to a proposed redesign known internally as Lighthouse. The early test could improve discovery, but Amazon has not said how it would reconcile personalized rankings with the paid placement and subscriptions that make Prime Video a complex storefront.
A report says some Chinese cities are gradually issuing robotaxi permits again after a review prompted by Baidu’s Wuhan outage. Shenzhen’s revised rules show that road testing, demonstrations and unmanned trials remain separately governed—and do not by themselves establish broad commercial service.
Databricks says it will run core operations and analytics on Azure Databricks, deepen Microsoft product integrations and expand use of Azure Cobalt chips. The partnership runs into the 2030s, but the companies have not disclosed its value, workload scale or customer results.
OpenAI is beginning a U.S. rollout of Health in ChatGPT for logged-in adults on web and iOS. The feature can bring optional Apple Health and supported medical-record context into ordinary chats, but connected data may be incomplete and requires user permission before use.
Yelp will supply ChatGPT with reviews, ratings, photos and business details, with branded links back to Yelp and a planned quote-request feature. The non-exclusive deal extends an existing licensing strategy, but neither company has disclosed its financial terms or how ChatGPT referrals will be measured and sold.
U.S. stocks fell as attacks on Saudi tankers added a second threat to Middle East oil flows and earnings from Alphabet and Tesla highlighted the cash demands of the AI buildout. The selloff was concentrated in megacap technology rather than a uniform retreat from equities.
Blackstone’s second-quarter distributable earnings rose 26% as realizations and fee-related earnings increased. Its data-center and other AI-linked investments are becoming a larger part of the firm’s portfolio narrative, but the results do not separately show how much of the gain came from AI while retail private-credit fundraising slowed.