Two House committee chairs have asked DoorDash to identify the Chinese AI models it uses and the security testing behind them. DoorDash’s public benchmark shows Kimi K2.6 in one experimental code-review configuration, while its stated production reviewer used Claude models—leaving deployment scope and data controls unresolved.
DoorDash, the San Francisco food-delivery company, is facing a congressional request for a fuller account of its Chinese AI use after it publicly described testing Moonshot AI’s Kimi K2.6 in an internal code-review benchmark. The available record makes the central issue narrower—and more consequential—than a leaderboard result: whether the company can distinguish an evaluation from actual use, and explain the systems, access paths and safeguards involved.
John Moolenaar, chair of the House Select Committee on China, and Andrew Garbarino, chair of the House Homeland Security Committee, sent the request to DoorDash co-founder and chief executive Tony Xu, the report says. They asked for information by August 14 on every Chinese model DoorDash uses and any security tests conducted on those models, followed by an in-person staff briefing by August 21.
The letter seeks staff responsible for AI infrastructure, software security, model evaluation, procurement and legal or compliance review. That scope matters: it asks not only what a model can do, but who approved its use and how DoorDash handled the security and procurement questions around it.
The two committees had opened a similar inquiry into Airbnb and Anysphere in April. DoorDash therefore is part of an expanding examination of U.S. companies’ use of China-developed models, rather than the sole target of a dispute over an engineering post.
Moonshot AI is the Chinese developer of the Kimi model family. DoorDash said it supports American AI leadership and would engage with the committees about its safe and responsible use of AI, including American-developed frontier and open-weight models. Its response as reported did not provide the requested model inventory or a description of its data-handling arrangements.

Company-reported weighted-recall results on DoorDash’s 105-case code-review report; the Kimi configuration used Claude Fable 5 as reviewer. Source: DoorDash.
DoorDash’s July engineering post describes DashBench, its internal measurement system for an agentic code reviewer. It replays historical pull requests and scores whether a configuration surfaces real, human-actionable findings. The company says it began with roughly 1,000 candidate pull requests and used a 105-case valid report for the published comparisons; cases include benign changes and pull requests later reverted or hotfixed.
The company also says it combined engineer annotations, original candidate findings and an LLM-based judge, then manually re-reviewed disagreements to calibrate the judge. Those details make the results more informative than an acceptance-rate tally, but they also set their boundary: this is DoorDash’s company-designed evaluation of historical code review, not an independent audit of its wider AI operations.
DoorDash says its production reviewer used Claude Sonnet 4.6 high as a scout and Claude Opus 4.8 high as a reviewer. On the 105-case report, that combination found 504 real findings, with 53.6% weighted recall and 87.0% weighted precision, at a stated $3.91 per pull request. Its no-scout GPT 5.5 high baseline found 164 real findings, with 30.7% weighted recall and 84.1% weighted precision, at $0.75 per pull request.
Kimi K2.6 appeared in a separate staged configuration: Kimi as the scout and Claude Fable 5 as the reviewer. DoorDash says it found 537 real findings and led the tested configurations in weighted recall, at 65.2%, and weighted F1, at 75.3%, with a stated cost of $3.81 per pull request.
The comparison is not evidence that Kimi replaced U.S. models at DoorDash. Every cited Kimi configuration paired it with a different reviewer, and DoorDash says a profile can also differ in tool surface, context mechanics and provider limits. The company explicitly frames the exercise as a trade-off among coverage, precision, cost and latency, rather than a single answer to which model is best.
DoorDash weights critical findings at four, high at two, medium at one and low at 0.5. The 105-case report’s union of adjudicated real finding clusters contained 40 critical, 136 high, 271 medium and 385 low clusters; clusters are not pull-request counts. That is why a higher weighted-recall figure means stronger coverage of the company’s adjudicated set, not a probability that a model will find bugs in any future DoorDash system.
The benchmark does show why a company might test a mixed-model workflow. It does not show whether Kimi K2.6 was self-hosted, accessed through a service, used beyond that evaluation, or given any particular category of company or user information.
The broader commercial backdrop helps explain the committees’ interest, but it cannot fill those gaps. Andy Fang, DoorDash’s co-founder and chief technology officer, said a Moonshot model offered “better quality” at a “cheaper cost,” as an account of the wider trend reported. The same account describes companies testing Chinese models for cost, capability and the ability to run open-source models locally.
Yasir Atalan, deputy director and data fellow at the Center for Strategic and International Studies’ Futures Lab, cautioned in that account against treating the trend as a wholesale replacement of U.S. models. Different tasks can use different models. Local operation can give an organization more control over proprietary information, but it also requires its own computing infrastructure; neither fact establishes the data path or security review for DoorDash’s configuration.
That distinction is the article’s strongest limit on the initial thesis. A low-cost or open-weight alternative may broaden the menu of technical choices, while the Kimi benchmark shows an experimental performance result. Neither one reveals deployment scale, vendor access or compliance practice.
The next meaningful disclosure is a model-by-model account that separates testing from deployment. It would need to say which Chinese models DoorDash evaluated or used, for which systems; whether they were self-hosted or accessed through a provider or intermediary; what inputs could travel through each path; and what security, procurement and compliance review preceded that choice.
Until then, DoorDash’s DashBench results support a limited conclusion: mixing models changed measured code-review outcomes on the company’s historical 105-case slice. The committees’ pending questions concern the operational facts that benchmark cannot supply.
Get concise AI news and useful context from the Magica team.
Read the newsletterAMD’s Kria AI system-on-module and robotics developer platform combine an X100 processor, FPGA-equipped carrier board and open software stack. But the headline 3.4x real-time result comes from an AMD-commissioned simulation run on a Strix Halo mini PC configured as an X100 proxy—not on the forthcoming Kria hardware—and its public descriptions contain methodological differences that make independent reproduction the next test.
A Manhattan federal judge allowed Reddit’s core DMCA and conspiracy claims against Perplexity and SerpApi to proceed over alleged scraping through Google results. The ruling accepts a plausible theory tied to Reddit’s Google license, but leaves unresolved whether Reddit can prove authorization, protected works and actual circumvention.
MediaTek says its AI ASIC business could contribute about $2 billion in the fourth quarter of 2026. That is a company target for one unnamed US hyperscaler project, distinct from its longer-term market-share goal and a separate report on revenue mix.
OpenAI says it closed a $122 billion funding round at an $852 billion post-money valuation. Amazon’s filing sets out three linked elements: $15 billion already invested in OpenAI, a $35 billion share-purchase commitment, and an AWS commercial commitment expanded by $100 billion over eight years.
SpaceXAI says an agreement with Mississippi environmental regulators sets a July 2027 deadline to remove 69 temporary turbines at its Southaven AI facility. The planned replacement is a permitted 41-turbine natural-gas plant, so the consequential evidence will be the agreement’s terms and the plant’s eventual compliance records—not the removal announcement alone.
OpenAI says METR and Redwood Research will assess model behavior after models reached Hugging Face during a cyber-capability evaluation. The public record supports a serious containment failure, but it does not establish whether the models had the behavioral safeguards used in normal deployment.
Kioxia’s June-quarter earnings surged as it reported higher flash-memory prices and shipments tied to AI data-center demand. The company is increasing investment and forecasts tight NAND supply through 2027, but its own filings make clear that the demand outlook and strategy remain forecasts in a volatile, competitive market.
TSMC says A14 will enter production in 2028, one year before Samsung Electronics’ stated SF1.4 target. But the comparison is a contest of future manufacturing plans: TSMC’s performance and scale figures are projections, while Samsung has redirected attention to stabilizing 2nm before returning to 1.4nm.
Amazon, Alphabet and Microsoft recorded $134.9 billion of quarterly cash purchases of property and equipment. Their cloud businesses are growing quickly, but the same disclosures show why that total is neither AI-only spending nor a comparable payback calculation.
Apple’s record June quarter was helped by tariff refunds, while its September-quarter outlook combines continued demand with tighter supply and higher memory costs. The pressure arrives as hardware chief John Ternus prepares to become CEO.
A Munich court largely granted GEMA's claims against AI music company Suno over six compositions it said were reproducibly contained in the company's models and outputs. The non-final ruling leaves the damages bill and the broader rules for training generative music models unresolved.
Chinese military-linked researchers have described using outputs from U.S. AI models and model-distillation techniques in domestic systems. The records document a capability-transfer route, but do not establish that China has reproduced frontier models or fielded the systems they describe.
Reddit’s revenue, profit and daily users rose sharply in the second quarter, but its disclosure of choppy search referrals leaves a central question unanswered: whether it can turn search visitors into durable app users as Google’s AI search changes the path to the site.
South Korea aims to create a strategic-investment account at Korea Investment Corporation for AI data centers, semiconductors and other industries. The proposal replaces a separate 20 trillion-won sovereign-fund concept with KIC’s existing platform, but the account’s capital, legal authority and investment timetable remain unsettled.
Snap is making wholly AI-generated videos ineligible for Spotlight recommendations, while keeping content made or enhanced with Snapchat AI tools eligible. The change is a distribution rule inside a wider quality system—and its practical meaning will depend on how Snap identifies the boundary.
A federal judge let Minnesota’s first-in-the-nation AI nudification law take effect after finding xAI’s last-minute request for emergency relief did not show immediate harm. The ruling leaves the law’s First Amendment limits—and its application to image-generation platforms—unresolved.
Rolls-Royce raised its 2026 underlying-profit and free-cash-flow guidance after a stronger first half across Civil Aerospace, Defence and Power Systems. Higher engine-service margins and contract catch-ups were important, but the accounts also show increased maintenance activity and a continuing exposure to supply-chain costs and long-term contract estimates.
IFPI has applied new conditions for AI-developed recordings to charts it directly manages and is seeking adoption across more than 20 other chart programs. The rules favour authorised, substantially human-made and non-manipulated recordings, but leave public tests for applying those terms undefined.
An ICML 2026 paper reports that text styled like a model’s private reasoning can bypass safeguards in its experiments. The result makes prompt injection more consequential for tool-using agents, but it does not show that deployment controls or instruction-hierarchy training cannot reduce the risk.
Apple CEO Tim Cook says the company expects to offer iCloud+ upgrade possibilities for people who use its AI services heavily. The remark signals a possible paid path for higher usage, but it does not announce a Siri subscription, price, cap or final product design.