OpenAI and Anthropic Differ on Who Can Stop Unsafe AI | Magica
OpenAI and Anthropic Agree on Safety Disclosure. Authority Is the Harder Question.
Editorial Team
••📖7 min read
OpenAI and Anthropic both support disclosure, incident reporting and outside scrutiny for the largest frontier-AI developers. Their proposals—and Illinois’ new law—show that the more consequential unresolved question is whether an evaluator or government agency can decide when a model’s risk is too high for deployment.
OpenAI and Anthropic both want the largest AI developers to disclose safety work, report serious incidents and face scrutiny beyond their own assessments.
Anthropic’s proposal would apply only to developers above both a computing threshold and a large revenue-or-spending threshold; it also concedes that a mature independent-evaluator market does not yet exist.
Illinois’ new law offers an early state test of disclosure and compliance audits. It does not make an audit a verdict that a model is safe to release.
OpenAI and Anthropic, companies developing frontier AI models, agree on a substantial part of the policy agenda: the biggest developers should document how they assess catastrophic risks, disclose findings, report serious incidents and submit to review. The disagreement that matters most is further downstream. Who gets to determine that a model has too much residual risk to deploy—and what power would that determination carry?
That distinction separates a reporting regime from one that can stop or restrict a model. It also makes Illinois’ Artificial Intelligence Safety Measures Act a useful, narrower test case. Gov. J.B. Pritzker, the state’s chief executive, signed the law on July 6; it requires qualifying developers to describe their process and undergo annual third-party audits, but the audit is aimed at compliance with that process.
Anthropic’s proposed scope for covered frontier-AI developers. Source: Anthropic.
A common baseline, with different ambitions
In a policy statement, Chris Lehane, OpenAI’s chief global affairs officer, calls the state-to-federal route “reverse federalism.” OpenAI says California, New York and Illinois have begun to align around three elements: a documented safety framework and public risk assessments, reporting of serious incidents, and independent, objective audits.
OpenAI’s case for national rules is also a case for federal control of the most technical work. It says state laws can supply an interim direction, while federal experts—with resources, access to classified systems and the capacity to work with developers—should test the most advanced systems. The company says the administration is developing a framework for U.S. government cyber testing of the most capable models, with a target of early August.
Mexico’s exports of computer equipment have surged as U.S. data-center construction and tariff differences reshape trade. The figures point to a bigger role in assembling and shipping equipment, but imported components, near-full factories and limited investment leave the higher-value parts of the chain elsewhere.
Editorial Team
The proposal is not simply a call for tougher requirements. OpenAI argues that an inconsistent state patchwork could divert resources, particularly at startups and small companies, and slow deployment of advanced tools to government, critical-infrastructure defenders, allies and other trusted partners. That competitive and national-security rationale limits any claim that the company is advocating a single-purpose safety regime.
Anthropic’s Advanced AI Framework is more explicit about scope and enforcement. It proposes obligations only for developers whose models require more than 10<sup>25</sup> training FLOP and who either earn more than $500 million in annual AI-derived revenue or spend more than $1 billion annually on AI research and development. The thresholds are meant to reach very large developers, not every business that uses or builds AI; the framework says the capability threshold may need to change if dangerous capabilities become cheaper to train.
Those covered developers would test and describe risks in four categories: biological weapons, offensive cyber operations, loss of control and automated research and development. The company proposes a public safety framework, a risk report at least every six months, and a system card when a covered model released for general access is materially more capable—or released with materially weaker safeguards—than comparable prior models. It would require reporting a defined critical safety incident to the designated agency within 15 days of discovery or facts that give the developer reasonable grounds to believe one occurred.
Anthropic’s proposed conditions for agency review and remedies. Source: Anthropic.
Outside review is proposed, not yet a settled institution
Anthropic says self-assessment is insufficient, but it also writes that “a mature independent evaluation ecosystem does not yet exist.” Within six months of a regulation taking effect, it proposes that covered developers regularly engage at least one qualified independent evaluator.
The proposed evaluator would receive unredacted risk reports and system cards, access to the developer’s most capable models, and the opportunity to ask questions. It would publish a review addressing the quality of the developer’s analysis, the materiality of public redactions and any disagreement with its risk conclusions. The framework also recognizes the incentives that can defeat that design: it suggests standards or licensing, pooled or public funding, conflict disclosures, and even random assignment of highly rated evaluators in high-stakes cases to reduce evaluator shopping.
That is a material constraint, not a procedural footnote. A rule can require comparable disclosures, but the disclosures do not independently establish whether safeguards are adequate. The proposal’s credibility depends on reviewers having access, technical competence and financial independence from the companies they examine.
Anthropic goes further than a disclosure-only model by outlining possible enforcement. An agency could review missing or inadequate reports, an evaluator’s independence and access, or a finding that a deployed model poses significant catastrophic risk even after safeguards. Possible remedies include fines, a bar on deploying further covered models until violations are corrected, and—in extreme cases—restrictions on models already deployed. But the framework presents those as options and pairs them with limits such as court enforcement, defined grounds for action, consistent treatment and judicial review. It does not resolve the tradeoff between weak enforcement and overly broad regulatory power.
Illinois tests compliance, not deployment readiness
Illinois’ law, as reported at the signing, applies to large model developers that generate more than $500 million in annual revenue and use massive computing power. It requires a published framework for identifying and assessing “catastrophic risk,” defined in the report as the likelihood of incidents causing death or serious injury to more than 50 people or more than $1 million in property damage.
Qualifying developers must report an incident that could harm Illinois within 72 hours of identifying it, or within 24 hours when it presents an imminent risk of death or serious physical injury. The attorney general can seek civil penalties of up to $1 million for a first violation and up to $3 million for subsequent violations. The law takes effect Jan. 1, 2028.
Its annual third-party audit requirement is the distinguishing feature among the state measures discussed in the report: New York’s comparable law required a single independent audit once developers became large enough to qualify. But annual review does not turn the Illinois audit into an assessment of whether a model should be in the world. Scott Wisor, policy director at Secure AI, which helped shape the Illinois bill, described the present test this way:
“Right now, the evaluation in this bill is, are you complying with your safety framework?”
Wisor’s point was that a developer could follow its stated steps and still release a risky system. TechNet, a coalition of technology executives, raised a different implementation concern during legislative debate: the law would ask private actors to make subjective safety-compliance determinations without established national standards, certifications or clear regulatory guardrails.
The state law therefore supplies a test of disclosure, timelines and audit practice—not evidence that either an audit or a published framework can settle the deployment decision. Its broader influence is also an estimate, not a measured national share: lawmakers said California, New York and Illinois represent roughly 40% of the U.S. AI market while accounting for roughly 20% of the population.
The federal choice is about consequences
OpenAI and Anthropic do share a federal-state boundary. OpenAI wants national capacity for the most demanding frontier-model testing. Anthropic says Congress should not preempt state law unless it enacts a rigorous federal regime that meets or exceeds the framework’s strongest catastrophic-risk measures; even then, it recommends narrow preemption limited to specified frontier-governance functions, including testing, evaluator licensing and closely related reporting. Outside those functions, it says states should retain authority.
The next evidence to watch is more concrete than competing statements of principle: the scope, testing standards, timelines and consequences in the promised federal cyber-testing framework; whether evaluators can obtain and protect the access they need; and whether lawmakers give an agency authority to act when an evaluation finds inadequately mitigated risk. Until those decisions are made, disclosures and compliance audits can make safety claims easier to examine, but they cannot answer who is empowered to prevent deployment.
AMD’s Kria AI system-on-module and robotics developer platform combine an X100 processor, FPGA-equipped carrier board and open software stack. But the headline 3.4x real-time result comes from an AMD-commissioned simulation run on a Strix Halo mini PC configured as an X100 proxy—not on the forthcoming Kria hardware—and its public descriptions contain methodological differences that make independent reproduction the next test.
A Manhattan federal judge allowed Reddit’s core DMCA and conspiracy claims against Perplexity and SerpApi to proceed over alleged scraping through Google results. The ruling accepts a plausible theory tied to Reddit’s Google license, but leaves unresolved whether Reddit can prove authorization, protected works and actual circumvention.
MediaTek says its AI ASIC business could contribute about $2 billion in the fourth quarter of 2026. That is a company target for one unnamed US hyperscaler project, distinct from its longer-term market-share goal and a separate report on revenue mix.
OpenAI says it closed a $122 billion funding round at an $852 billion post-money valuation. Amazon’s filing sets out three linked elements: $15 billion already invested in OpenAI, a $35 billion share-purchase commitment, and an AWS commercial commitment expanded by $100 billion over eight years.
Two House committee chairs have asked DoorDash to identify the Chinese AI models it uses and the security testing behind them. DoorDash’s public benchmark shows Kimi K2.6 in one experimental code-review configuration, while its stated production reviewer used Claude models—leaving deployment scope and data controls unresolved.
A proposed class action alleges Granola recorded and transcribed meeting participants without their consent and used their data to improve its AI. Granola says its product captures audio locally, offers optional notices and lets users opt out of anonymized-data model improvement; the case turns on how those controls worked in practice.
SpaceXAI says an agreement with Mississippi environmental regulators sets a July 2027 deadline to remove 69 temporary turbines at its Southaven AI facility. The planned replacement is a permitted 41-turbine natural-gas plant, so the consequential evidence will be the agreement’s terms and the plant’s eventual compliance records—not the removal announcement alone.
OpenAI says METR and Redwood Research will assess model behavior after models reached Hugging Face during a cyber-capability evaluation. The public record supports a serious containment failure, but it does not establish whether the models had the behavioral safeguards used in normal deployment.
Kioxia’s June-quarter earnings surged as it reported higher flash-memory prices and shipments tied to AI data-center demand. The company is increasing investment and forecasts tight NAND supply through 2027, but its own filings make clear that the demand outlook and strategy remain forecasts in a volatile, competitive market.
TSMC says A14 will enter production in 2028, one year before Samsung Electronics’ stated SF1.4 target. But the comparison is a contest of future manufacturing plans: TSMC’s performance and scale figures are projections, while Samsung has redirected attention to stabilizing 2nm before returning to 1.4nm.
Amazon, Alphabet and Microsoft recorded $134.9 billion of quarterly cash purchases of property and equipment. Their cloud businesses are growing quickly, but the same disclosures show why that total is neither AI-only spending nor a comparable payback calculation.
Apple’s record June quarter was helped by tariff refunds, while its September-quarter outlook combines continued demand with tighter supply and higher memory costs. The pressure arrives as hardware chief John Ternus prepares to become CEO.
A Munich court largely granted GEMA's claims against AI music company Suno over six compositions it said were reproducibly contained in the company's models and outputs. The non-final ruling leaves the damages bill and the broader rules for training generative music models unresolved.
Chinese military-linked researchers have described using outputs from U.S. AI models and model-distillation techniques in domestic systems. The records document a capability-transfer route, but do not establish that China has reproduced frontier models or fielded the systems they describe.
Reddit’s revenue, profit and daily users rose sharply in the second quarter, but its disclosure of choppy search referrals leaves a central question unanswered: whether it can turn search visitors into durable app users as Google’s AI search changes the path to the site.
South Korea aims to create a strategic-investment account at Korea Investment Corporation for AI data centers, semiconductors and other industries. The proposal replaces a separate 20 trillion-won sovereign-fund concept with KIC’s existing platform, but the account’s capital, legal authority and investment timetable remain unsettled.
Snap is making wholly AI-generated videos ineligible for Spotlight recommendations, while keeping content made or enhanced with Snapchat AI tools eligible. The change is a distribution rule inside a wider quality system—and its practical meaning will depend on how Snap identifies the boundary.
A federal judge let Minnesota’s first-in-the-nation AI nudification law take effect after finding xAI’s last-minute request for emergency relief did not show immediate harm. The ruling leaves the law’s First Amendment limits—and its application to image-generation platforms—unresolved.
Rolls-Royce raised its 2026 underlying-profit and free-cash-flow guidance after a stronger first half across Civil Aerospace, Defence and Power Systems. Higher engine-service margins and contract catch-ups were important, but the accounts also show increased maintenance activity and a continuing exposure to supply-chain costs and long-term contract estimates.