Anthropic’s book shredding was fair use; its pirate library was not | Magica
Anthropic’s book shredding was fair use; its pirate library was not
Editorial Team
••📖6 min read
A federal judge treated Anthropic’s purchased-book scanning and its model training as distinct fair uses, while leaving its earlier pirate-library copying exposed. Newly unsealed Project Panama records show how the company tried to replace that source of books at industrial scale.
A federal judge found fair use in Anthropic’s training copies and in its one-for-one digitization of print books it had bought.
The same order left a separate claim over more than 7 million pirated library copies for trial; Anthropic later settled that case for $1.5 billion without admitting wrongdoing.
Newly unsealed Project Panama documents make the operation’s scale visible: an industrial effort to obtain book-quality text without relying on the earlier shadow-library archive.
Anthropic’s decision to cut the bindings from purchased books, scan every page and recycle the paper was legally unusual but not, on the record before the court, the copyright violation. The more consequential distinction was between a paid-for print copy converted into one internal digital replacement and an unauthorized copy downloaded to build a permanent general-purpose library.
Anthropic PBC, the AI company founded in 2021 by former OpenAI employees, operates Claude, the text-generating service it first released publicly in March 2023. In a June 2025 order, U.S. District Judge William Alsup treated the company’s book pipeline as separate uses with separate fair-use outcomes. That is a narrower conclusion than a general approval of an AI company’s data collection: it depends on how the copies were obtained, what was done with each copy and what the plaintiffs had actually alleged about outputs.
The order separated Anthropic’s conduct into three categories.
Copies selected to train particular large language models: fair use. The judge found the training use transformative on this record, which did not include an allegation that Claude had given users exact copies or substantial knockoffs of the plaintiffs’ books.
Print books Anthropic bought, then converted to internal digital replacements: fair use for a different reason. The company destroyed the print original, made one searchable digital replacement and, the order said, did not share or sell that replacement outside the company.
Enigma has emerged from stealth with a $71 million seed round and a public test involving at least 100 real robots. The startup says the interactions could improve robot interfaces and models, but it has not disclosed a specific commercial use case, customers or performance evidence.
Editorial Team
More than 7 million copies acquired from Books3, LibGen and Pirate Library Mirror: not cleared as fair use when held in a permanent, general-purpose library. The court described a plan to store everything “forever,” including books Anthropic had decided not to use for model training.
That final category matters because a later fair use does not retroactively justify the initial acquisition. Alsup wrote that a separate justification was required for the library use; the order also declined to bless any additional copies engineers made from the library for purposes other than model training because the record was underdeveloped.
The case was not a finding that every model trained on books is lawful. It was a ruling on summary judgment, drawn from the claims and evidence in this lawsuit. In particular, a future claim based on allegedly infringing model outputs was outside the record, not rejected on the merits.
Project Panama was an industrial replacement for a risky source
The court record says Anthropic valued books for their well-curated facts, organized analysis and strong prose. It had initially obtained at least 196,640 books from Books3, at least 5 million from LibGen and at least 2 million from Pirate Library Mirror. The company later became concerned about using pirated ebooks for legal reasons, the order says, while retaining the copies.
In February 2024, Anthropic hired Tom Turvey, formerly Google Books’ head of partnerships, to help obtain “all the books in the world.” That prior role was material: Google Books had built a major digitization program around authorized library holdings and search functions, a comparison the court discussed when assessing Anthropic’s own format conversion. An contemporaneous account of the court record notes that Google Books largely scanned borrowed library books non-destructively and returned them; Anthropic instead bought bulk volumes and used a faster destructive process.
The newly unsealed Project Panama documents add operational detail but not a definitive total. A report based on those filings says Anthropic spent tens of millions of dollars within about a year to acquire and process millions of books. The documents redact the ultimate book count and cost. A proposal from a vendor that worked with Anthropic contemplated converting 500,000 to 2 million books over six months, using a hydraulic cutting machine, production scanners and recycling pickup.
The point was not merely that destruction was faster. By purchasing a print copy, then replacing it with a single internal digital copy, Anthropic created the factual pattern on which the court found space-saving and searchability transformative. It did not receive a ruling that a company may make an unrestricted digital archive merely because a model may later use some books within it.
The economic trade-off was collection cost, not a settled license price
The record supports a concrete comparison but not a complete cost calculation. Anthropic spent “many millions of dollars” buying millions of often-used print books, according to the order; the later report characterizes the wider Project Panama effort as tens of millions. Neither retained source provides an unredacted final count and price, so a per-book cost cannot be responsibly calculated.
What is clear is that Anthropic considered licensing. Turvey contacted major publishers about training licenses in spring 2024, but the conversations lapsed, according to the order. It then contacted distributors and retailers to bulk-buy books for a “research library.” The court’s analysis did not turn on whether a publisher license would have been cheaper than purchasing used copies. It turned on the lawful acquisition of the particular copies and the one-for-one, non-distributed conversion that followed.
That leaves an important limit on the apparent business lesson. Documentation and access controls help establish what a company actually did, but they are not a standalone safe harbor. The order contrasted Google’s authorized holdings, technical safeguards and legal agreements with Anthropic’s uncontrolled pirated library; it did not hold that internal controls alone cure an unauthorized acquisition.
What the settlement leaves unresolved
The court had set the pirate-library claims, including damages and possible willfulness, for trial. That trial did not happen. Anthropic paid $1.5 billion to settle in August without admitting wrongdoing, the unsealed-filings report says; it estimated eligible authors could receive about $3,000 per title. Anthropic deputy general counsel Aparna Sridhar said the settlement concerned how some materials were acquired, while the June 2025 ruling on training remains intact.
The settlement therefore resolves the authors’ monetary dispute but does not expand the holding. Anthropic has said in legal filings that it did not train a revenue-generating commercial AI model using LibGen data and did not use Pirate Library Mirror to train any complete model. The order nonetheless treated the retained pirate archive as a distinct use and left open claims involving copies made from the library for other purposes.
The next cases will need to supply evidence this record did not: whether a developer can trace each dataset from acquisition through training and retention; whether its systems expose copyrighted expression in outputs; and whether an archive is actually restricted to the use offered as its justification. Those facts—not the spectacle of a cutting machine—will determine how far this ruling travels.
Starbucks has retired Automated Counting, a computer-vision tool for counting milk and beverage components in North American coffeehouses. The company says it is standardizing inventory counts; reported miscounts leave the economics and operational value of the replacement unproven as Starbucks pursues more frequent replenishment.
HSBC plans to open a Global AI Centre of Excellence in Singapore in the second half of 2026 and hire more than 100 specialists. The bank is adding a deployment hub to an existing Google Cloud programme; the unresolved question is whether it can substantiate the promised gains while retaining human accountability in wealth, payments and treasury work.
Microsoft says MAI-Cyber-1-Flash lets its MDASH vulnerability system route routine work to a smaller in-house model and reserve GPT-5.4 for difficult cases. The resulting benchmark and cost claims apply to the combined system, leaving customer deployment, pricing and remediation outcomes to test in Project Perception’s preview.
Meta has raised its planned investment in the Hyperion campus in Richland Parish from $10 billion to more than $50 billion. Its promises of jobs, local investment and lower power bills now depend on tax terms and a 20-year utility arrangement whose long-run protections are still being tested.
AT&T has signed an agreement to explore D-Wave's annealing technology in more network operations after an early optimization workload dropped from about an hour to under 15 seconds. The commercial question is whether that company-reported result can deliver better decisions at the cost, scale and reliability of a live carrier network.
Yuyuantantian, an account linked to China Central Television, has proposed matching access to AI-model capabilities and risks rather than treating models as simply open or closed. The commentary is not a regulation, and separate reported discussions of limits on overseas access remain preliminary.
NVIDIA has made an undisclosed investment in Safe Superintelligence and promised access to its Vera Rubin platform that the companies say can increase the lab’s compute tenfold. The partnership gives Ilya Sutskever’s closed AI lab another hardware route, but leaves its research, capacity allocation and commercial terms unexamined in public.
NVIDIA has formed the Open Secure AI Alliance with more than 30 technology and security partners, using Hugging Face’s response to an AI-enabled intrusion as its case for deployable open-weight defensive tools. Members have identified work on identity, scanning and patching, but the coalition has not yet set out a common process for evaluating, disclosing or remediating failures.
CXMT's 466% Shanghai debut made the DRAM maker China's most valuable listed company, but the price was set with only 6.73% of its enlarged share capital freely tradable and before the company has proved how far its new capital can take it against tooling restrictions and established memory rivals.
Investigations and a TikTok search study found synthetic or impersonated clinicians reaching large audiences with dubious health claims. The evidence is not a platform-wide measure, but it shows why an AI label alone may not tell viewers whether a medical recommendation is credible.
Alphabet, Amazon and Meta have disclosed 2026 capital-spending plans that add to $520B to $550B at their stated ranges. The total conveys the scale of their infrastructure push, but it is not a comparable measure of AI-only spending or of the obligations each company is taking on.
Anthropic has made Claude Fable 5 generally available with classifiers that redirect certain risky requests to Claude Opus 4.8, while the same underlying model, Claude Mythos 5, remains available with cyber safeguards lifted only to approved organizations. The models share a listed token price, making vetting, monitoring and capacity—not a purchasable premium tier—the immediate constraint on less-restricted access.
Nvidia is reportedly discussing a roughly $250 billion financing backstop that could help OpenAI lease a proposed 10GW Ohio data-center campus. But no deal has been announced, other AI companies are pursuing the federally controlled power, and the reported 800MW first phase would be only a fraction of the planned buildout.
MSI's China distributor sheet shows week-over-week increases of roughly 8% to 20% for listed RTX 50 cards. Colorful's sheet shows substantial premiums to China MSRP, but without earlier prices it cannot establish how much—or how recently—those models rose.
Moonshot AI has released the weights and technical report for Kimi K3, its 2.8T-parameter mixture-of-experts model. The release gives researchers and operators a deployable checkpoint and more of its training infrastructure, but Moonshot still recommends large supernodes and its performance claims are not a uniform, like-for-like comparison.
xAI has added a built-in /deep-research command to Grok Build, turning a coding agent's existing delegation controls into a bounded, source-backed reporting process. The documentation is unusually specific about workflow limits and failure reporting, but it does not establish that the process improves accuracy, speed or cost.
An independently reported account says President Donald Trump posted AI-generated images of an Iranian tanker seizure, a burning tanker and a strike on Kharg Island. The posts did not announce a new operation, but they followed reported U.S. strikes on the island’s military assets and renewed attention to its oil terminal.
Alphabet’s proposed $80 billion equity package is partly tied to AI infrastructure, but its $40 billion at-the-market program is primarily intended to cover employee-equity tax obligations. Together with Microsoft’s $190 billion 2026 capital-expenditure forecast, the disclosures show a more complicated shift in how hyperscalers fund constrained compute capacity.
Intel is putting €5 billion into equipment and connections at its operating Leixlip campus to increase output of Intel 3-based Xeon processors. The investment strengthens an existing European manufacturing base, but Intel has not disclosed its added wafer capacity, the split between its own products and foundry work, or an outside customer.