Goodfire’s Silico packages mechanistic-interpretability experiments into software for organizations that can access a model’s internal parameters. The launch broadens a specialist workflow, but its price, operating cost and ability to deliver reliable control outside demonstrations remain unproven.
Goodfire, a San Francisco AI research company founded in June 2024, has released Silico, software that plans, runs and monitors experiments on how an AI model works internally. The product turns mechanistic interpretability—a research approach that maps a neural network’s neurons and their connections during a task—into a workflow for research and engineering teams, as described in a company profile.
The significance is distribution, not a solution to AI’s black-box problem. Goodfire says Silico can help inspect a model, investigate failures and run post-training experiments; the launch report says the company is packaging techniques that had been used in-house. That could make a scarce specialty more usable outside frontier labs. It does not show that a model’s behavior can be completely specified or controlled.

Goodfire’s demonstration of a DINOv2-base atom explorer. Source: Goodfire.
Silico uses agents to develop an experimental plan, run work in parallel and monitor progress, according to Goodfire’s product page. Its listed functions include visualizing architecture, training sparse autoencoders and probes, testing causal hypotheses, tracing regressions, and comparing checkpoints while running supervised fine-tuning, direct preference optimization and reinforcement-learning experiments.
Those are distinct stages of a model-development workflow. The page describes post-training methods rather than a claim that Silico builds a foundation model from scratch. It also promotes separate uses for life sciences, robotics and vision, and language models.
The intended workflow is to move from an observed failure to a testable internal hypothesis. A team might inspect which features or pathways correlate with the behavior, make a targeted parameter or data intervention, and then measure the result. That is more specific than prompt-level testing, but it is still experimental work; the product page’s promise to coordinate jobs does not establish that every internal correlation is a dependable cause.
Silico cannot inspect a typical closed AI service from the outside. The reporting says users need access to a model’s inner workings, which excludes probing closed systems such as ChatGPT or Gemini in this way but allows work on many open-source models.
That condition redistributes, rather than removes, the expertise problem. The most plausible users are teams training a model, adapting an open model, or operating a specialized model with control over its parameters and training process. Organizations that only call an external model API do not obtain that access through Silico.
It also puts Goodfire in a competitive field broader than a single product launch. The same report identifies Anthropic, OpenAI and Google DeepMind as other organizations working on mechanistic interpretability. A separate profile names Conjecture and Guide Labs among emerging companies building explainable foundation models. Silico’s differentiation is therefore its attempt to offer an off-the-shelf experimental environment, not ownership of interpretability research itself.

Goodfire’s illustration for the Silico for life sciences use case. Source: Goodfire.
Goodfire has recruited about 50 interpretability researchers from labs including OpenAI and Google DeepMind and raised more than $200 million, the profile reports; it says the company was valued at $1.25 billion and that Anthropic was its first startup investment. Those details matter because interpretability requires specialized researchers as well as substantial experimentation capacity. They also underscore the commercial pressure to make the methods usable by teams that do not employ their own interpretability group.
The same profile says Mayo Clinic, Rakuten and Microsoft are using Goodfire’s product. It describes Mayo using it to check a DNA model that studies rare genetic mutations, while Prima Mente used Goodfire’s tools during work on a model intended to predict whether a patient would develop Alzheimer’s. These are reported uses, not independently reported performance validations of Silico.
The operating economics remain unavailable. Goodfire declined to provide pricing details; the launch coverage says fees are set case by case according to customer requirements. The product page offers a macOS download or discussions about a team’s infrastructure, but supplies no public rate card. That leaves buyers unable to compare the cost of a Silico-assisted investigation with staffing specialists or using internal tooling.
Goodfire’s examples support the narrower claim that a model’s output can sometimes be changed through targeted internal work. In one, the company found a neuron in the open-source model Qwen 3 associated with the trolley problem; activating it made responses frame outputs as explicit moral dilemmas, according to the report. In another company demonstration, raising neurons associated with transparency and disclosure changed an answer from refusing disclosure to favoring it in nine out of 10 trials.
Neither result supplies a denominator beyond that demonstration, a measure of side effects, or evidence that the intervention carries across models, retraining runs or production settings. Leonard Bereska, a University of Amsterdam researcher who has worked on mechanistic interpretability, described the tool as potentially useful but warned that it adds precision to existing trial and error rather than making model development as principled as conventional engineering.
Silico’s next test is not whether it can locate striking individual neurons. It is whether customers can reproduce useful interventions on the model classes and workloads they actually operate—and whether the gains survive retraining without unacceptable trade-offs elsewhere.
The decision-relevant evidence would include independent results on diagnostic success rates, unintended behavior changes, the model access required, experiment compute costs and the price paid by comparable customers. Until then, Goodfire has made a specialized research workflow easier to acquire for suitably equipped teams; it has not established a general route from seeing inside a model to reliably engineering its behavior.
Get concise AI news and useful context from the Magica team.
Read the newsletterSamsung has shown an HBM5 mock-up built around a 2nm base die and a new heat-management design, but it has not provided a production timetable or comparable performance data. Its nearer challenge is turning HBM4E samples into qualified, high-volume supply as rivals ship competing HBM products and expand capacity.
Anthropic is reported to have agreed to spend $10 billion over six years on Volta computing capacity. The disclosed Norwegian lease linked to Volta clarifies the opportunity—and the construction, financing and delivery conditions that still stand between a headline commitment and usable AI compute.
SoftBank has drawn $20 billion of a $40 billion bridge facility for its OpenAI investment and expects another $10 billion draw in October. Its Aug. 6 earnings release should show whether the group can turn a stated menu of asset-backed financing, bonds and asset sales into longer-term funding.
A newly published review of Phoenix Ikner’s yearlong ChatGPT record describes crisis-support replies alongside detailed answers in the hours before the 2025 Florida State University shooting. The record raises hard questions about OpenAI’s safeguards and account-level risk detection, but it does not establish causation or reveal what the company’s internal systems knew.
Patrick Steven Yaroch, a former FBI supervisory special agent who worked counterintelligence, is accused of using bureau-held wallet information to transfer cryptocurrency into accounts he controlled. The charging records make ChatGPT conversations part of the government’s evidence, but not the alleged means of moving the money.
Caterpillar raised its 2026 sales outlook after quarterly revenue first topped $20 billion. A surge in power-generation demand points to the data-center buildout, but stronger construction sales, pricing and a growing dealer inventory make the company’s record backlog an incomplete measure of that boom.
Cloudflare is taking reservations for cloudflare.pay handles and has outlined stablecoin wallets that would let account owners constrain AI-agent spending. The wallet is not yet available, and Mastercard’s broader machine-payments program shows that identity, limits and settlement are already a competitive infrastructure race.
Apple is seeking expedited discovery before a ruling on its preliminary-injunction request against OpenAI and former Apple hardware employees. The filing adds allegations involving other former employees, while OpenAI calls the case factually wrong; the court must decide whether Apple’s proposed evidence-gathering is justified and proportionate.
HappyRobot's $150 million Series C values the enterprise AI-agent company at $1.2 billion and funds expansion beyond logistics. The company says its deployment model can automate complex operational work, but the disclosures leave its reliability, pricing and cross-sector repeatability unmeasured.
Bending Spoons has agreed to acquire Airtable in an all-cash transaction with a $1.285 billion enterprise value. Against Airtable’s company-reported $480 million in annual recurring revenue, that is about 2.7 times ARR—a more useful, though still limited, view of the transaction than a simple comparison with the company’s 2021 funding valuation.
Wisconsin Assembly member and gubernatorial candidate Francesca Hong proposes a statewide pause on new AI data-center construction while lawmakers write rules on utilities, environmental protections, subsidies, labor and local approval. The proposal puts a clear choice before the state: set those terms before more projects advance, or keep local dealmaking open while regulations are debated.
Microsoft paid more than $20 million to 562 security researchers in the year ended June 30, a company record. But the total also reflects a broader eligibility policy, a live hacking event and a higher volume of AI-assisted research, leaving the program's operational performance unmeasured.
Arista Networks reported $3.036 billion in Q2 2026 revenue and forecast about $3.3 billion for Q3. The filing confirms rapid product-led growth and strong operating margins, but it does not break out sales from AI fabrics, customers or other end markets—leaving the breadth and durability of the AI thesis unproven.
AMD reported $11.536 billion in second-quarter revenue as data-center sales more than doubled, but investors’ focus has shifted to whether its Helios rack-scale systems can turn that demand into shipments and cash generation on the timetable AMD projects.
The Justice Department says OpenAI and its subsidiary Statsig will pay $3.2 million to resolve allegations that they disadvantaged U.S. workers during recruitment for fewer than 10 visa-linked jobs. The agreement combines penalties, a back-pay fund and required changes to job postings and applications.
A 64-year-old Pauillac photographer has been placed under formal investigation and remanded in custody after a partial device review found several thousand suspected child sexual abuse files. The central unanswered question is which children, if any, can be matched to the alleged AI-altered material.
A reported White House framework would offer U.S. officials a 30-day pre-release review of some closed frontier AI models while exempting open models. The government has not published the rules, leaving its thresholds, participants and consequences unclear.
The Ninth Circuit lifted a preliminary injunction that had barred Perplexity’s Comet assistant from using Amazon.com, holding that Amazon was unlikely to show Perplexity itself accessed its computers. The decision is limited to the current record and leaves Amazon’s lawsuit, its terms of service and wider questions about AI-agent liability unresolved.
An Aug. 1 Arena snapshot put Alibaba's Qwen3.8 Max fourth in front-end web development at 1,668 points. The preliminary, task-specific result is a useful signal, but it cannot by itself verify Alibaba's broader claim for Qwen3.8—especially because the source record separately names a July Preview release and a Qwen3.8 Max leaderboard entry.
SpaceX’s first public quarterly report showed 92% revenue growth and a smaller-than-expected loss, but six-month cash flows show that its expanding launch, connectivity and AI plans still depend far more on financing than operations.