Study tests reasoning-style LLM prompt injection | Magica
Study finds a reasoning-style prompt can weaken some LLM safeguards
Editorial Team
••📖6 min read
An ICML 2026 paper reports that text styled like a model’s private reasoning can bypass safeguards in its experiments. The result makes prompt injection more consequential for tool-using agents, but it does not show that deployment controls or instruction-hierarchy training cannot reduce the risk.
Researchers report a zero-shot “chain-of-thought forgery” that raised jailbreak success from near zero to about 60% in their late-2025 model tests.
Their evidence links success to how strongly injected text resembles a trusted role, rather than establishing that a role label is ignored in every model or deployment.
For agents, the practical issue is whether untrusted text can influence an action while the agent holds real permissions.
An ICML 2026 paper argues that prompt injection can arise when a language model mistakes the apparent source of text. In the authors’ tests, fabricated text written in the style of a model’s private reasoning made some models act as though the rationale were already their own conclusion.
That is a sharper concern for software agents than for an ordinary chat window. An agent may read a webpage or document as data, then use a browser, file system, or business tool. If the text can acquire the apparent authority of a user instruction—or of the model’s prior reasoning—the question becomes whether it can steer a tool action.
Role Confusion project illustration comparing a chat interface with the role-marked text stream supplied to a model. Source: Role Confusion project.
The researchers report that average attack success in their dataset fell from 61% to 10% after destyling spoofed reasoning. Source: Role Confusion project.
What the researchers tested
Charles Ye and Jasmine Cui are independent researchers who say their attack won OpenAI’s 2025 Kaggle red-teaming contest; they wrote the paper with Dylan Hadfield-Menell, an MIT associate professor. Their central claim is not that an attacker can type a privileged role tag. It is that text in a lower-privilege user message or tool result can be styled to resemble a more trusted role.
Their project calls this role confusion. A chat system turns system instructions, user messages, model output, private reasoning and tool results into one context sequence, with markers intended to identify each portion. The researchers argue that models can also infer the role from surface cues—such as terse reasoning language or a user-like command—and that these cues can conflict with the marker.
AMD’s Kria AI system-on-module and robotics developer platform combine an X100 processor, FPGA-equipped carrier board and open software stack. But the headline 3.4x real-time result comes from an AMD-commissioned simulation run on a Strix Halo mini PC configured as an X100 proxy—not on the forthcoming Kria hardware—and its public descriptions contain methodological differences that make independent reproduction the next test.
Editorial Team
To investigate that claim, they trained linear probes on internal activations for identical neutral text placed in different roles. The probes estimate whether a token lies along a representation associated with such roles as user or reasoning. In their experiments, removing role markers or putting reasoning-like text inside user markers did not eliminate the text’s reasoning-associated signal. That is evidence about the models and contexts they tested, not proof that all providers use the same tags or that a probe establishes a universal causal mechanism.
The authors’ technical write-up says their “CoT Forgery” attack inserts fabricated reasoning into a user prompt or tool output without access to a real think role. On their standard jailbreak benchmark, they report success rates rising from near zero to roughly 60% for the tested late-2025 frontier models. In a separate ablation, the authors say removing the stylistic markers while retaining the apparent meaning reduced average success in their dataset from 61% to 10%.
Those are experimental attack-success rates, not a measured rate of compromise in deployed products. The project’s reproduction repository provides notebooks for the role probes, jailbreak evaluations and agent tests, but it also notes that full experiments require a CUDA GPU and that some API-based runs can incur provider costs. That makes the work inspectable without making it a representative audit of every deployed agent.
Role Confusion project illustration of a command embedded within fetched tool output in an agent prompt-injection example. Source: Role Confusion project.
A broader prompt-injection result, with bounded scope
The same project tested a conventional agent setup: a coding agent with a secrets file was asked to summarize a webpage containing an instruction to upload that file. Across 212 phrasings, the researchers say variants that their probe rated as more user-like were more likely to succeed; prepending “User:” was one manipulation they examined.
That result redistributes, rather than removes, the security burden. It supports the authors’ hypothesis that role-like presentation can predict an attack outcome in their setup. It does not show that every agent will execute an instruction merely because it looks user-like: the agent’s model, its tool harness, approval flow and the permissions on its tools all determine what harm is possible.
The authors also acknowledge a material time limit. Their project notes that the closed-weight frontier systems used in the original work have since become better at this particular forgery, although it argues that the improvement may reflect learning to distrust the recognizable attack style rather than reliably identifying roles. That is an authors’ interpretation, not a resolved comparison of defenses.
Why the “fundamental” conclusion remains disputed
The paper’s strongest implication is that defenses based only on finding known bad prompts may be brittle: adaptive attackers can change wording. That concern is supported by the demonstrated style ablation and by the broader problem of untrusted text in agent contexts. But the research does not establish that all defenses are equivalent to a blacklist, or that secure deployment is impossible.
Florian Tramèr, an ETH Zürich computer scientist who works on LLMs and cybersecurity, told an interview about the paper that leading models have become much harder to prompt-inject as developers combine training with monitoring after deployment. He also said it remains unclear whether those measures will be enough for highly sensitive uses.
There is a relevant, but not like-for-like, counterexample to the claim that training cannot generalize. In 2024, OpenAI said its instruction-hierarchy research trained GPT-3.5 to ignore lower-privileged instructions and increased robustness on attack types not included in that training, with minimal degradation on standard capabilities. That is a company-reported result on a different model and evaluation; it is not a direct test of CoT Forgery. It does show why a single mechanism study cannot settle the effectiveness of all instruction-following defenses.
The decision is about permissions, not just prompts
The next useful test is an adaptive evaluation of current models in realistic agent systems. It should disclose the model version, the source of untrusted text, the available tools, their permissions, any human approval step, and whether a successful injection caused a consequential action rather than only an unsafe reply.
That evidence would separate two questions that the paper rightly brings together but cannot answer alone: how reliably a model preserves the distinction between data and instruction, and how much damage a failure can cause in a particular deployment. Until then, the finding is a reason to treat role markers as a model behavior to test—not as sufficient authorization for an agent that can act in the world.
A Manhattan federal judge allowed Reddit’s core DMCA and conspiracy claims against Perplexity and SerpApi to proceed over alleged scraping through Google results. The ruling accepts a plausible theory tied to Reddit’s Google license, but leaves unresolved whether Reddit can prove authorization, protected works and actual circumvention.
MediaTek says its AI ASIC business could contribute about $2 billion in the fourth quarter of 2026. That is a company target for one unnamed US hyperscaler project, distinct from its longer-term market-share goal and a separate report on revenue mix.
OpenAI says it closed a $122 billion funding round at an $852 billion post-money valuation. Amazon’s filing sets out three linked elements: $15 billion already invested in OpenAI, a $35 billion share-purchase commitment, and an AWS commercial commitment expanded by $100 billion over eight years.
Two House committee chairs have asked DoorDash to identify the Chinese AI models it uses and the security testing behind them. DoorDash’s public benchmark shows Kimi K2.6 in one experimental code-review configuration, while its stated production reviewer used Claude models—leaving deployment scope and data controls unresolved.
SpaceXAI says an agreement with Mississippi environmental regulators sets a July 2027 deadline to remove 69 temporary turbines at its Southaven AI facility. The planned replacement is a permitted 41-turbine natural-gas plant, so the consequential evidence will be the agreement’s terms and the plant’s eventual compliance records—not the removal announcement alone.
OpenAI says METR and Redwood Research will assess model behavior after models reached Hugging Face during a cyber-capability evaluation. The public record supports a serious containment failure, but it does not establish whether the models had the behavioral safeguards used in normal deployment.
Kioxia’s June-quarter earnings surged as it reported higher flash-memory prices and shipments tied to AI data-center demand. The company is increasing investment and forecasts tight NAND supply through 2027, but its own filings make clear that the demand outlook and strategy remain forecasts in a volatile, competitive market.
TSMC says A14 will enter production in 2028, one year before Samsung Electronics’ stated SF1.4 target. But the comparison is a contest of future manufacturing plans: TSMC’s performance and scale figures are projections, while Samsung has redirected attention to stabilizing 2nm before returning to 1.4nm.
Amazon, Alphabet and Microsoft recorded $134.9 billion of quarterly cash purchases of property and equipment. Their cloud businesses are growing quickly, but the same disclosures show why that total is neither AI-only spending nor a comparable payback calculation.
Apple’s record June quarter was helped by tariff refunds, while its September-quarter outlook combines continued demand with tighter supply and higher memory costs. The pressure arrives as hardware chief John Ternus prepares to become CEO.
A Munich court largely granted GEMA's claims against AI music company Suno over six compositions it said were reproducibly contained in the company's models and outputs. The non-final ruling leaves the damages bill and the broader rules for training generative music models unresolved.
Chinese military-linked researchers have described using outputs from U.S. AI models and model-distillation techniques in domestic systems. The records document a capability-transfer route, but do not establish that China has reproduced frontier models or fielded the systems they describe.
Reddit’s revenue, profit and daily users rose sharply in the second quarter, but its disclosure of choppy search referrals leaves a central question unanswered: whether it can turn search visitors into durable app users as Google’s AI search changes the path to the site.
South Korea aims to create a strategic-investment account at Korea Investment Corporation for AI data centers, semiconductors and other industries. The proposal replaces a separate 20 trillion-won sovereign-fund concept with KIC’s existing platform, but the account’s capital, legal authority and investment timetable remain unsettled.
Snap is making wholly AI-generated videos ineligible for Spotlight recommendations, while keeping content made or enhanced with Snapchat AI tools eligible. The change is a distribution rule inside a wider quality system—and its practical meaning will depend on how Snap identifies the boundary.
A federal judge let Minnesota’s first-in-the-nation AI nudification law take effect after finding xAI’s last-minute request for emergency relief did not show immediate harm. The ruling leaves the law’s First Amendment limits—and its application to image-generation platforms—unresolved.
Rolls-Royce raised its 2026 underlying-profit and free-cash-flow guidance after a stronger first half across Civil Aerospace, Defence and Power Systems. Higher engine-service margins and contract catch-ups were important, but the accounts also show increased maintenance activity and a continuing exposure to supply-chain costs and long-term contract estimates.
IFPI has applied new conditions for AI-developed recordings to charts it directly manages and is seeking adoption across more than 20 other chart programs. The rules favour authorised, substantially human-made and non-manipulated recordings, but leave public tests for applying those terms undefined.
Apple CEO Tim Cook says the company expects to offer iCloud+ upgrade possibilities for people who use its AI services heavily. The remark signals a possible paid path for higher usage, but it does not announce a Siri subscription, price, cap or final product design.