skip to content
The Weighted Average

Wire

Role confusion makes prompt injection a 61% problem

Role-confusion researchers cut chain-of-thought-forgery attack success from 61% to 10% merely by stripping stylistic cues from the injected text. Their ICML research write-up also reports testing 212 role-signaling variations and argues that models often infer who is speaking from prose style rather than trust the surrounding role tags; MIT Technology Review’s account puts the result in the context of adaptive attacks against newer frontier systems. For teams adopting continuous agent red-teaming, the practical implication is governance, not a cleverer prompt: keep authorization, least privilege, and consequential approvals outside the model.