Prompt Injection Works Because the Model Can't Tell Who Is Talking
New research traces prompt injection to a single mechanism: LLMs decide who is speaking from writing style, not from role tags. Forged reasoning takes attacks from near-zero to 60% success. Remove the style, and it collapses to 10%.
5 min read