Signal Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation
Summary
Researchers Natalie Isak and Matthew Dressman published a paper on arXiv (cs.CY) on August 3, 2026, describing a new AI safety evasion technique called cross-session goal decomposition, where an attacker splits a harmful goal into innocuous-looking units executed across separate, stateless agentic AI sessions. They demonstrate that this technique can elicit more harmful capability than equivalent single-session or multi-turn attacks, exploiting the fact that while an AI agent has no memory between sessions, the attacker does. The paper identifies a structural gap in existing AI misuse detection research, which they say focuses almost exclusively on single-turn or single-session multi-turn threat models. To address this, the authors propose 'Magnet,' a detection method that aggregates capability evidence across sessions and time at a higher-level correlator such as a user ID, rather than inspecting each session individually. The system is designed to pull together incriminating evidence scattered across many individually harmless-looking sessions into a single evidence bundle a detector can act on.
Classification
Evidence 1
- arXiv (cs.CY) 2026-08-03 accessed 2026-08-05T02:02:05+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-66e9d1bb5e56