Future Monitor 한국어

Signal Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation

Summary

Researchers Natalie Isak and Matthew Dressman published a paper on arXiv (cs.CY) on August 3, 2026, describing a new AI safety evasion technique called cross-session goal decomposition, where an attacker splits a harmful goal into innocuous-looking units executed across separate, stateless agentic AI sessions. They demonstrate that this technique can elicit more harmful capability than equivalent single-session or multi-turn attacks, exploiting the fact that while an AI agent has no memory between sessions, the attacker does. The paper identifies a structural gap in existing AI misuse detection research, which they say focuses almost exclusively on single-turn or single-session multi-turn threat models. To address this, the authors propose 'Magnet,' a detection method that aggregates capability evidence across sessions and time at a higher-level correlator such as a user ID, rather than inspecting each session individually. The system is designed to pull together incriminating evidence scattered across many individually harmless-looking sessions into a single evidence bundle a detector can act on.

Classification

Main topicAI & Computing
Region menusGlobal
Impactscope:global
Time horizon0-3 years (2026-08-05)
Last updated2026-08-05T02:11:17.902688+00:00

Evidence 1

Part of trends 0

No objects.

Directly linked issues 0

No objects.

Public id: fm-66e9d1bb5e56