Signal When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
Summary
Yongli Xiang and colleagues published a paper on arXiv (cs.CY) on August 4, 2026, introducing AntiSkillBench, an end-to-end benchmark for evaluating privacy and impersonation risks in 'persona skills' that distill personal interaction histories into portable, executable artifacts for AI agents. The benchmark comprises a dataset of 7,500 persona-grounded dialogue traces built from 50 behaviorally rich profiles, an evaluation suite measuring skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies, and a defense evaluation covering four configurations spanning online and post-hoc interventions. Experiments across three frontier agents found that persona-skill risks persist regardless of agent backbone or distillation protocol, extending beyond explicit personal attributes to communication style and personality traits. Existing defenses showed limited and distillation-dependent effectiveness, failing to generalize across different risk types and distillation strategies. The authors present AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.
Classification
Evidence 1
- arXiv (cs.CY) 2026-08-04 accessed 2026-08-05T02:34:16+00:00
Part of trends 0
No objects.
Directly linked issues 0
No objects.
Public id: fm-510b37dfcf9f