A recent publication by the SPY Lab ask how good large language models are at re-identifying anonymous online users. This is a collaboration between Daniel Paleka, Joshua Swanson, Michael Aerni and Florian Tramèr from ETH Zurich’s Department of Computer Science (D-INFK), together with Simon Lermen (MATS) and Nicholas Carlini (Anthropic).
Although deanonymization is a decades-old idea, pseudonymous accounts on Reddit, Hacker News and similar forums have stayed largely safe — such attacks traditionally needed structured data or hours of manual investigation.
The main insight of the paper is that LLMs change the economics of deanonymization. It is previously known that LLMs can extract identity-relevant features from online comments. Our pipeline extends this to search for candidate matches across platforms, and reasons over them to confirm a match. LLM-based methods massively outperform classical ones, linking accounts to real people for less than $5 each. The takeaway: the more you post, the easier you are to unmask.
The findings received wide press coverage: the ETH D-INFK spotlight, plus Bloomberg, The Guardian, The Verge, El País, CyberScoop and Bruce Schneier internationally, and 20 Minuten, Tages-Anzeiger and Le Temps in Switzerland.
