FINDING · EVALUATION

WFP models trained and tested on scripted browser traffic achieve near-perfect accuracy (NetCLR: 98.9%, TF: 98.3%, Var-CNN: 98.9%), but accuracy collapses to below 10% — with F1-scores near zero (TF: 0.034, Var-CNN: 0.022, ARES: 0.025) — when trained on scripted traffic and tested on real human traffic (CrossDomain). This structural mismatch means scripted datasets overestimate the real-world WFP threat by more than 80 percentage points.

From 2026-song-redefining-website-fingerprintingRedefining Website Fingerprinting Attacks with Multi-Agent LLMs · §5.3 / Table 1 · 2026 · PoPETs 2026

Implications

Tags

censors
generic
techniques
website-fingerprintml-classifiertraffic-shape

Extracted by claude-sonnet-4-6 — review before relying.