Phishing is extensively studied, yet it keeps reaching all-time highs. Attacks increasingly rely on modern web design patterns to look legitimate and, at the same time, to evade phishing detectors and security crawlers. A crawler that loads only the first page misses most of what a victim actually experiences.
A crawler that plays along
We built an intelligent crawler that combines browser automation, machine learning, and visual analysis to simulate the interactions phishing sites expect from their victims:
- A field parser and field classifier find input fields and work out what each one asks for (email, name, card number, and so on).
- A page interactor fills the fields with appropriate fake data and moves through multi-page flows.
- A trait analyzer then studies what each site did: CAPTCHAs, multi-factor authentication prompts, multi-stage data collection, and the campaigns the sites belong to.
Key findings
- Across 51,859 phishing sites we identified 8,472 campaigns.
- 45% of sites collected data across multiple pages, mimicking the experience of legitimate sites.
- 5.6% of sites used click-through gating to hide data-collection pages behind preliminary interactions.
- Phishing sites often impersonate a brand without closely copying its design, embed modern user-verification systems such as CAPTCHAs, and sometimes end by reassuring victims that their data is safe.
Why it matters
These behaviors directly undermine detectors that assume phishing pages are single-page clones of a brand’s login form. Understanding phishing from the user’s perspective points toward more robust detection. It also shaped our later work on CAPTCHA-bypass services and in-browser defenses.