Skip to main content
Scientific Reports logoLink to Scientific Reports
. 2026 May 21;16:23234. doi: 10.1038/s41598-026-51242-2

Detection and mitigation of abusive web traffic using convolutional neural networks

Farkhanda Athar 1, Akmal Shahbaz 2, Mansoor Qadir 3, Anandhavalli Muniasamy 4, Sana Munir 1, Hend Khalid Alkahtani 5,
PMCID: PMC13402595  PMID: 42168277

Abstract

Illicit websites depend upon abusive Traffic Distribution Systems (TDSs) to generate user traffic for malicious. Traffic Distribution Systems are the intermediate websites that redirect the HTTP traffic from online advertisements. However, such systems also started to promote abusive activities such as phishing, scams, ad frauds, malicious downloads, and social engineering attacks. In this study, we present Online Abusive Traffic Finder (OATF), an enhanced web security protection system designed to investigate and evaluate abusive TDSs and their associated threats. A total of 10,746 webpages were collected over a one-month period (May 15, 2024–June 14, 2024) from four diverse traffic sources, including advertisement-based URL shortening services, typosquatting websites, unlicensed online pharmacy sites, and the PhishTank dataset. We use these sources due to their diverse nature to redirect users toward abusive and malicious sites. During data collection process, we collect destination web pages screenshots, browser, and content logs. We semi automatically label collected pages and use labeled data to automatically examine page content to understand the threats from these traffic sources. To protect users from abusive TDSs, a Convolutional neural network (CNN) based classifier is integrated as a supporting component for automated detection of abusive webpages using visual features. The CNN model achieved the highest accuracy (91.92%) within the proposed framework. The proposed approach provides deeper insights into the operational behavior of abusive traffic ecosystems and contributes toward improving web security against evolving malicious distribution strategies.

Subject terms: Engineering, Mathematics and computing

Introduction

Illicit websites increasingly rely on abusive Traffic Distribution Systems (TDSs)1 to generate user traffic for malicious activites. TDSs act as intermediate websites that redirect HTTP traffic from online advertisements towards final destination pages. Although TDSs were originally designed to optimize traffic distribution in advertising campaigns, they are increasingly exploited to promote abusive activities in order to maximize user traffic24. For example, TDSs automatically redirect visitors toward malicious destination pages without their consent, such as phishing59, scams7,1, malicious downloads10, ad frauds11,12, botnets13, and social engineering attacks14. Although the research community has conducted numerous experiments on abusive practices such as redirection from fake news domains2,11,15,16, attacks using URL-shortener services6, domain squatting abuse17,18, DNS poisoning19, cryptocurrency giveaway scams13, etc., there remains a need to investigate how TDSs operate within advertising networks and advertiser ecosystems to facilitate the distribution of abusive content.

In parallel, the technological revolution has made the Internet dominant channel for online advertising, including promotions and marketing2026. As compared to traditional advertisement approaches, online advertisement is very economical and convenient and has become a billion dollar industry27. However, the same infrastructure has also enabled the proliferation of malicious and abusive activities such as scamming7, click frauds10,26, ad fraud11, phishing6,8,9,28,29, malicious downloads10,30, botnets13, social engineering attacks3133, etc. Attackers can easily create accounts on major advertising platforms such as DoubleClick10,22,34, allowing malicious messages to be distributed to a large population. Consequently, cybercriminals increasingly exploit web advertising platforms as a low-cost and highly effective medium for distributing malicious advertisements and conducting fraudulent activities3538.

Displaying ads through Traffic Distribution Systems (TDSs) consists of three major parties: ad publishers (who publish ads on their web pages), ad advertisers (who create ads), and targeted audiences (the visitors who visit publishers’ page)36,38,39. TDSs store the statistics of each visitor (date and time, redirection info, IP address, browser type, operating systems), settings and rules to redirect each user through malicious advertisements. The redirected pages are mostly the landing pages associated with the affiliate program40. The affiliate marketers select different landing pages of different designs and themes to redirect maximum traffic. The incoming traffic from abusive TDSs comes from two main sources: hacked websites that redirect the user from legitimate sites to abusive TDSs and malicious ads that secretly redirect the user towards abusive TDS38,41.

Traffic Distribution Systems (TDSs) operate under two models: Pay Per Click (PPC) and Pay Per Redirect (PPR)1. In the PPC model, ad publishers are paid according to clicks on their websites22,24,25. As the clicks demand participation from active users, a more intrusive approach is used to auto redirects the user to the targeted page. This method is known as Pay Per Redirect because the ad hosting websites receive money according to the number of redirections on targeted websites. Abusive TDSs use both PPC and PPR models4,42,43. However, Pay Per Redirect (PPR) is more destructive than Pay Per Click (PPC) and more actively involved in malicious and abusive activities44.

Despite extensive prior research on abusive advertising ecosystems and Traffic Distribution Systems (TDSs), existing studies primarily focus on isolated abuse scenarios, such as technical support scams, survey scams, domain squatting, or fake news redirection4553. As these studies focus on single traffic source which limits their ability to capture the broader working behavior of abusive TDSs. In real world environment, Traffic Distribution systems consider several factors for redirection such as device type, browser configuration, geographic location, and crawler detection mechanisms. Moreover, recent advancements in machine learning and deep learning have improved the detection of malicious URLs and phishing websites. However, most of existing techniques rely primarily on URL features, textual content, or static webpage characteristics5458,86. Consequently, they often fail to capture dynamic web-level behaviors such as cloaking techniques, user differentiation strategies, and multi-stage redirection chains used by modern abusive Traffic Distribution Systems.

Recent data-driven detection and localization frameworks25,5962 have demonstrated strong performance in identifying sophisticated attacks within controlled cyber-physical environments. Although these approaches are effective for domain-specific settings, they are not designed for the highly dynamic nature of open web ecosystems. In web environments, adversaries frequently exploit adaptive user differentiation, IP-based cloaking, and multi-profile redirection techniques to evade detection mechanisms. As a result, there remains a lack of a unified web measurement framework capable of jointly analyzing user differentiation, cloaking behaviors, and large-scale abusive redirection patterns across diverse traffic sources. Consequently, there remains a lack of a comprehensive measurement and detection framework capable of analyzing large-scale abusive redirection behaviors across heterogeneous traffic sources and diverse user personas.

To investigate abusive Traffic Distribution Systems (TDSs) and their associated threats, a framework Online Abusive Traffic Finder (OATF) is designed. The proposed framework collects large-scale web measurement data and employs CNN based classification model to automatically detect abusive webpages using webpage screenshots. Data was collected from four different traffic sources; typosquatting domains, PhishTank dataset, e-pharmacy websites, and advertisement-based URL shortening services. These diverse traffic sources allow us to analyze how advertising campaigns and redirection vary based on user characteristics such as browser type, operating system, installed Java plugins, and geographic location (IP address). While previous studies focus was to examine abusive advertising within specific ecosystems such as technical support scams63, survey scams64, social engineering attacks32,65,66, fake news websites67, and domain squatting abuse40,51,68,69; there remains a lack of systematic analysis of how TDSs dynamically adapt their redirection strategies across different user profiles, including desktop users, mobile devices, and web crawlers. To address this limitation, our framework adopts broader analysis strategies such as IP cloaking and user differentiation, extending the work of1.

Using OATF, a total of 10,746 pages were scraped over a one-month period (May 15, 2024–June 14, 2024). Screenshots, browser logs, and content logs were collected from the destination web pages. To examine the complexity of abusive Traffic Distribution Systems (TDSs), multiple user profiles were employed, including desktop users, emulated mobile users, and web crawlers. The analysis shows that ad-based URL shortening services primarily redirect users to the same malicious destination pages and exhibit the highest number of malicious content. Furthermore, our findings indicate that attackers primarily target desktop users through scam survey campaigns, such as prize announcements requiring users to complete surveys or download applications via social media platforms. In contrast, deceptive landing pages were rarely observed in the pharmacy dataset, while keyword-stuffed and illicit pages were more prevalent in typosquatting datasets. Emulated mobile users exhibited different redirection patterns compared to desktop users, and GoogleBot was redirected to entirely different destination pages. These observations demonstrate that abusive TDSs do not exclusively target desktop users, but dynamically adapt their attacks across multiple access channels, including mobile devices and web crawlers.

To support the analysis of abusive traffic distribution systems (TDSs), an automated detection component for identifying abusive web pages is introduced within the proposed framework. This component is based on a multilayer Convolutional Neural Network (CNN), consisting of stacked Conv2D and MaxPooling2D layers followed by a fully connected layer for feature extraction and classification. The CNN model operates on webpage screenshots and achieved an accuracy of 91.92%, proving its effectiveness as a supporting tool within the framework.

The key contributions of this work includes:

  • To investigate the complexity of abusive Traffic Distribution Systems (TDSs), the Online Abusive Traffic Finder (OATF) was employed to visit URLs using different user profiles, including desktop, mobile, and web crawlers.

  • To evaluate the effect of proxies on abusive TDSs, each URL was accessed both with and without proxies. Screenshots, browser logs, and content logs of the destination pages were collected and semi-automatically labeled to categorize different types of attacks.

  • To support abusive OATF an automated web page detection model based on a multilayer Convolutional Neural Network (CNN) was developed using the collected webpage screenshots. The CNN model achieved a classification accuracy of 91.92%.

The organization of this paper is as follows. “Related work” section presents an overview of related studies. “Methodology” section describes the methodology employed in this research. “Results” section reports the results and provides an analysis of the findings. “Challenges and future research directions” section discusses the limitations of the study and offers suggestions for future work. Finally, “Conclusion” secton presents the concluding remarks.

Related work

This section highlight prior work relevant to our study. It describes the related work in terms of malicious advertisement, illicit traffic sources, ad blockers and cloaking techniques.

Malicious advertisement

Usha et al.70 proposed a third-party server to validate advertisements before loading them into different applications. The advertisement libraries worked within an independent sandbox with their environments and authorization. Moreover, the system guarantees that none of the ad material extracted sensitive information while accessing the host device. Subramani et al.44 proposed a system called PushAdMiner for automated collection and analysis of Web Push Notifications from various publishing platforms The system discovered the WPNs from different publisher sites and identifies the malicious ad campaigns. The results demonstrate that 51% of ads were malicious, and traditional filters and ad blockers and filters were uneffective against them. Singh et al.,71 proposed multilayer Convolutional Neural Network (CNN) offers a robust solution by capturing complex features in URLs through multiple layers, significantly improving detection accuracy. This method provides a more reliable approach as compared to previous studies to enhance user security in the ever-evolving digital landscape.

Chen and Frere,11 proposed a system to uncover redirection ad networks in real time. The experiments verified that 91% of entry points were expired and re-registered. Nowroozi et al.4 introduced a machine-learning technique to detect fraudulent URLs. Using different statistical properties, 12 different datasets were used for detection, classification and prediction. They achieved 99.63% accuracy with 0.0037 false positive rate. Moreover, they used an unsupervised data clustering approach for visual analysis. The analysis was performed to measure how adversarial attacks affect classification vulnerability.

Birthriya et al.59 introduced a detection framework for phishing websites that combines a dual-layer CNN with a GRU enhanced by an attention mechanism. By incorporating lexical NLP features, the model is able to capture both structural and contextual patterns of phishing URLs, improving its ability to detect evolving attacks. Nevertheless, the framework depends on large labeled datasets and requires further evaluation against adversarial attack scenarios. Zhou et al.60 proposed a detection framework for malicious URLs through logic constraint neural networks and hybrid spatial sequences. The framework used both spatial patterns and sequential dependencies of URLs that ensure higher semantic consistency in classification. Table 1 presents a detailed comparison of different works related to malicious advertisement, their contributions, limitations and future directions.

Table 1.

Comparison of recent work related to Malicious advertisement.

Papers Year Data sources Evaluation metrics Methodology Limitations Future work
Srinivasan et al.63 2018 Domain name registrar data Typosquatting domains features Domain registrar data Stateless crawler Improve clustering efficiency
Usha et al.70 2019 Android Apps Detection rate, false positive rate Feature selection Not applicable to all types of malware Improve detection accuracy
Subramani et al.44 2020 Push notification ads Detection rate, false positive rate Machine learning classification Unable to block attacks Automatic ad detector
Singh et al.44 2021 Malicious URLs Detection accuracy Multilayer Convolutional Neural Network (CNN) Limited to URL-based attacks Evaluate against dynamic URL patterns
Chen and Frere2 2021 Web traffic data Click through rate Statistical modeling Limited to specific websites Explore user interventions
Nowroozi et al.4 2022 Malicious URLs Attack success rate Machine learning malicious ads detection Not generalized to all detection models Vulnerability comparison
Birthriya et al.59 2024 ISCXURL-2016 dataset K fold cross validation word embedding Required labeled data Extend adversarial validation
Zhou et al.60 2025 Hybrid Spatial Sequence Attention Captures both spatial Sequential dependencies High computational cost Real time scalability

Abusive traffic sources

Page et al.72 examined clicks through URL shortener analytics to compare malware and phishing attacks. Overall, about half of the malware attacks have persisted for several years, while the remaining half occurred within the past three months. Tahir et al.73 introduced a mechanism to better investigate why some URLs become extra dangerous due to typing mistakes. They investigated the relationship between human hand anatomy, keyboard layouts, and typing errors across various URL datasets. The cross-validation against DNS and URL datasets determines that it is an effective and meaningful defensive mechanism.

Zhang et al.74 systematically investigated click interception tactics on different websites. They introduced a framework called browser-based to explore the click behavior. The results show that most users are redirected towards malicious content such as scamware using click interception. Oest et al.75 proposed a framework to identify phishing websites. The findings indicates that extended empirical measurements effectively reveal weaknesses in anti-phishing systems. The results further highlight that improved mobile security protection and evidence-based protocols are effective to counter phishing attacks

Hao et al.76 measured the perceptual hashing robustness and security applications rely on it. The results show that attacks are highly successful on Bing and TinEye while moderately successful on Yandex, and Google. Chen et al.2 performed a statistical analysis of web traffic data from The Gateway Pundit (TGP). Their geo-location study revealed that TGP is particularly popular in regions that predominantly supported Trump in the 2020 election. Their unique web traffic insights and dataset help researchers to measure the social media platforms, and fake news URLs more efficiently.

Zeng et al.40 performed a large scale investigation about the adversaries’ URL redirections to distribute malicious content and security checks in practice. For this purpose, they designed a framework to find out domains used for malicious redirections. They provided detailed information about the malicious redirections but still, there are certain limitations. Srinivasan et al.63 analyzed the performance of Technical Support Scams (TSS) that redirect users during online searches. Their investigation of search and advertisement abuse provided new insights into TSS strategies, helping to uncover previously unknown methods used to propagate these scams. Table 2 presents the comparison of different methodologies, data sources, techniques adopted to perform research, their contributions, limitations and future directions.

Table 2.

Comparison of recent works related to abusive traffic sources.

Papers Year Collected data Methodology Targeted attacks Limitations Future work
Tahir et al.73 2018 DNS traffic Empirical study Typosquatting Spelling mistakes issues Typosquatting on mobile devices
Zhang et al.74 2019 Alexa top websites Empirical study Click Interception Visited only main pages Defense against click interception
Oest et al.75 2020 Images, Perceptual hashes Experimental study Perceptual Hashing attacks Appearance amendments of phishing sites Implementation at large scale
Hao et al.76 2021 Browser based experiments Framework development Phishing Limited evaluation scale More countermeasure experiments
Chen et al.,2 2022 Web traffic analysis Data analysis News consumption Datasets permission issues Collaboration with Industry partners
Zeng et al.40 2022 client-side and server-side redirections Empirical study Regional Restrictions Limited websites Relationship investigation between intermediaries
Garg et al.77 2025 11M redirecting URIs Web redirection patterns URI terminations vs. errors No identification of malicious entities Security considerations
Tian et al.78 2025 Publicly available datasets 2016-2020 Repository analysis Detection techniques Old detection methods Detection framework

Machine learning in web security

Amouri et al.79 studied intrusion detection techniques in both Industrial IoT (IIoT) and Internet of Things (IoT) systems using optimization techniques. They investigated the swarm based optimization algorithms to enhance feature selection and detection performance across complex, high-dimensional network traffic. Dobrojevic et al.80 applied optimized gradient boosting classifiers to detect cyberbullying, sexism, and harassment in social media text datasets. The results indicate that metaheuristic tuning significantly improves classification accuracy for abusive content. Zivkovic et al.81 used NLP models combined with machine learning to identify insider threats through sentiment and behavioral analysis in organizational communication data. Their findings indicates that optimized models can more effectively capture subtle behavioral indicators of malicious intent.

Rao et al.82 proposed a machine learning framework with arithmetic optimization algorithms to detect misinformation and fake news at large scale online information streams. Their study highlights that metaheuristic based optimization improves feature selection and classification performance in the context of online content verification. Savanovic et al.83 introduced a hybrid CNN–XGBoost model optimized with a modified sine cosine algorithm. This framework uses convolutional layers for feature extraction along with ensemble classifiers, achieving high detection accuracy on the TON_IoT dataset.

Table 3 shows a detailed comparison of different works related to machine learning applications in web security.

Table 3.

Recent machine learning approaches in cybersecurity applications.

Paper Year Data sources Evaluation metrics Methodology Limitations Future work
Amouri et al.79 2024 IoT network traffic datasets Accuracy, Detection Rate, False Positive Rate Metaheuristic-based feature selection, machine learning classifiers Sensitive feature engineering Development of lightweight IDS models
Dobrojevic et al.80 2024 Social media text datasets Precision, Recall, F1-score NLP feature extraction, gradient boosting classifiers Dependence on labeled training datasets Integration of multimodal social media signals
Rao et al.82 2024 Online news datasets Accuracy, Precision, Recall Machine learning classifiers, arithmetic optimization algorithm Sensitive to data imbalance Hybrid deep learning for misinformation detection
Savanovic et al.83 2025 TON_IoT Dataset Accuracy, Precision, Recall, F1-score Hybrid CNN feature extraction, XGBoost classifier, cosine algorithm Computational complexity Explore additional real-world datasets
Zivkovic et al.81 2026 Organizational communication datasets Accuracy, AUC, F1-score NLP-based sentiment analysis Limited organizational datasets Cross-domain threat detection

Exploration of ad networks

Kharraz et al.64 introduced a system called SUREVYLANCE to auto-detect survey scams performed through machine learning techniques. The analysis revealed that Alex’s top 30K websites are mostly involved in survey scams. These websites exposed the users to different security issues such as identity fraud, potentially unwanted programs, deceptive downloads, malware and malicious extensions. Hashmi et al.27 introduced a security metric to measure the changes in advertisements and ad-tracking domains that are blocked by well-known blacklist services. The analysis showed that ads and ad tracking domains vary from time to time, and certain ad tracking domains are more informed about ads and ad tracking, but their change rate is slower than other blacklists. The results show that top Alexa websites in Canada, UK and the US have the most number of ads and ad-blocking websites and have the highest proactive scores.

Vadrevu & Perdisci32 introduced a system to track and discover Social Engineering Attack Campaigns launched through Malicious Advertisement (SEACMA). The results show that SEACMA ads use different methods to evade ad blockers and URL blacklists. Therefore, the results could be used to enhance defensive mechanisms against malicious ads and social engineering attacks. Gao et al.72 performed an experimental study to assess user concerns about ad performance cost. They conducted a correlation analysis to explore user complaints about higher performance costs. They conclude that the performance costs with the version of the ad are higher than those ads that have no version.

Zhao et al.84 conducted experiments on Forbes, a leading US media website. They categorized ad blockers into two groups to assess their impact, and the results indicated that the use of paywall strategies negatively affects user engagement.

Szurdi et al.1 developed the Observatory of Dynamic Illicit ad-Networks (ODIN) to analyze user differentiation, cloaking techniques, and business integration by simultaneously analyzing four different traffic sources. The results show that these traffic sources redirect users towards the same malicious advertisement or domain names. Moreover, the cloaking techniques decreased half of the malicious pages they observed, and the most popular blacklist services were unable to block them.

Samarasinghe and Mannan85 presented an automated crawler to collect and analyze data from the top 1,00,000 malicious domains. They used a headless Chromium browser to identify the cloaking behaviour of content. The results indicate that the current blacklist tools are ineffective against malicious cloaked websites. Todri86 investigated the impacts of ad blocking on online searching and purchasing. The analysis of the consumer-level panel data set revealed that online purchasing is significantly impacted by ad blocking. They examined that ad blockers can minimize spending on brand pages. The results show unintended consequences as they minimize user search activities across different channels. Papadogiannakis et al.67 developed an ad detection technology for fake news websites identification and the intermediary entities involved to facilitate those ad revenues. The results show that fake news sites are part of entertainment websites, politics, and businesses.

Ahmad et al.87 ad networks and found that digital advertising continues to promote misinformation by algorithmic placement, even when advertisers are unaware. It shows that consumer backlash occurs when companies learn their ads appeared on misinformation sites, and supports low-cost, transparency-based interventions. Papadogiannakis et al.88 analyze the impact of the EU’s 2022 Strengthened Code of Practice on ad-based misinformation monetization, finding little reduction in ad relationships for highly trafficked misinformation sites—even though some low-traffic sites saw declines. Garg et al.77 conducted an extensive analysis of around 11 million unique redirecting URIs, uncovering typical and non-canonical redirection behaviors—including canonical HTTP-to-HTTPS transitions and deceptive “sink” URI clusters—revealing significant risks to usability, SEO, and digital preservation. Tian et al.78 systematically reviewed approaches for detecting malicious URLs, introducing a modality-based taxonomy (URL, HTML, Visual, etc.), and curating datasets and implementations to promote reproducibility in the field. Table 4 shows a detailed comparison of different works related to ad-blocking software and browsers.

Table 4.

Comparison of recent work related to ad-networks.

Papers Year Data sources Methodology Evaluation metrics Limitations Future work
Kharraz et al.64 2019 Online survey platforms Machine learning models Precision, Recall, F1-score Supervised Dependent selection Extension to other scam types
Vadrevu & Perdisci32 2019 SE attack campaigns Detection and tracking system Detection rate, False positive rate Evasion from defense Integration with other security systems
Hashmi et al.27 2019 Ad-blocking blacklists Longitudinal study Changes in ad-blocking blacklists over time Encounter errors Varied frequencies Analysis of other ad-blocking techniques
Zhao et al.84 2020 Browser extensions and web pages Measurement and analysis User engagement metrics Limited websites Analysis of user attitudes towards ad-blockers
Gao et al.72 2021 Mobile apps User study User engagement metrics Experimented limited apps Analysis of other types of in-app ads
Szurdi et al.43 2021 Malicious websites Experimentation and analysis Detection rate, False positive rate Visited limited traffic sources Deploy crawlers from multi vantage points
Samarasinghe & Mannan85 2021 Web traffic data Measurement and analysis Abusive traffic metrics Dynamically lower bound Different search engines cloaking
Todri86 2022 Online consumer behavior Measurement and analysis Online consumer behavior metrics Analyzed only consumers and brands Consumer level analysis
Papadogiannakis et al.67 2022 Fake news sites Survey Ad revenue sources, Profit margins Limited to online consumer behavior Identify business relationship
Ahmad et al.87 2024 5,000 sites, 42K advertisers, 9M ads Data analysis, consumer experiment Ad prevalence, consumer demand shifts Only US focused Global scaling
Papadogiannakis et al.88 2024 Misinformation site dataset Temporal Analysis Persistence of ad relationships Short dataset Network specific response.

Multi vantage point web measurement

Mi et al.89 explored Residential IP proxy as a service prevision. They used servers and the clients for RESIP to discover 6 million RESIP IPs across 52K+ Internet Service Providers (ISPs) across 230+ countries. They further fingerprinted collected results using new profiling systems. The results shows that compromised hosts, including smart devices, mostly serve as proxies. They also discovered many Potentially Unwanted Programs (PUPs) including illegal promotions, malware hosting, phishing, and others are provided through top IT companies. Their research also proved as a first step to exploring new Internet services and efficiently contributing to controlling security risks. Dao et al.90 investigated CNAME cloaking used for web tracking. They crawled the top Alexa 300,000 websites and studied CNAME cloaking. The results show that 1,762 websites consist of CNAME cloaking on shopping and business websites in the US.

Zhang et al.91 introduced the CrawlPhish framework, which automatically detects and classifies client-side cloaking employed by known phishing websites. The result shows several websites hijack user clicks with the collaboration of third-party cookies. Zeng et al.40 studied domain-squatting abuse. They selected 786 most queried domains and hunted them against ISP-level DNS traffic. They found that although most squatting domains are related to typo-squatting, combo-squatting can also attract more traffic. Teoh et al.92 introduces PhishDecloaker, a detection framework targeting phishing sites that employ CAPTCHA-based cloaking. It addresses both server-side and client-side cloaking by forcing the execution of JavaScript and simulating human interactions to reveal hidden phishing content that evades traditional detection methods. Nakano et al.93 introduced PhishParrot, a novel crawling system that leverages Large Language Models (LLMs) to dynamically simulate user environments and circumvent client-side cloaking. It constructs optimal browsing profiles that mimic targeted users, significantly enhancing the detection of cloaked phishing attacks. Table 5 presents a detailed comparison of these studies.

Table 5.

Comparison of recent works related to multi vantage point web measurement.

Papers Year Data sources Methodology Limitations Future work
Mi et al.89 2019 Residential IP proxies Residential IP Proxy Measurement inaccuracies Detailed analysis of IP proxies
Dao et al.90 2020 Alexa Top 1M domains CNAME cloaking based tracking Limited coverage of non HTTP traffic CNAME analysis in different contexts
Zhang et al.91 2021 Crowdsourcing and phishing kits Client side cloaking measurement Data collection legal and ethical limitations Examine only one type of malicious redirection
Zeng et al.40 2022 Large scale HTTPS traces Malicious redirection analysis Experimented only correlated domains Malicious redirections and their intermediaries
Teoh et al.92 2024 Phishing websites Detection of masking techniques through runtime behavior Focused only CAPTCHA technique Expansion of cloaking techniques
Nakano et al.93 2025 Phishing sites Usage of LLM to bypass crawling Heavier computational overhead Environment optimization

Although the previous studies worked on detection methods such as URL pattern analysis, traffic flow monitoring, supervised learning models, and distributed crawling from multiple vantage points. However, these traditional methods often fail to capture evasive behaviors such as IP-based cloaking and multi-stage redirection chains. Because these methods are dependent on URL-based features, including lexical patterns, domain characteristics, and hosting infrastructure properties. In contrast, the Online Abusive Traffic Finder (OATF) framework provides a comprehensive approach by integrating multiple traffic sources and simulating diverse end user profiles. Additionally, a CNN-based classifier is employed as a supporting component to analyze webpage screenshots and capture visual patterns associated with abusive content. This combination enables the detection of attacks that intentionally disguise malicious behavior to evade traditional URL-based methods.6

Table 6.

Comparison of OATF framework with representative existing detection approaches.

Paper Year Data source Data set type Technique Features Limitations
Singh et al.71 2021 URL Dataset Static Dataset Multilayer CNN on URL features Lexical, web scraped Limited real-world evaluation
Nowroozi et al.4 2022 6 benign & 6 malicious URLS datasets Static Datasets Traditional ML (RF, XGBoost, AdaBoost) Lexical, Web-scraped Vulnerable to adversarial attacks
Teoh et al.92 2024 CAPTCHA phishing sites Specific Phishing sites Hybrid Vision, Interaction Models Visual, CAPTCHA interaction Limited to phishing websites
OATF (Proposed) 2025 Typosquatting, URL shortening, PhishTank, online pharmacy sites Real-world multi-source traffic ML, CNN (supporting), Data-driven framework HTTP logs, screenshots, browser data Computational cost

Methodology

This section presents the Online Abusive Traffic Finder (OATF) framework as illustrated in Fig. 1. To analyze the abusive behavior of TDSs, the framework is designed to collect and analyze the traffic from four different sources with user agents variations and IP based cloaking.

Fig. 1.

Fig. 1

Overview of online abusive traffic finder (OATF).

Data collection

This section describes the data collection process. Data was collected from four different types of traffic sources: unregulated pharmacy sites, typosquatting domains, ad-driven URL shorteners, and PhishTank sites. A scheduler was employed to sequence the URLs and prevent detection by Traffic Distribution Systems (TDSs) triggered by multiple requests from the same IP address. To examine how different user types are targeted by various attacks, each URL was accessed using five distinct user agents: three desktop users, GoogleBot, and an emulated mobile device. Using OATF, a total of 10,746 pages were scraped over a one-month period (May 15, 2024–June 14, 2024)

Target creation and selection

Typosquatting dataset For typosquatting analysis, variants of Alexa’s top domains were generated to create a comprehensive set of typo domains. From this set, 200 domains were randomly selected and visited using five different user profiles. Experiments were conducted only on domain names with valid NS and A records to ensure authenticity40.

URL shortening services dataset For the URL shortening dataset, URLs were collected from ten URL shortening services that redirect to Alexa top-ranked domains. For each experimental run, 200 URLs were selected.

Illicit pharmacies dataset To determine strong coverage of pharmacy related domain names, a set of pharmacy-related search keywords was used with the Google Search API. For each experimental run, 200 URLs were generated and selected for analysis.

PhishTank dataset To analyze IP-based cloaking behavior, phishing URLs were collected from the PhishTank website and accessed using five different user agents. For each experimental run, 200 URLs were selected for analysis.

A total of 200 URLs were selected from each dataset to balance the computational cost and capture diverse redirection, labeling efforts and content variation within each traffic source. Moreover, this approach allows extending the framework easily at large scale without any additional changes to measurement infrastructure or detection pipeline.

User emulation

The main objective is to examine how Traffic Distribution Systems (TDSs) differentiate between different user types. To achieve this, 5 different user profiles were simulated, and each URL was accessed using five distinct profiles. A fully featured headless Chromium browser94, controlled via Selenium95,96, was used for all profiles. Table 7 summarizes the user profiles employed in the experiments. All experiments were conducted using Linux Chrome with different user agents, except for mobile emulation, where Android Chrome was used.

Table 7.

Overview of user profiles used for experiments.

Profile Category Browser/Agents Mobile Simulation Source/Referrer Proxy Usage
Sdandard (No Proxy) Windows Chrome No Google No
Referrer Windows Chrome No Google Yes
Baseline Vanilla Windows Chrome No None Yes
Mobile Emulation Android Chrome Yes Google Yes
Google Crawler Google Bot No None Yes

Desktop users To emulate the behavior of desktop users, a standard desktop crawler was employed. The crawler operated using Google Chrome with a default Windows user agent to simulate a typical browsing environment. To mitigate cloaking techniques that rely on referrer headers, a customized crawler configuration was implemented in which the HTTP referrer was set to https://google.com during the preliminary crawling process. Additionally, each URL was accessed both with and without using a proxy to assess the impact of proxy on the crawling results.

Mobile phone users To investigate how Traffic Distribution Systems (TDSs) target the mobile users differently as compared to desktop users, a simulated mobile environment was employed. Instead of using real phone due to some limitations, Chrome mobile emulation features were used by the corrected window size, pixel ratio, and a user agent according Android phones.

Google bot Malicious websites evade their activities or show Search Engine Optimization (SEO) pages while visiting using Google crawlers. To determine behavior of TDSs while using a search engine crawler, user agent was set to Google Crawler.

These five profiles were selected to reflect the most common web access scenarios. The three desktop profiles represent the different desktop browsing behavior during human interaction. The mobile user simulates mobile device -specific behavior and screen configuration. While the GoogleBot stimulates the legitimate search engine crawling that enables the system to differentiate between benign and malicious traffic. Together, these provide a realistic analysis of traffic distribution systems.

Cloaking identification and prevention

To investigate IP address based cloaking, each URL was accessed through two different proxy configurations. In the first setup, a single IP address was used to simulate a consistent user. In the second setup, IP addresses were rotated across 240 different proxies to represent multiple users and examine variations in the delivered content. A more significant feature of our system is to determine adversarial behaviour of Traffic Distribution Systems (TDSs) without detecting cloaking.These experiments were conducted using four distinct user agents, and the objectives were addressed through the following approaches.

Self-rate limiting Some traffic sources, especially typosquatting domains, conceal their IP addresses when multiple requests originate from the same client. To counter this form of IP-based cloaking, the scheduler is designed to prioritize crawling these URLs in a shorter interval and classify URLs from a common origin as related. Anti-browser Fingerprinting The most common methods to check automation is to detect user agent, and handle cookies. We modify our browser properties by adding extensions, changing windows size, and adding a default language to communicate browser fingerprinting.

Proxy detection avoidance Proxies were employed as the primary method for rotating IP addresses. To bypass proxy detection, the forwarded_for HTTP header was utilized. Additionally, a dataset was collected from pages accessed without proxies to investigate whether some attackers deploy more advanced proxy techniques. To examine user behavior–based detection, mouse movements were also emulated during browsing.

Due to the presence of potentially harmful URLs and copyright restrictions associated with collected webpages, the complete dataset cannot be publicly released. However, the dataset structure, labeling taxonomy, and crawling methodology are described in detail to facilitate reproducibility of the study. Researchers interested in accessing the dataset for academic purposes may contact the authors.

Data labeling

This section presents an overview of the labels and label categories used in the dataset to classify the different types of abuse encountered by OATF. The labels are categorized into five primary categories: illicit, suspicious, error, benign, and malicious.

Error labels Errors in the dataset are categorized based on their origin. The errors are labeled as “crawl” errors that occur due to infrastructure. Mostly, the crawling errors occur due to the proxy not working. When URLs are blocked during crawling, labeled as “blocked”. We label all other errors as “errors”.

Benign labels Web pages are labeled as “empty” if they consist of a small amount of content or no content. Pages are labeled as “parked” if the page is the HTTP server default page and is underdeveloped, under construction, consists of ads, or tries to sell domain names. Webpages were labeled as “pharmacy” if pharamacy related sites does not consists of any compromised data. Moreover, the pages that does not fit into any benign category are labeled as “original content”.

Illicit labels Pharmacy-related websites are classified as illicit pharmacies if the online pharmacies are compromised for black hat search engine optimization. While visiting these pages using Google bots, they deliver pages full of keywords to attract more visitors. We label pages as “affiliate abuse” if the sites consist of abusive affiliate content that automatically redirects users towards advertisers. In some cases, OATF redirects toward adult content and is labeled as “adult”. Suspicious Labels Pages are labeled as “downloads” or “surveys” that involve suspicious content related to downloadable files or survey forms, respectively. However, when the pages engage in suspicious activities, such as the empty page consisting of notification permission, they are labeled as “other” because their intentions cannot be determined.

Malicious labels When the pages consist of deceptive downloads or surveys, they are labeled as “deceptive downloads” and “deceptive surveys”. Deceptive download pages try to engage users to download files by giving notifications that the Flash player is out of date or consists of viruses or vulnerabilities. If the downloaded file consists of malware, then pages labeled as “malicious download”. We label pages as “scams” if the pages inform about the selection to receive free products or money after filling out specific surveys. Mostly these pages force users towards different actions such as filling out surveys, downloading a particular application, and asking for personal information. Pages are also labeled as “scams” that offer high-paying jobs and high-yield investments. Pages are labeled as “impersonating” if these pages consist of content to steal users’ sensitiveivee information and labeled as “tech scams” if the pages are scaring users to believe that their machine is infected or they are paying for technical support that is mandatory to clean their machines. Finally, pages are labeled as “malicious” if users receive an error message or deception warning.

To ensure consistent annotation and minimize the overlap between categories, operational definitions are defined in Table 9 for each label during the manual labeling process. These operational definitons provide a clear difference between conceptually related categories. Moreover, Table 8 summarizes label distributions across different traffic sources. This shows that advertising-driven URL shortening services contains maximum malicious pages, whereas pharmacy-related datasets have benign or illicit but non-malicious content.

Table 9.

Operational definitions of label categories and subclasses.

Category Label Description
Error Crawling error Webpages that could not be retrieved due to crawler failures, network issues, or incomplete page loading during the crawling process.
Blocked Pages where access was denied by the server, typically due to bot detection mechanisms or access restrictions imposed by the website.
Error page Webpages returning server-side errors such as HTTP 404 or 500 that prevent normal webpage content from being displayed.
Benign Parked Domains registered but not actively used, often containing placeholder advertisements or domain parking services.
Empty Webpages containing minimal or no meaningful content.
Original Legitimate websites containing genuine content without evidence of abusive or deceptive behavior.
Gambling Webpages offering online betting, casino services, or gambling-related activities.
Adult Websites containing explicit adult content but not necessarily engaging in malicious activities.
Online pharmacy Websites selling pharmaceutical products online that appear to operate as commercial pharmacy services.
Defensive Domains registered by organizations or individuals to protect brand names from misuse or impersonation.
Illicit Illicit pharmacy Webpages selling pharmaceutical products without proper authorization, regulation, or medical verification.
Illicit adult Adult content distributed through unauthorized, illegal, or deceptive channels.
Affiliate abuse Pages designed to manipulate affiliate marketing systems for fraudulent revenue generation.
Keyword stuffed Webpages containing excessive or manipulated keywords aimed at artificially improving search engine rankings.
Suspicious Download Pages offering downloadable content that may contain potentially harmful or misleading software but cannot be conclusively classified as malicious.
Survey Webpages attempting to lure users into completing surveys or providing personal information for questionable purposes.
Other Pages exhibiting suspicious characteristics that do not clearly fall into predefined subclasses.
Malicious Crypto scam Webpages attempting to defraud users through fake cryptocurrency investments or giveaway schemes.
Malicious download Pages distributing malware, trojans, or unwanted software disguised as legitimate downloads.
Impersonating Webpages mimicking legitimate organizations or brands in order to deceive users.
Phishing Pages that try to get sensitive information, such as login credentials or financial data.
Other scam Fraudulent webpages involved in online scams that do not fall into other predefined categories.
Tech support scam Pages falsely claiming technical issues and tricking users into paying for fake support services.
Other malicious Other forms of malicious activity that cannot be classified into the above subclasses.

Table 8.

Overview of data labels and label categories.

Label categories Labels
Error Crawling Error, Blocked, Error
Benign Parked, Empty, Original, Gambling, Adult, Online Pharmacy, Defensive
Illicit Illicit Pharmacy, Illicit Adult, Affiliate Abuse, Keyword Stuffed
Suspicious Download, Survey, Other
Malicious Crypto Scam, Malicious Download, Impersonating, Phishing, Other Scam, Tech Support Scam, Other Malicious

Clustering and dataset labeling

The collected pages were semi-automatically labeled using the labels described above. Pages were clustered using multiple approaches, including grouping by matching perceptual hashes and textual similarity. The k-Nearest Neighbour (KNN) algorithm was used for clustering. For each unique hash, a random selection of pages was manually labeled to ensure labeling accuracy. To ensure consistency, labels were propagated to every page that has an identical perceptual hash, and a subset of webpages was verified accordingly. For final validation, up to one hundred screenshots per label were manually reviewed. Only a small number of screenshots were found to be incorrectly labeled. These mislabeled pages typically corresponded to errors, blocked content, parked domains, or empty pages. However, this level of inaccuracy is considered acceptable, as our focus was not on distinguishing between underdeveloped and error pages. Figure 2 presents labeled images according to the different classes used in the dataset

Fig. 2.

Fig. 2

Labeled images corresponding to different classes.

To further ensure the reliability, the semi-automatic labeling process was double-checked through manual validation across all label categories. In cases where inconsistencies were observed between automated clustering results and visual page content, labels were corrected through consensus-based review. This validation process ensured that the labeled dataset maintained consistency and high quality for training and evaluating the CNN-based classifier.

Results

In this section, the defined labels are applied to determine which type of attacks user face. Then, we discuss the overlapping between different Traffic Distribution Systems (TDSs). Finally, we evaluate the performance of our trained classifier. Figures and tables presented in this section provide evidence of how abusive content varies as per traffic source and crawling profile.

Label analysis

This section presents the results analysis for each label class and provide detailed insights into how the content types vary according to user profiles. Table 10 and Table 11 summarize the number of labeled pages of each class.

Table 10.

Label classification per traffic source.

Types PhishTank Pharmacy Typosquatting URL Shortening All
Num. % Num. % Num. % Num. % Num. %
Error 311 10.4% 54 2.2% 310 10.6% 152 7.85% 827 8%
Benign 1712 57.4% 2089 85% 1559 52.63% 1038 53.7% 6,398 61.9%
Illicit 56 1.8% 131 5.33% 571 19.27% 0 0% 758 7.33%
Suspicious 253 8.48% 62 2.52% 283 9.55% 64 3.3% 662 6.4%
Malicious 651 21.8% 119 4.84% 239 8.06% 680 35.3% 1689 16.34%
All 2983 2455 2962 1934 10,334

Table 11.

Label classification per crawl profile.

Types Vanilla Desktop No Proxy Android Google Bot
Num. % Num. % Num. % Num. % Num. %
Error 141 7.1% 148 4.7% 135 6.43% 149 7.37% 307 15.42%
Bnign 1272 64.1% 2538 80% 1335 63.7% 1269 62.8% 1251 62.9%
Illicit 150 7.5% 155 4.9% 203 9.6% 133 6.5% 117 5.8%
Suspicious 132 6.7% 118 3.7% 135 6.43% 138 6.8% 150 7.5%
Malicious 289 14.5% 195 6.18% 289 13.8% 332 16.42% 165 8.29%
All 1984 3154 2097 2021 1990

Phone vs desktop users

Figure 3 shows the distribution of labeled web pages according to crawling profiles and traffic sources. Our analysis reveals that attackers predominantly target desktop users with scam surveys, prize announcements, and impersonation campaigns, often propagated via social media or illicit websites. Desktop users are also more exposed to deceptive download pages and technical support scams. In contrast, mobile users encounter fewer surveys and technical support scams, indicating that attackers tailor content based on device type. As both mobile and desktop experiments were carried out simultaneously, so the results show that attackers target mobile and desktop users with different attacks.

Fig. 3.

Fig. 3

Label counts to present most well known labels in each profile types and traffic source.

Traffic sources per common malicious destination pages

Analysis of traffic sources demonstrates that advertising-driven URL shorteners redirect users to a large number of common malicious landing pages. The results show that different users are targeted with different malicious pages, surveys, technical scams and downloads. Moreover, several attack campaigns redirect users toward similar landing pages. In contrast, pharmacy-related redirections appear to be largely distinct from other traffic sources.

Malice in our datasets

Table 10 presents the results of the various categories used in data collection and analysis. The findings indicate that pharmaceutical-related queries exhibit behaviors distinct from other traffic sources. In the pharmacy dataset, malicious landing pages were rarely observed, whereas the typosquatting dataset primarily contained illicit and keyword-stuffed pages. Furthermore, OATF detected numerous malicious files while accessing pharmacy-related URLs. The results also varied across user profiles: emulated phone users displayed different browsing patterns compared to desktop users, and the GoogleBot crawler was redirected to entirely different pages.

Typosquatting domains mostly lead towards parked pages. Although, the typosquatters also engage the users in various malicious activities and affiliate abuse campaigns. Most of the suspicious and malicious content consists of downloads, surveys, technical support scams, Google searches, impersonating pages and many other scams. Several malicious web pages are specifically used for typosquatting pages including financial phishing pages, surveys without deceptions, and forced Google searches.

Cloaking and bot detection

To investigate cloaking behavior in abusive Traffic Distribution Systems (TDSs), the destination pages delivered to different user profiles such as desktop users, mobile emulation, and automated crawlers. Cloaking is identified when the same source URL provides different destination pages depending on the requesting user agent or environment.

To quantify this behavior cloaking rate is defined as the proportion of URLs that generate inconsistent redirection outcomes across different user profiles. Formally, the cloaking rate is defined as:

graphic file with name d33e2582.gif

where Inline graphic represents the number of URLs that redirect the traffic to different destinations using different profiles, and Inline graphic represents the total number of visited URLs.

The results indicate that when OATF acted as a GoogleBot crawler, it encountered fewer suspicious, illicit, and malicious pages as compared to other profiles. Specifically, automated crawler profiles detected approximately 8% fewer malicious pages than desktop and mobile users. This observation shows that few TDSs selectively modify redirection behavior to evade from detection.

Table 11 highlights how different crawling profiles lead towards different content. The results show when an OATF acted as a Google bot, it only observes a few types of suspicious, illicit and malicious content. This behavior shows how different Traffic Distribution Systems (TDSs) behave differently while visiting using an automated crawler. The analysis shows that automated crawlers are blocked 8% more than the other users. Interestingly, no evidence of cloaking were found on the URLs visited using proxies. Moreover, the “No proxy” profile results in the lowest error rate. This happened due to errors using proxies. Illicit pharmacy sites use HTTP referrer headers to cloak their malicious activities. Contrarily, setting referrer headers in other profiles has the opposite effect. In that profiles, it decreases the number of malicious pages we discover. Finally, some adversaries may employ time-based cloaking strategies, where malicious content is delivered only during specific time intervals in order to evade detection. Due to the limited duration of our measurement period, a detailed analysis of time-based cloaking was beyond the scope of this study. Investigating temporal cloaking behavior remains an important direction for future work.

IP-cloaking

The outcomes of experiments conducted using a single IP address were compared with those obtained using multiple IP addresses. The use of different IP addresses frequently directed us to distinct malicious landing pages. However, this approach also introduced errors and revealed suspicious or illicit pages. In several cases, abusive TDSs displayed error or benign pages instead of malicious ones, indicating cloaking behavior, and at times blocked our crawler altogether. Further it is observed that typosquatting sources were more likely to block our crawler when accessed from a single IP address, compared to pharmacy and URL-shortening datasets. Notably, the mobile crawler encountered fewer blocks than the desktop crawlers.

Data classification using CNN

During visits to the collected URLs, screenshots of destination web pages were captured and used as input images for a Convolutional Neural Network (CNN)-based classification model. CNN was originally designed to solve image recognition problems, now widely adopted in domains such as malicious web page detection due to their capability to auto learn spatial hierarchies of visual features. The most significant feature of the CNN algorithm is convolutional weight sharing, which minimizes the model complexity. We adopt CNN for malicious web page classifications due to the following two reasons.

  • As compared to the artificial feature based model, CNN presents great image classification performance at a large scale.

  • Although CNN takes a long time to train, its detection speed is very high.

Figure 4 represent the CNN architecture used in our model. The proposed CNN architecture is inspired by classical deep learning models including AlexNet, VGGNet, GoogLeNet, and ResNet.

Fig. 4.

Fig. 4

CNN model structure to detect malicious web pages.

The model consists of three convolutional layers followed by a fully connected layer. All input images were resized to 256 × 256 × 3. Each convolutional layer applies learnable filters with increasing channel depth (e.g., 32, 64) followed by non-linear activation using the Rectified Linear Unit (ReLU): The ReLU function is computed as follows.

graphic file with name d33e2636.gif 1

A fully connected layer is added following several convolutional and max-pooling layers. The convolutional and pooling layers serve as an automated feature extraction process for the input images. These extracted features are then supplied to fully connected layer, which performs the final classification.

The output prediction layer carries out the final classification by assigning each input image to its corresponding web page category. As this is a multiclass classification problem, the softmax activation function is used to compute the probabilities for each class. The softmax function is defined as follows:

graphic file with name d33e2650.gif 2

Moreover, we calculated separate losses of each class for multi-class classification and then summed them up. So, the cross-entropy can be calculated as follows.

graphic file with name d33e2655.gif 3

To address class imbalance, class weighting was applied during training, while no resampling techniques were used. The dataset was split into training, validation, and testing sets using an 70%, 15%, 15% strategy respectively. K-fold cross-validation was not applied due to computational constraints and is considered for future work. To reduce overfitting, dropout layers and batch normalization were incorporated into the model. Early stopping (patience = 5 epochs) was used based on validation loss. No data augmentation was applied in order to preserve original webpage layout characteristics. To evaluate training stability, the model was trained five independent times. The average accuracy was 91.92%, with a standard deviation of ±0.37%, indicating moderate variability across runs.

Measurement metrics

Confusion Matrix: A confusion matrix, as shown in Table 12 is a measurement matrix that consists of information about the predicted and actual classifications, and is used to evaluate the accuracy and effectiveness of a classification algorithm19.

Table 12.

Confusion matrix.

Actual class
Positive Negative
Positive True Positive (TP) False Positive (FP)
Negative False Negative (FN) True Negative (TN)

Accuracy defined as ratio of correctly predicted labels to the total number of predictions.

graphic file with name d33e2724.gif 4

Precision (P) defined as a ratio of true positive labels to predicted positive labels and determined by

graphic file with name d33e2731.gif 5

Recall (R) also called detection is defined as the ratio of true positive labels with total predictive labels and determined by

graphic file with name d33e2738.gif 6

False Positive rate (FP) is defined as ratio of negative label that are incorrectly classified as positive and determined by

graphic file with name d33e2746.gif 7

F1-score is defined as negative labels that are incorrectly labeled as positive and determined by

graphic file with name d33e2753.gif 8

Here Inline graphic has value from zero to infinity and used to control the weight assigned by Precision and True Positives.

Loss is computed on both the training and validation sets during the training of the CNN model. It represents how well or poorly the model performs at each optimization step. A lower loss value indicates better model performance, whereas a higher loss value reflects greater deviation from the expected output.

Model evaluation

The proposed CNN model was implemented using TensorFlow and Keras. Additional libraries used include NumPy, Pandas, Scikit-learn, Matplotlib, and OpenCV. The model was trained using the Adam optimizer with a learning rate of 0.001. The batch size was set to 32, and the network was trained for 50 epochs. Early stopping was employed based on validation loss with a patience of 5 epochs to prevent overfitting. Model checkpoints were used to retain the best-performing model during training. All experiments were conducted on a workstation with Intel Core i7 processor and 24 GB RAM. The architecture of the CNN used in our experiments is shown in Fig. 4. Convolutional kernel sizes, layer dimensions, pooling sizes, and other hyperparameters were configured according to task-specific requirements. The CNN model was trained using categorical cross-entropy as the loss function. The Adam optimizer was used to update model weights, and the learning rate was set to 0.001. The model was trained for 50 epochs with a batch size of 32.

The dataset of 4,544 images was randomly divided into training (70%), validation (15%), and testing (15%) sets while maintaining sufficient data for training. Figure 5 presents the average accuracy and loss during training and validation, demonstrating that the proposed CNN model effectively detects malicious web pages, achieving an accuracy of 91.92%. The model was trained five independent times with different initializations. Figure 6 presents precision, recall, and F1-scores of CNN model. These results indicate that the model performs consistently across most classes, although performance variation is observed among different categories, likely due to dataset complexity. Figure 7 compares the proposed approach with a previous study. The results indicate that the proposed framework achieves competitive performance. To further evaluate model reliability, standard evaluation metrics including accuracy, precision, recall, and F1-score were used. These metrics provide a comprehensive view of classification performance in a multi-class setting. The confusion matrix in Table 12 further illustrates the distribution of correct and incorrect predictions across classes.

Fig. 5.

Fig. 5

Accuracy and Loss of CNN model.

Fig. 6.

Fig. 6

Precision, recall, f1-score of our multi-class CNN model.

Fig. 7.

Fig. 7

Comparison of Model Accuracies.

Although the results show promising performance, it should be noted that statistical significance tests (e.g., paired t-test or Wilcoxon signed-rank test) were not conducted in this study. Therefore, the observed performance differences should be interpreted as indicative rather than statistically confirmed improvements. Additionally, ROC curves and AUC analysis were not included but are planned for future work to further assess the discriminative ability of the model. The dataset size of 4,544 images provides a reasonable basis for evaluating the proposed method; however, it may not fully capture all real-world variations of malicious web pages. Therefore, the results should be interpreted in the context of the dataset scope and experimental design. Overall, the experimental findings suggest that the proposed CNN-based approach, combined with the OATF framework, can effectively learn discriminative visual patterns for malicious web page detection, while further improvements and broader validation remain open for future research.

Although the primary focus of this study is the evaluation of the proposed OATF framework combined with a CNN-based detection model, it is relevant to discuss the roles of individual components in the pipeline. The framework integrates a multi-profile crawling strategy (desktop, mobile, and GoogleBot), visual screenshot-based representation of web pages, and a CNN-based classifier. The multi-profile crawling strategy is intended to improve content coverage and expose potential cloaking behavior that may not be observable through a single crawler configuration. Similarly, screenshot-based representations enable the model to utilize visual layout and structural features of web pages. However, a formal ablation study isolating each component was not conducted due to the complexity of the data collection pipeline. Therefore, the individual contribution of each component cannot be quantitatively determined in this study. This remains an important direction for future research.

Challenges and future research directions

With the rise of digital advancements, online abuse is increasing day by day. Law enforcement agencies, the security industry and academic researchers, all are assisting security practitioners to discover online questionable content. To mitigating different cloaking techniques and emulating different forms of factors (desktop and mobile), they are deploying automated crawling systems to detect abusive web activities. In this study, abusive domain registrations and Traffic Distribution Systems (TDSs) were investigated using a multi-dimensional measurement approach. This includes the analysis of typosquatting activities, malicious redirection behaviors, protocol usage, and user emulation to identify cloaking mechanisms. While these efforts provide valuable insights into abusive ecosystems, there remain significant opportunities to further understand the evolving nature of abusive ecosystems. Despites these contributions, the paper work still have some limitations

First, the data collection infrastructure limits the scale of active measurements. Due to this limitation, each webpage was visited five times to capture variations in cloaking behavior and to compare responses between desktop and mobile users. However, a more comprehensive analysis could incorporate additional user profiles, including variations in browser type, device configuration, geographic location, and browsing history.

Second, the completeness of the dataset remains limited due to the enormous scale of the web, which contains billions of webpages across numerous languages and millions of domain names. Despite advancements in the analysis of online crime infrastructures, there is still plenty of scope to further explore the the complexities of different abusive ecosystems. Therefore, we limit the resources and carefully selected the most common resources.

Third, the CNN-based classification component, while effective, has inherent limitations. The model achieved an accuracy of 91.92%, but misclassifications remain due to visual similarities between categories such as suspicious, illicit, and malicious webpages. Additionally, dynamic content variations across user profiles and minor inconsistencies in semi-automatic labeling may affect classification performance. Furthermore, the semi-automatic labeling process, although carefully validated, may introduce minor inconsistencies in certain cases such as parked pages, error pages, or empty domains.

Future research can extend this work in several directions. First, larger and more diverse datasets could be collected to enhance the generalization capability of deep learning models for abusive web traffic detection. Techniques such as k-fold cross-validation and expanded datasets may also be employed to further validate the robustness of the proposed model. Second, incorporating additional contextual features—such as HTML structure, network behavior, and URL-based attributes—alongside visual screenshot analysis has the potential to significantly improve detection accuracy.

Moreover, future studies may explore additional traffic sources and sampling strategies to improve measurement completeness across diverse online ecosystems. Another promising direction is to investigate how typosquatting domains are exploited across different application contexts and how domain registration ecosystems can strengthen policy enforcement to reduce malicious registrations.

Finally, it is worth noting that there is currently no universally agreed-upon definition of online abuse or malicious activity. In this study, a composite definition of malicious activity was adopted, including threats such as phishing and malicious downloads. However, this definition does not encompass the full spectrum of emerging threats. Future research should therefore aim to develop more comprehensive threat taxonomies and detection frameworks that capture the evolving landscape of online abuse.

Conclusion

To investigate and evaluate abusive behavior of Traffic Distribution Systems (TDSs) across different environments and their associated threats, this study introduced the Online Abusive Traffic Finder (OATF). Four different types of traffic sources and 5 different types of user profiles were integrated to identify cloaking strategies, targeted redirections and adaptive behavior by adversaries. The analysis demonstrates that these systems mostly operated in interconnected manners that redirect the users towards malicious and abusive destination pages. The usage of different user agents and proxy configurations highlights the dynamic and evasive nature of TDSs. A CNN-based classification component was introduced to identify the abusive web pages through visual feature extraction as a supporting automated analysis. The CNN model achieved accuracy of 91.92%, showing its effectiveness in the proposed framework. Although the proposed systems provide valuable insights, but still certain limitations discussed in detail in the Challenges and Future Directions section. Overall, this work offers a comprehensive perspective on abusive web traffic ecosystems and highlights the importance of combining large-scale data collection with automated detection techniques to better understand and mitigate evolving web-based threats.

Acknowledgements

The authors extend their appreciation Princess Nourah bint Abdulrahman University Researchers Supporting Project number (PNURSP2026R384), Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia. The authors extend their appreciation to the Deanship of Research and Graduate Studies at King Khalid University for funding this work through Small Research Project under grant number RGP1/80/46.

Author contributions

Farkhanda Athar Conceptualization; investigation; methodology; validation; writing—original draft., Akmal Shahbaz Data curation; formal analysis; project administration; resources; supervision; writing—original draft., Sana Munir Formal analysis; software; investigation; validation; visualization; writing—original draft., Mansoor Qadir Data curation; investigation; methodology; resources; software; visualization; writing-review and editing., Anandhavali Muiasamy Project administration; resources; software; validation; visualization; writing-review and editing., Hend Khalid Alkahtani Conceptualization; funding acquisition; project administration; resources; supervision; validation; writing-review and editing. All authors have read and agreed to the published version of the manuscript.

Funding

The authors extend their appreciation Princess Nourah bint Abdulrahman University Researchers Supporting Project number (PNURSP2026R384), Princess Nourah bint Abdulrahman University, Riyadh, Saudi Arabia. The authors extend their appreciation to the Deanship of Research and Graduate Studies at King Khalid University for funding this work through Small Research Project under grant number RGP1/80/46.

Data availability

The data that support the findings of this study are available from the corresponding author upon reasonable request.

Declarations

Competing interests

The authors declare no competing interests.

Footnotes

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

References

  • 1.Szurdi, J., & Christin, N.: Email typosquatting. In: Proceedings of the 2017 internet measurement conference, pp. 419–431. https://doi.org/10.1145/3131365.3131399 (2017).
  • 2.Chen, Z., Chen, H., Freire, J., Nagler, J., & Tucker, J. A.: Understanding how people consume low quality and extreme news using web traffic data. https://doi.org/10.48550/arXiv.2201.04226 (2022).
  • 3.Leontiadis, N., Moore, T., & Christin, N.: A nearly four-year longitudinal study of search-engine poisoning. In: Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security, pp. 930–941. https://doi.org/10.1145/2660267.2660332 (2014).
  • 4.Nowroozi, E., Mohammadi, M., Conti, M., et al.: An adversarial attack analysis on malicious advertisement url detection framework. https://doi.org/10.48550/arXiv.2204.13172 (2022).
  • 5.Umamageswari, B., Anandhi, K., & Sindhuja, M. Real-Time Phishing URL Detection by using XGBoost and Google Safe Browsing API. In 2025 5th International Conference on Soft Computing for Security Applications (ICSCSA) (pp. 186-191). IEEE. https://doi.org/10.1109/ICSCSA66339.2025.11171104 (2025)
  • 6.Le Page, S., Jourdan, G.-V., Bochmann, G. V., Flood, J., & Onut, I.-V.: Using url shorteners to compare phishing and malware attacks. In: 2018 APWG Symposium on Electronic Crime Research (eCrime), pp. 1–13. https://doi.org/10.1109/ECRIME.2018.8376215 (2018).
  • 7.Miramirkhani, N., Starov, O., & Nikiforakis, N.: Dial one for scam: A large scale analysis of technical support scams. In: 24th Annual Network and Distributed System Security Symposium. https://doi.org/10.14722/ndss.2017.23163 (2017).
  • 8.Oest, A., Safaei, Y., Zhang, P., Wardman, B., Tyers, K., Shoshitaishvili, Y., & Doupé, A.: Phishtime: Continuous longitudinal measurement of the effectiveness of anti-phishing blacklists. In: 29th USENIX Security Symposium (USENIX Security 20), pp. 379–396. (2020).
  • 9.Koide, T., Nakano, H. & Chiba, D. Chatphishdetector: Detecting phishing sites using large language models. IEEE Access.10.1109/ACCESS.2024.3483905 (2024). [Google Scholar]
  • 10.Li, Z., Alrwais, S., Xie, Y., Yu, F., & Wang, X.: Finding the linchpins of the dark web: A study on topologically dedicated hosts on malicious web infrastructures. In: 2013 IEEE Symposium on Security and Privacy, pp. 112–126. https://doi.org/10.1109/SP.2013.18 (2013).
  • 11.Chen, Z., & Frere, J. Discovering and measuring malicious url redirection campaigns from fake news domains. In: 2021 IEEE Security and Privacy Workshops (SPW). https://doi.org/10.1109/SPW53761.2021.00008 (2021).
  • 12.Saric, K., Savins, F., Ramachandran, G.S., Jurdak, R. and Nepal, S. May. Hyperlink Hijacking: Exploiting Erroneous URL Links to Phantom Domains. In Proceedings of the ACM Web Conference 2024 (pp. 1724-1733). https://doi.org/10.1145/3589334.3645510 (2024)
  • 13.Li, Xigao and Yepuri, Anurag and Nikiforakis, Nick. Double and Nothing: Understanding and Detecting Cryptocurrency Giveaway Scams. https://doi.org/10.14722/ndss.2023.24584 (2023).
  • 14.R, Upendra Shetty D and Patil, Anusha and Mohana. Malicious URL Detection and Classification Analysis using Machine Learning Models. In 2023 International Conference on Intelligent Data Communication Technologies and Internet of Things (IDCIoT) 470-476. http://doi.org/10.1109/IDCIoT56793.2023.10053422 (2023).
  • 15.Garimella, K., Kostakis, O., & Mathioudakis, M.: Ad-blocking: A study on performance, privacy and counter-measures. In: 9th International ACM Web Science Conference, pp. 259–262. https://doi.org/10.1145/3091478.3091514 (2017).
  • 16.Gudla, S., Kumar, Bose, J., Sunkara, S., & Verma, S. A unified push notifications service for mobile devices. In: 2015 IEEE International Conference on Electronics, Computing and Communication Technologies (CONECCT), pp. 1–6. https://doi.org/10.1109/CONECCT.2015.7383922 (2015).
  • 17.Khan, M. T., Huo, X., Li, Z., & Kanich, C.: Every second counts: Quantifying the negative externalities of cybercrime via typosquatting. In: 2015 IEEE Symposium on Security and Privacy, pp. 135–150. https://doi.org/10.1109/SP.2015.16 (2015).
  • 18.Zeng, Y., Zang, T., Zhang, Y., Chen, X., & Wang, Y.: A comprehensive measurement study of domain-squatting abuse. In: IEEE International Conference on Communications (ICC). pp. 1–6. https://doi.org/10.1109/ICC.2019.8761388 (2019).
  • 19.Berger, H., Dvir, A. Z. & Geva, M. A wrinkle in time: A case study in DNS poisoning. Int. J. Info. Security.19, 313–329. 10.48550/arXiv.1906.10928 (2021). [Google Scholar]
  • 20.Barford, P., Canadi, I., Krushevskaja, D., Ma, Q., & S. Muthukrishnan.: Adscape: Harvesting and analyzing online display ads. In: Proceedings of the 23rd international conference on World wide web, pp. 597–608. https://doi.org/10.1145/2566486.2567992 (2014).
  • 21.Bashir, M., Ahmad, Arshad, S., Kirda, E., Robertson, W., & Wilson, C.: How tracking campanies circumvent ad blockers using websockets. In: Proceedings of the Internet Measurement Conference 2018, pp. 471–477. https://doi.org/10.1145/3278532.3278573 (2018).
  • 22.Bashir, M., Ahmad, Arshad, S., Robertson, W., & Wilson, C. Tracing information flows between ad exchanges using retargeted ads. In: 25th USENIX Conference on Security Symposium, pp. 481–496. Preprint at: https://doi.org/10.48550/arXiv.1811.00920 (2016).
  • 23.Wang, D., Dai, S., Ding, Y., Li, T., & Han, X.: POSTER: AdHoneyDroid – Capture Malicious Android Advertisements. In: Proceedings of the 2014 ACM SIGSAC Conference on Computer and Communications Security, pp. 1514–1516. https://doi.org/10.1145/2660267.2662395 (2014).
  • 24.Le Pochat, V., Ballard, C., Desmet, L., Joosen, W., McCoy, D. and Lauinger, T. Partnerka in Crime: Characterising Deceptive Affiliate Marketing Offers. In International Conference on Passive and Active Network Measurement (pp. 405-436). Cham: Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-85960-1_17 (2025).
  • 25.Vekaria, Y., Shafiq, Z. & Zannettou, S. Auditing the Compliance and Enforcement of Twitter’s Advertising Policy. Soc. Med. Soc.11(1), 20563051251319676. 10.48550/arXiv.2309.12591 (2025). [Google Scholar]
  • 26.Zhang, R., Sridhar, R.P., Yao, M., Yang, Z., Oygenblik, D., Xu, H., Dave, V., Herley, C., England, P. and Saltaformaggio, B. Identifying Incoherent Search Sessions: Search Click Fraud Remediation Under Real-World Constraints. In 2025 IEEE Symposium on Security and Privacy (SP) (pp. 93-111). IEEE. https://doi.org/10.1109/SP61157.2025.00111 (2025)
  • 27.Hashmi, S. S., Ikram, M., & Kaafar, M. A.: A longitudinal analysis of online ad-blocking blacklists. In: 2019 IEEE 44th LCN Symposium on Emerging Topics in Networking (LCN Symposium), pp. 158–165. https://doi.org/10.1109/LCNSymposium47956.2019.9000671 (2019).
  • 28.Li, W., Laghari, S. U. A., Manickam, S. & Chong, Y. W. Exploration and evaluation of human-centric cloaking techniques in phishing websites. KSII Trans. Internet Inform. Syst. (TIIS)19(1), 232–258. 10.3837/tiis.2025.01.011 (2025). [Google Scholar]
  • 29.Choo, E. et al. A large scale study and classification of virustotal reports on phishing and malware urls. Proceedings of the ACM on Measurement and Analysis of Computing Systems7(3), 1–26. 10.1145/3626790 (2023). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 30.Huang, C.-T., Sakib, M. N., Kamhoua, C. A., Kwait, K. A., & Njilla, L.: A bayesian game theoretic approach for inspecting web-based malvertising. In: IEEE Transactions on Dependable and Secure Computing, 17(6), pp. 1257–1268. https://doi.org/10.1109/TDSC.2018.2866821 (2018).
  • 31.Salahdine, F. & Kaabouch, N. Social engineering attacks: A survey. Future Internet11(4), 89. 10.3390/fi11040089 (2019). [Google Scholar]
  • 32.Vadrevu, P., & Perdisci, R.: What you see is not what you get: Discovering and tracking social engineering attack campaigns. In: Proceedings of the Internet Measurement Conference, pp. 308–321. https://doi.org/10.1145/3355369.3355600 (2019).
  • 33.Sheikhalishahi, S. M., Martinelli, A., La, M. F., Mejri, A. & Tawbi, M. N. Digital Waste Disposal: An automated framework for analysis of spam emails. Int. J. Info. Security.19, 499–522. 10.1007/s10207-019-00470-x (2020). [Google Scholar]
  • 34.Vekaria, Y., Agarwal, V., Agarwal, P., Mahapatra, S., Muthiah, S., Balan, Sastry, N., & Kourtellis, N.: Differential tracking across topical webpages of indian news media. In: 13th ACM WEB Science Conference 2021, pp. 299–308. https://doi.org/10.1145/3447535.3462497 (2021).
  • 35.Kidmose, E., Stevanovic, M., & Pedersen, J. M.: Detection of malicious domains through lexical analysis. In: 2018 International Conference on Cyber Security and Protection of Digital Services (Cyber Security), pp. 1–5. https://doi.org/10.1109/CyberSecPODS.2018.8560665 (2018).
  • 36.Kumavat, P., & Khandare, N.: Threats involved with internet advertisements and attacks on botnet networks. In International Journal of Advance Research, Ideas and Innovations in Technology, vol. 4. (2018).
  • 37.Li, B., Vadrevu, P., Lee, K., Hyung, & Perdisci, R.: Jsgraph: Enabling reconstruction of web attacks via efficient tracking of live in-browser javascript executions. In: 25th Annual Network and Distributed System Security Symposium (NDSS), pp. 18–21. (2018).
  • 38.Li, Z., Zhang, K., Xie, Y., Yu, F., & Wang, X. Knowing your enemy: Understanding and detecting malicious web advertising. In: Proceedings of the 2012 ACM conference on Computer and communications security, pp. 674–686. https://doi.org/10.1145/2382196.2382267 (2013).
  • 39.Zarras, A., Kapravelos, A., Stringhini, G., Holz, T., Kruegel, C., & Vigna, G.: The dark alleys of madison avenue: Understanding malicious advertisements. In: Proceedings of the 2014 Conference on Internet Measurement Conference, pp. 373–380. https://doi.org/10.1145/2663716.2663719 (2014).
  • 40.Zeng, Y., Liu, Z., Tian, J., Chen, X., & Zang, T.: Hidden path: Understanding the intermediary in malicious redirections. In: IEEE Transactions on Information Forensics and Security. https://doi.org/10.1109/TIFS.2022.3169923 (2022).
  • 41.Saini, A., Gaur, M. S., Laxmi, V. & Conti, M. You click, I steal: Analyzing and detecting click hijacking attacks in web pages. Int. J. Info. Security.18, 481–504. 10.1007/s10207-018-0423-3 (2019). [Google Scholar]
  • 42.Rafique, M. Z., Van Goethem, T., Joosen, W., Huygens, C., & Nikiforakis, N.: It’s free for a reason: Exploring the ecosystem of free live streaming services. In: Proceedings of the 23rd Network and Distributed System Security Symposium (NDSS 2016), pp. 1–15. https://doi.org/10.14722/ndss.2016.23030 (2016).
  • 43.Szurdi, J., Luo, M., Kondracki, B., Nikiforakis, N., & Christin, N. Where are you taking me? understanding abusive tribution systems. In: Proceedings of the Web Conference 2021, pp. 3613–3624. https://doi.org/10.1145/3442381.3450071 (2021).
  • 44.Subramani, K., Yuan, X., Setayeshfar, O., Vadrevu, P., Lee, H., Kyu, & Perdisci, R.: When push comes to ads: Measuring the rise of (malicious) push advertising. In: Proceedings of the ACM Internet Measurement Conference, pp. 724–737. https://doi.org/10.1145/3419394.3423631 (2020).
  • 45.Agten, P., Joosen, W., Piessens, F., & Nikiforakis, N.: Seven months’ worth of mistakes: A longitudinal study of typosquatting abuse. In Proceedings of the 22nd Network and Distributed System Security Symposium (NDSS 2015). https://doi.org/10.14722/ndss.2015.23058 (2015).
  • 46.Alrwais, S., Yuan, K., Alowaisheq, E., Li, Z., & Wang, X.: Understanding the dark side of domain parking. In: 23rd USENIX Security Symposium (USENIX Security 14), pp. 207–222. https://doi.org/10.1007/978-3-642-15257-3_7 (2014).
  • 47.Lulli, R. R. R. W. D., & Levin, D.: Deceiving users with generic top-level domains. (2020).
  • 48.Masri, R., & Aldwairi, M.: Automated malicious advertisement detection using virustotal, urlvoid, and trendmicro. In: 2017 8th International Conference on Information and Communication Systems (ICICS), pp. 336–341. https://doi.org/10.1109/IACS.2017.7921994 (2017).
  • 49.Nikiforakis, N., Maggi, F., Stringhini, G., Rafique, M. Z., Joosen, W., Kruegel, C., Piessens, F., Vigna, G., & Zanero, S.: Stranger danger: Exploring the ecosystem of ad-based url shortening services. In: Proceedings of the 23rd international conference on World wide web, pp. 51–62. https://doi.org/10.1145/2566486.2567983 (2014).
  • 50.Papadopoulos, P., Ilia, P., Polychronakis, M., Markatos, E. P., Ioannidis, S., & Vasiliadis, G. Master of web puppets: Abusing web browsers for persistent and stealthy computation. In: 26th Annual Netwok and Distributed System Security Symposium, NDSS 2019, pp. 24–27. https://doi.org/10.48550/arXiv.1810.00464.
  • 51.Moodi, M. & Ghazvini, M. A new method for assigning appropriate labels to create a 28 Standard Android Botnet Dataset (28-SABD). J. Ambient Intell. Human Comput.10, 4579–4593. 10.1007/s12652-018-1140-5 (2019). [Google Scholar]
  • 52.Singh, M., Singh, M. & Kaur, S. Identifying bot infection using neural networks on DNS traffic. J. Comput. Virol. Hack. Tech.10.1007/s11416-023-00462-5 (2023). [Google Scholar]
  • 53.Wang, Y., Guo, C., Yan, J., Zhang, Z. & Cheng, Y. Unmasking hidden threats: Enhanced detection of embedded malicious domains in pirate streaming videos. Comput. Electr. Eng.123, 110087. 10.1016/j.compeleceng.2025.110087 (2025). [Google Scholar]
  • 54.Chaudhari, S., Thakur, A., & Rajan, A. An Efficient Malicious URL Detection Approach Using Machine Learning Techniques. In International Conference on Women Researchers in Electronics and Computing (pp. 485-495). https://doi.org/10.1007/978-981-99-7077-3_48 (2023).
  • 55.Liu, R., Wang, Y., Guo, Z., Xu, H., Qin, Z. & Zhang, F. PyraTrans: Attention-Enriched Pyramid Transformer for Malicious URL Detection. Preprint at: https://arxiv.org/abs/2312.00508 (2023).
  • 56.Xie, L., Zhang, H., Yang, H., Hu, Z. & Cheng, X. A scalable phishing website detection model based on dual-branch TCN and mask attention. Comput. Networks263, 111230. 10.1016/j.comnet.2025.111230 (2025). [Google Scholar]
  • 57.Yoon, J.-H., Buu, S.-J. & Kim, H.-J. Phishing webpage detection via multi-modal integration of HTML DOM graphs and URL features based on graph convolutional and transformer networks. Electronics13(16), 3344. 10.3390/electronics13163344 (2024). [Google Scholar]
  • 58.Kibriya, H. et al. Lightweight malicious URL detection using deep learning and large language models. Scientific Reports15, 43044. 10.1038/s41598-025-26653-2 (2025). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 59.Birthriya, S. K., Ahlawat, P. & Jain, A. K. Enhanced phishing website detection using dual-layer CNN and GRU with attention mechanism and lexical NLP features. SN Comput. Sci.5(7), 929. 10.1007/s42979-024-03282-6 (2024). [Google Scholar]
  • 60.Zhou, J. et al. A malicious URL detection framework based on custom hybrid spatial sequence attention and logic constraint neural network. Symmetry17(7), 987. 10.3390/sym17070987 (2025). [Google Scholar]
  • 61.Li, Y., Wang, Y., Xu, H., Guo, Z., Cao, Z. & Zhang, L. URLBERT: A Contrastive and Adversarial Pre-trained Model for URL Classification. arXiv preprint arXiv:2402.11495. Preprint at: https://arxiv.org/abs/2402.11495 (2024).
  • 62.Yang, Y. et al. A data-driven detection and localization framework for false data injection attacks in DC microgrids. IEEE Trans. Smart Grid14(3), 2154–2166. 10.1109/TSG.2022.3223456 (2023). [Google Scholar]
  • 63.Srinivasan, B., Kountouras, A., Miramirkhani, N., Alam, M., Nikiforakis, N., Antonakakis, M., & Ahamad, M. Exposing search and advertisement abuse tactics and infrastructure of technical support scammers. In: Proceedings of the 2018 World Wide Web Conference, pp. 319–328. https://doi.org/10.1145/3178876.3186098 (2018).
  • 64.Kharraz, A., Robertson, W., & Kirda, E.: Surveylance: Automatically detecting online surevy scams. In 2018 IEEE Symposium on Security and Privacy (SP), pp. 70–86. https://doi.org/10.1109/SP.2018.00044 (2018).
  • 65.Misquitta, J., & Kannan, A. A Comparative Study of Malicious URL Detection: Regular Expression Analysis, Machine Learning, and VirusTotal API. In International Congress of Electrical and Computer Engineering (pp. 219-232). hrrps://doi.org/10.21203/rs.3.rs-3685949/v1 (2023).
  • 66.Vissers, T., Barron, T., Van Goethem, T., Joosen, W., & Nikiforakis, N.: The wolf of name street: Hijacking domains through their nameservers. In: Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, pp. 957–970. https://doi.org/10.1145/3133956.3133988 (2017).
  • 67.Papadogiannakis, E., Papadopoulos, P., Markatos, E. P., & Kourtellis, N.: Who funds misinformation? A systematic analysis of the ad-related profit routines of fake news sites. https://doi.org/10.1145/3543507.3583443 (2022).
  • 68.Vissers, T., Joosen, W., & Nikiforakis, N.: Parking sensors: Analyzing and detecting parked domains. In: Proceedings of the 22nd Network and Distributed System Security Symposium (NDSS 2015), pp. 53–53. https://doi.org/10.14722/ndss.2015.23053 (2015).
  • 69.Vu, D.-L., Pashchenko, I., Massacci, F., Plate, H., & Sabetta, A.: Typosquatting and combosquatting attacks on the Python ecosystem. In: 2020 IEEE European Symposium on Security and Privacy Workshops (EuroS&PW), pp. 509–514. https://doi.org/10.1109/EuroSPW51379.2020.00074 (2020).
  • 70.Usha, D., Niveditha, V. & Priya, K. S. Anomaly detection on android system. Int J. Innov. Technol. Explor. Eng.9(1), 4310–4313. 10.35940/ijitee.A4951.119119 (2019). [Google Scholar]
  • 71.A. Singh and P. K. Roy. Malicious URL Detection using Multilayer CNN. In International Conference on Innovation and Intelligence for Informatics, Computing, and Technologies (3ICT), Zallaq, Bahrain, 2021, pp. 340-345. https://doi.org/10.1109/3ICT53449.2021.9581880 (2021).
  • 72.Gao, C. et al. Do users care about ad’s performance costs? Exploring the effects of the performance costs of in-app ads on user experience. Inform. Softw. Technol.10.1016/j.infsof.2020.106471 (2021). [Google Scholar]
  • 73.Tahir, R., Raza, A., Ahmad, F., Kazi, J., Zaffar, F., Kanich, C., & Caesar, M.: It’s all in the name: Why some urls are more vulnerable to typosquatting. In: IEEE INFOCOM 2018-IEEE Conference on Computer Communications, pp. 2618–2626. https://doi.org/10.1109/INFOCOM.2018.8486271 (2018).
  • 74.Zhang, M., Meng, W., Lee, S., Lee, B., & Xing, X.: All your clicks belong to me: Investigating click interception on the web. In: 28th USENIX Security Symposium (USENIX Security 19). pp. 941–957. (2019).
  • 75.Oest, A., Safaei, Y., Doupé, A., Ahn, G.-J., Wardman, B., & Tyers, K.: Phishfarm: A scalable framework for measuring the effectiveness of evasion techniques against browser phishing blacklists. In: 2019 IEEE Symposium on Security and Privacy (SP), pp. 1344–1361. https://doi.org/10.1109/SP.2019.00049 (2019).
  • 76.Hao, Q., Luo, L., Jan, S. T., & Wang, G. It’s not what it looks like: Manipulating perceptual hashing based applications. In: Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security, pp. 69–85. https://doi.org/10.1145/3460120.3484559 (2021)
  • 77.Garg, K., Alam, S., Ayala, D.C., Weigle, M.C., Nelson, M.L., 2025. Not Here, Go There: Analyzing Redirection Patterns on the Web. Preprint at: https://arxiv.org/abs/2507.22019.
  • 78.Tian, Y., Yu, Y., Sun, J., Wang, Y., 2025. From Past to Present: A Survey of Malicious URL Detection Techniques, Datasets and Code Repositories. Preprint at: https://arxiv.org/abs/2504.16449.
  • 79.Amouri, A., Al Rahhal, M. M., Bazi, Y., Butun, I., & Mahgoub, I. Enhancing intrusion detection in IoT environments: An advanced ensemble approach using Kolmogorov-Arnold networks. In 2024 International Symposium on Networks, Computers and Communications (ISNCC) (pp. 1-6). IEEE. https://doi.org/10.48550/arXiv.2408.15886 (2024)
  • 80.Dobrojevic, M. et al. Cyberbullying Sexism Harassment Identification by Metaheuristics-Tuned eXtreme Gradient Boosting. Comput. Mater. Continua.10.32604/cmc.2024.054459 (2024). [Google Scholar]
  • 81.Zivkovic, T. et al. Applying metaheuristic optimization for insider-threat detection using natural language processing. Int. J. Data Sci. Anal.22(1), 68. 10.1007/s41060-025-00996-5 (2026). [Google Scholar]
  • 82.Rao, T. S., Shariff, S. M., Francis, S. & Rohan, R. Advanced ML techniques for systematic fake news detection a comprehensive review. Macaw Int. J. Adv. Res. Comput. Sci. Eng.10(1), 86–93. 10.70162/mijarcse/2024/v10/i1/v10i1s10 (2024). [Google Scholar]
  • 83.Savanovic, N. et al. Hybrid CNN XGBoost intrusion detection approach tuned by modified sine cosine algorithm towards better cloud security. Connect. Sci.37(1), 2549581. 10.1080/09540091.2025.2549581 (2025). [Google Scholar]
  • 84.Zhao, S., Kalra, A., Borcea, C., & Chen, Y.: To be tough or soft: Measuring the impact of counter-ad-blocking strategies on user engagement. In: Proceedings of The Web Conference 2020, pp. 690–2696. https://doi.org/10.1145/3366423.3380025 (2020).
  • 85.Samarasinghe, N., & Mannan, M. On cloaking behaviors of malicious websites. Computer and Security, 101. https://doi.org/10.1016/j.cose.2020.102114 (2021).
  • 86.Todri, V. Frontiers: The impact of ad-blockers on online consumer behavior. Market. Sci.41(1), 7–18. 10.1287/mksc.2021.1309 (2022). [Google Scholar]
  • 87.Ahmad, W., Sen, A., Eesley, C. & Brynjolfsson, E. Companies inadvertently fund online misinformation despite consumer backlash. Nature630(8015), 123–131. 10.1038/s41586-024-07404-1 (2024). [DOI] [PMC free article] [PubMed] [Google Scholar]
  • 88.Papadogiannakis, E., Papadopoulos, P., Markatos, E.P. and Kourtellis, N., 2024. Before & After: The Effect of EU’s 2022 Code of Practice on Disinformation. Preprint at https://arxiv.org/abs/2410.11369
  • 89.Mi, X., Feng, X., Liao, X., Liu, B., Wang, X., Qian, F., Li, Z., Alrwais, S., Sun, L., & Liu, Y.: Resident evil: Understanding residential ip proxy as a dark service. In: 2019 IEEE symposium on security and privacy (SP), pp. 1185–1201. https://doi.org/10.1109/SP.2019.00011 (2019).
  • 90.Dao, H., Mazel, J. & Fukuda, K. Characterizing cname cloaking-based tracking on the web. IEEE/IFIP TMA20, 1–9. 10.1109/TNSM.2021.3072874 (2020). [Google Scholar]
  • 91.Zhang, P., Oest, A., Cho, H., Sun, Z., Johnson, R., Wardman, B., Sarker, S., Kapravelos, A., Bao, T., Wang, R., et al.: Crawlphish: Large-scale analysis of client-side cloaking techniques in phishing. In: 2021 IEEE Symposium on Security and Privacy (SP), pp. 1109–1124. https://doi.org/10.1109/SP40001.2021.00021 (2021)
  • 92.Teoh, X. PhishDecloaker: Detecting CAPTCHA-cloaked Phishing. In: Proceedings of the USENIX Security Symposium (2024).
  • 93.Nakano, H., Koide, T. & Chiba, D. PhishParrot: LLM-Driven Adaptive Crawling to Unveil Cloaked Phishing Sites. Preprint at: https://arxiv.org/abs/2508.02035 (2025).
  • 94.Chromium blog. https://blog.chromium.org/2020/01/ (2022).
  • 95.Selenium: Getting started with selenium. https://www.selenium.dev/documentation/getting_started/ (2021).
  • 96.Puppeteer vs selenium. https://www.browserstack.com/guide/puppeteer-vs-selenium (2022). [Online; accessed 2-May-2022]

Associated Data

This section collects any data citations, data availability statements, or supplementary materials included in this article.

Data Availability Statement

The data that support the findings of this study are available from the corresponding author upon reasonable request.


Articles from Scientific Reports are provided here courtesy of Nature Publishing Group

RESOURCES