Skip to main content
NIHPA Author Manuscripts logoLink to NIHPA Author Manuscripts
. Author manuscript; available in PMC: 2026 Apr 23.
Published in final edited form as: Res Nurs Health. 2025 Sep 26;48(6):750–762. doi: 10.1002/nur.70021

Addressing Survey Fraud in Online Health Research: A Case Study of Latine Sexual Minority Men

Lisvel A Matos 1, Susan Silva 1, Michael V Relf 1,2, Rosa Gonzalez-Guarda 1
PMCID: PMC13101259  NIHMSID: NIHMS2160039  PMID: 41001781

Abstract

Introduction.

Online survey research has become an increasingly popular and effective method in the social sciences for exploring and addressing health-related issues. However, the increasing prevalence of fraudulent activities, particularly survey bots, threatens data integrity and can compromise health research by generating misleading data.

Purpose.

The purpose of this paper was to describe the implementation of bot detection strategies in an online survey with Latine sexual minority men (SMM).

Methods.

Eleven bot detection indicators, including AI-detection software for open-ended responses, were used in two approaches to differentiate bot-generated from human responses. In the first approach, bot detection indicators were applied stepwise to identify valid entries. In the second approach, a fraud detection algorithm was used to identify three fraud categories. Key demographics and study variables were compared across fraud groups using chi-square/Fisher’s Exact tests for categorical data and Kruskal-Wallis tests for continuous data (significance set at 0.05).

Results.

Of the 1,147 total survey entries, 837 (73%) completed at least 20% of the survey (814 completed all items). A total of 739 (88%) of the 837 completed surveys were classified as fraudulent. Among the 837 completed surveys, 333 (40%) had an AI-generated open-ended response and fast completion time (≤ 20 minutes) and 234 (28%) entries were flagged for all three of these indicators. Sociodemographic characteristics and HIV prevention outcomes were largely similar across bot-generated and human responses.

Conclusion.

Findings suggest that survey bots are a pervasive threat to online research and are effective at providing human-like responses. To protect data integrity and ensure the development of effective health policies and interventions, health science researchers should adopt comprehensive bot detection and prevention strategies.

Keywords: Data integrity, bot detection, AI bots, surveys

1 |. INTRODUCTION

The utilization of online survey research has become increasingly widespread among social science researchers as an effective method to explore and address critical public health issues. Online methods, including the use of email campaigns, dedicated webpages, and social media advertising, offer a scalable, affordable means to engage large numbers of participants (Christensen et al., 2017). Online surveys are also regarded as an effective strategy to engage hard-to-reach groups, including historically marginalized groups such as racial and ethnic minoritized populations and sexual and gender minoritized populations (Bauermeister et al., 2012; Dalessandro, 2018; Ramo & Prochaska, 2012). The anonymity of online surveys facilitates the exploration of sensitive health topics in the area of sexuality and sexual health (Bybee et al., 2022), including HIV research among Latine sexual minority men (SMM) (Andrade et al., 2023; Fitch et al., 2022; Garcia & Saw, 2019).

An urgent concern in online surveys is the surge in fraudulent activities caused by survey bots. Although programmers’ motivations for deploying malicious bots vary, they are often tied to fraudulent enterprises like survey farms, where individuals or groups exploit research study incentives (Pasternak, 2019; Pozzar et al., 2020). Additionally, the use of social media bots has become a widespread concern, with automated accounts being deployed to influence public perception, manipulate political discourse, and spread misinformation on health-related topics, such as COVID-19 (Himelein-Wachowiak et al., 2021; Pew Research Center, 2018). These same tactics have extended into the research space, where entities seeking to distort study outcomes may program survey bots to intentionally skew results (Pew Research Center, 2020). Academic research is especially at risk, with some researchers reporting that 60% to 90% of responses in their studies were bot-generated (Brainard et al., 2022; Cascalheira, 2023; Griffin et al., 2022). Recent reports indicate a rise in survey fraud on studies targeting studies on underrepresented minoritized groups (Keppler, 2023), with one research team finding that fraudulent participants were more likely to identify as a Latine and living with HIV (Grey et al., 2015).

Survey bot interference in online research can severely undermine data quality and validity, increasing the risk of Type I and Type II errors (Cascalheira, 2023; Storozuk et al., 2020). Despite the significant threats bots pose to academic research, investigators using online surveys rarely report on techniques used to address bot interference (Storozuk et al., 2020). Ensuring the integrity of research findings is foundational to evidence-based decision-making in health disciplines. This issue is particularly important to mitigate harm to minoritized groups, who are often the focus of online studies and may be disproportionately targeted by malicious attacks. The purpose of this paper was to describe the implementation of bot detection strategies in an online survey of Latine SMM.

2 |. BACKGROUND

As online studies became more widespread in the early 2000s, survey fraud was rare, with reported occurrence rates as low as 3% (Teitcher et al., 2015; Topp & Pawloski, 2002). In the 2010s, researchers began to notice a rise in prevalence of online survey fraud particularly through duplicate submissions (Teitcher et al., 2015). During this time, signs of bot-driven survey fraud began to emerge, including unusually fast completion times (Teitcher et al., 2015), repeated survey entries from the same IP address (Bauermeister et al., 2012), and identical response patterns (Teitcher et al., 2015). This rise in fraud has been observed across multiple recruitment platforms, including crowdsourcing sites and social media (Bybee et al., 2022; Cloud Research, 2018). Despite these concerning trends, the literature has rarely addressed bot infiltration in academic research (Pinzón et al., 2023).

2.1 |. Bot detection strategies

Many survey platforms implement bot prevention measures to distinguish between human users and automated bots. One widely used tool is reCAPTCHA, an enhanced version of CAPTCHA (Completely Automated Public Turing test to tell Computers and Humans Apart) developed by Google to verify that users are human (Google, 2025). However, bots with advanced programming can bypass reCAPTCHA, which limits its effectiveness in preventing bots from infiltrating online surveys (Griffin et al., 2022; Simone, 2019). Consequently, researchers have focused on developing bot-detection strategies that can be implemented during the data cleaning stage after data collection (Brainard et al., 2022; Bybee et al., 2022; Dewitt et al., 2018; Simone, 2019; Storozuk et al., 2020), such as monitoring IP addresses to identify duplicate entries, survey completion times, and responses to open-ended questions (Cascalheira, 2023; Pinzón et al., 2023; Storozuk et al., 2020).

A few researchers have provided empirical findings on the implementation and effectiveness of bot detection strategies in online surveys. Pinzón et al. (2023), tested 31 bot indicators across two online surveys and found that rapid completion times, consecutive survey start times, duplicate IP addresses, and suspicious email addresses were strong predictors of survey fraud. In a separate study, Project Queer Survivors of Trauma (QueST), employed an iterative process to test and determine degree of confidence in 11 bot detection strategies (Cascalheira, 2023). The researchers corroborated findings from Pinzón et al. (2023), and reported that discrepancies in redundant questions responses and participants who avoid contact with research team members were indicators highly predictive of survey fraud.

As researchers develop strategies to combat survey bot activity, programmers continue to advance bot algorithms, making them increasingly effective at mimicking human behavior. Recent studies indicate that these bots can learn from taking surveys and adapt their responses, posing a growing threat to data integrity (Cascalheira, 2023; King-Nyberg et al., 2023; Pinzón et al., 2023). Additionally, the rise of generative artificial intelligence (e.g., ChatGPT) may enable bots to produce coherent and contextually appropriate open-ended responses, further complicating detection efforts (King-Nyberg et al., 2023; Pinzón et al., 2023).

2.2 |. The present study

Building on previous work, this paper aims to evaluate the implementation of empirical bot detection methods in an online survey of Latine SMM. Specifically, we first describe the online study design and bot detection measures. We then: (a) analyze recruitment and response patterns, (b) apply bot detection strategies to identify fraudulent survey entries, and (c) compare key demographic and HIV prevention variables between valid and fraudulent responses. These findings provide empirical evidence on the effectiveness of current bot detection strategies and offer insights for refining future approaches in online survey research.

3 |. METHODS

3.1 |. Online study design

The online study was a cross-sectional survey that examined individual and social determinants of HIV self-protection among Latine SMM. The survey collected information about sociodemographic characteristics, cultural phenomena such as optimism in the American dream, experiences with discrimination, HIV prevention, and mental wellbeing. REDcap (Harris et al., 2019) was used to create the survey which included the eligibility screening form, electronic consent form, and survey measures. Researchers and four Latine SMM consultants pilot-tested the survey, estimating a completion time of up to one hour for screening, consent, and survey completion. Institutional Review Board approval was obtained from Duke University (IRB# Pro00112026) prior to data collection.

3.2 |. Sample and setting

Eligibility criteria included being 18–35 years old, Latine, biologically male, a U.S. resident, and having engaged in anal sex with a man in the past three months. To reach young Latinx (the term used in the original advertisement) SMM a Facebook advertising campaign was run from February 2 to February 12, 2024. Facebook ads targeted individuals who identified as Latinx/e, were male, and between the ages of 18 and 35, and living in the U.S. Eligible participants who completed the survey were offered $40 as compensation.

3.3 |. Study procedures

The Facebook ad included a direct survey link. Upon access, participants completed a reCAPTCHA before the eligibility questionnaire. Ineligible participants were redirected to a thank-you page, while eligible ones selected their language preference to review the informed consent form in English or Spanish. The survey was bilingual, with English and Spanish text displayed side by side on the same page. Responses were required, but a “prefer not to answer” option was available. At the end, participants completed a compensation form with their name, email, and phone number.

3.4 |. Bot prevention and detection strategies

Prior to data collection, a bot detection protocol was established using evidence from the literature (Cascalheira, 2023; Pinzón et al., 2023; Storozuk et al., 2020). This protocol outlined various measures (detailed below) designed to detect bot activity. The protocol specified that the study team would contact participants as needed to verify their identity via telephone call or email. Participants were also made aware that they may be contacted for verification when they provided their contact information. As bot detection strategies were implemented, survey entries were flagged for further inspection. Further inspection included looking at patterns and inconsistencies in entries across multiple measures for bot detection detailed in Table 1 and below.

TABLE 1.

Bot detection indicators

Bot indicators Point value
 1) Is the survey entry timestamp the same as a previous entry (same minute)? 2
 2) Is the survey entry timestamp one minute apart from the previous entry? 2
 3) Is the survey entry timestamp 2-3 minutes apart from the previous entry? 1
 4) Does the first and last name match a previous entry? 2
 5) Does the email address match a previous entry? 2
 6) Does the phone number match a previous entry? 2
 7) Are ages within the entries different? 2
 8) Was the duration of the survey ≤ 20.0 minutes? 2
 9) Open-ended response was flagged as AI generated by CopyLeaks? 2
 10) Do open-ended responses match a previous entry word for word? 2
 11) Do the first 25 characters of an open-ended response match a previous entry? 1

Note: Fraud algorithm- Each entry was assigned total fraud score. Entries with a total fraud score between 1-2 points were classified as suspicious. Entries with a score of ≥ 3 or more points were categorized as fraudulent.

3.5 |. Measures for bot detection

3.5.1 |. Survey timestamps

Survey timestamps were recorded for each respondent, capturing survey initiation, time spent on each section, and completion time (Cascalheira, 2023; Pinzón et al., 2023; Zhang et al., 2022). These timestamps were analyzed to identify anomalous patterns indicative of bot activity. Specifically, we examined start times to identify duplicate or irregular temporal patterns, such as consecutive survey entry times. Consecutive survey entries were defined as those occurring either: (a) on the same minute; (b) one minute apart; or (c) two to three minutes apart from the previous entry.

3.5.2 |. Survey duration

Survey entries with response times significantly deviating from the average human completion time were flagged for further review. Based on pilot testing, the expected completion time was 45 to 60 minutes; therefore, entries completed in ≤ 20 minutes were identified for inspection (Pinzón et al., 2023).

3.5.3 |. Personal information

Survey entries were reviewed for duplicate entries of names, email addresses, and phone numbers (Pinzón et al., 2023; Zhang et al., 2022). Unusual email address patterns, such as combinations of names (e.g., John.Smith@domain.com) or random alphanumeric sequences (e.g., xyz556abc@domain.com), have been used as indicators of bot activity (Ballard et al., 2019; Pinzón et al., 2023; Storozuk et al., 2020). However, given the subjectivity of this measure (Storozuk et al., 2020), email addresses were primarily used for participant verification and reviewed only when survey entries were flagged for further inspection.

3.5.4 |. Redundant questions

Redundant questions have been used to detect discrepancies in verifiable information, serving as indicators of bot activity (Cascalheira, 2023; Simone, 2019). In this study, age was asked in two separate sections using a free-text response format. One section restricted responses to the eligibility range of 18–35 years, while the other allowed unrestricted input. Entries with inconsistent age responses were flagged for review.

3.5.5 |. Open-ended response analyses

The survey included three open-ended questions exploring beliefs about the American Dream, such as, “What does the American Dream mean to you?” Responses were analyzed for AI-generated content using Copyleaks AI (Copyleaks, 2024), an AI detection tool trained to identify human writing and detect disruptions in human writing patterns. Additionally, responses were manually reviewed for duplication and nonsensical content. Survey entries containing AI-generated text or repetitive patterns were flagged for further review.

3.6 |. Implementation of bot detection strategies

Bot detection indicators in Table 1 were employed to identify fraudulent survey entries using two approaches. The first approach was a sequential bot detection approach, in which all survey entries (N = 1,147) were reviewed for the presence of any of the 11 bot indicators. This included incomplete survey entries, ineligible survey entries, and completed survey entries. Bot detection indicators were applied one at a time in a stepwise manner. Survey entries flagged by each indicator were removed sequentially, and the process continued until only entries with no indicators of bot activity remained. A key limitation of this approach is that it classifies survey entries based on a single source of evidence, rather than allowing multiple indicators to collectively support bot detection. This may increase the likelihood of false positives (i.e, the misclassification of human participants as bots).

In the second approach, we developed a fraud detection algorithm that was implemented manually and informed by prior research which guided the scoring of each entry based on the summation of the separate indicators present (Ballard et al., 2019; Cascalheira, 2023; Pinzón et al., 2023; Wang et al., 2023). We first applied bot detection indicators to all survey entries that were more than 20% complete (N=837). Then, each bot detection indicator was assigned a point value of 1 or 2 based on its predictive power for identifying bot activity, as established in previous research studies with large samples (Cascalheira, 2023; Pinzón et al., 2023). More specifically, the definition of predictive power used by Pinzon and colleagues (2023) was used to generate an estimate of precision or predictive accuracy in correctly identifying fraudulent responses. Indicators with high predictive power (i.e., measures that strongly predict bot interference in previous studies) received a score of 2, while those with low to moderate predictive power or that had not yet been empirically tested were assigned a score of 1. Although the use of Copyleaks to detect AI-generated open-ended responses has not been previously validated in the context of research survey bot detection, it was assigned a score of 2 because it is a software tool specifically designed and validated to identify AI-generated text. Survey entries were then classified into three fraud categories based on their total score: (a) fraudulent (met two or more indicators, totaling 3 or more points), (b) suspicious (met one indicator, totaling 1–2 points), and (c) potentially valid (no indicators met, 0 points). This classification strategy is supported by previous research, which recommends combining at least two bot detection indicators to reduce the risk of false positives and improve the accuracy of fraud detection compared to relying on a single indicator alone (Ballard et al., 2019; Buchanan & Scofield, 2018; Pozzar et al., 2020).

3.6 |. Analyses

Descriptive analyses were performed to describe characteristics of completed surveys and the three fraud categories established using the fraud detection algorithm. Demographic characteristics of the sample (e.g., age, nativity, education level) and HIV self-protection variables were summarized. Key demographic and study variables were compared across fraud categories (fraudulent, suspicious, and potentially valid) using chi-square/Fisher’s Exact tests for categorical variables. Kruskal Wallis tests were utilized to compare ranked values for continuous variables. If the overall group effect was significant at the 0.05 level, a posteriori pairwise comparisons were conducted using Wilcoxon Two-Sample tests for continuous variables and 2x2 chi-square/Fisher’s Exact tests for categorical variables. All quantitative analyses and non-directional statistical tests were conducted using SAS software, version 9.4 (SAS Institute Inc, 2015), with significance set at 0.05 for all tests. The primary purpose of comparing key demographic and study variables across fraud classification groups was to assess whether response patterns differed by classification group. This analysis was conducted to explore the assumption that form-filling bots generate random responses, which may produce normally distributed patterns across survey items. Open ended questions were uploaded to Copyleaks which provided results of text that was AI generated. Open-ended responses exported to Microsoft Excel and analyzed for duplication patterns. To identify duplicates, we applied conditions that highlighted both identical responses and those that began with the same first 25 characters.

4 |. RESULTS

We begin with an overview of advertising campaign metrics, survey entries, and initial observations suggesting bot activity. Next, we present findings from the implementation of the two bot detection approaches previously described. Finally, we examine differences in sociodemographic characteristics and key study across fraud classification groups.

4.1 |. Advertisement campaign

The advertising campaign ran on Facebook from February 2 to February 12, 2024. During this 10-day period, the ads generated a total of 16,962 impressions (the number of times the ad was displayed on users’ screens) and 401 clicks (the number of times users clicked on the ad). A total of $100 was spent on ad placement through Facebook’s advertising platform, which operates on a pay-per-click model. This resulted in an average cost per click of $0.25.

4.2 |. Survey entries

A total of 1,137 survey attempts were recorded from February 2 to 18, 2024. Notably, this number far exceeded the 401 clicks generated by the Facebook advertising campaign, and was one of the initial indicators that prompted further investigation into data validity. The survey link was temporarily deactivated from February 20 to 21, and an additional 10 attempts were made between March 14 and 18 before the survey was permanently closed on March 19, 2024. This brought the overall total to 1,147 survey entries. Of these, 33 participants accessed the survey but did not complete the eligibility section. Among the 1,114 remaining entries, 68 participants were deemed ineligible, primarily because they identified as non-Hispanic/Latino or reported living with HIV. Additionally, 175 eligible participants did not proceed beyond the consent form. After these exclusions, 837 (73%) surveys had at least 20% completion, of which 814 (71%) were fully completed.

4.3 |. Initial observations

The first author actively monitored incoming survey responses and quickly identified multiple red flags suggesting bot interference. The most immediate concern was the rapid submission of survey entries, which began on February 6, 2024, at 21:00 and continued until February 8, 2024, at 09:00. During this 36-hour period, a total of 922 survey entries were recorded (Figure 1). Notably, 717 of these entries occurred on February 7, 2024, despite only 34 clicks being recorded on the Facebook study ads that day. This stark discrepancy between ad clicks and survey submissions strongly suggested automated activity.

FIGURE 1.

FIGURE 1

Number of survey responses per day.

Additionally, approximately one-third of all survey entries (32%, n=352) were submitted between 00:00 and 06:00, a pattern previously identified as indicative of bot interference (Pozzar et al., 2020; Storozuk et al., 2020). Further suspicious patterns were observed in email addresses and participant names. Many email addresses followed formulaic structures (e.g., John.Smith@domain.com, xyz556abc@domain.com) (Cascalheira, 2023; Storozuk et al., 2020). The presence of female names, despite eligibility criteria requiring biological males, and the limited representation of Hispanic surnames raised additional concerns about the authenticity of these entries.

4.5 |. Bot detection strategies

4.5.1 |. Sequential bot detection

The sequential implementation of these steps is detailed in Figure 2. After removing survey entries that failed the bot detection indicators, there were a total of 23 potentially valid survey entries.

FIGURE 2.

FIGURE 2.

Flow diagram - sequential removal of fraudulent survey entries.

4.5.2 |. Fraud detection algorithm

Table 2 summarizes bot detection indicators, reporting the number and percentage of flagged records among participants who completed more than 20% of the survey (N=837).

TABLE 2.

Bot detection indicators in survey entries ≥20% complete (N = 837)

Indicators Description n (%)
1) Consecutive entry Survey entry timestamps the identical to a previous entry 378 (45.2)
2) Consecutive entry Survey entry timestamp one minute apart from another entry 136 (16.2)
3) Consecutive entry Survey entry timestamp 2-3 minutes apart from another entry 91 (10.9)
4) Duplication First and last name matches a previous entry 18 (2.2)
5) Duplication Email address matches a previous entry 2 (0.2)
6) Duplication Phone number matches a previous entry 67 (8.0)
7) Redundant Question Different ages reported within the same entry 35 (4.2)
8) Speed Survey entries completed in ≤ 20.0 minutes 457 (54.6)
9) Open-ended response AI generated text 572 (68.3)
10) Open-ended response Identical to another survey entry 296 (35.4)
11) Open-ended response The first 25 characters of response identical to another survey entry 358 (44.7)

Note: N=837 represents survey entries with 20% or more of the survey items completed. Consecutive survey entries were categorized into three mutually exclusive indicators: 1) duplicate timestamp, 2) one minute apart, and 3) 2-3 minutes apart. Open-ended responses were categorized into three indicators:1) AI-generated text, 2) duplicate open-ended response, 3) duplicate first 25 characters of open-ended responses, with the latter two being mutually exclusive classifications.

The fraud detection algorithm was then applied to 837 survey entries to classify fraud categories. A total of 739 (88%) of 837 survey entries were classified as fraudulent, 82 (10%) were classified as suspicious, and 16 (2%) had no flags or indicators of fraud. Scores ranged from 0 to 11 points, with a median score of 5.0 (25th-75th percentile:4.0-7.0) points across survey entries. Table 3 presents the number and percent with each bot indicator for the 821 entries classified as either fraudulent (N=739) or suspicious entries (N=82).

TABLE 3.

Bot detection indicators by fraud category (N=821)

Fraudulent survey entries
(N=739)
n (%)
Suspicious survey entries
(N=82)
n (%)
Time and duration
 Consecutive entry timestamp (duplicate) 371 (50.2) 7 (8.5)
 Close proximity timestamp (1 minutes apart) 128 (17.3) 9 (11.0)
 Close proximity timestamp (2-3 minutes apart) 79 (10.7) 12 (14.6)
Personal information
 Duplicate name 17 (2.3) 0 (0)
 Duplicate phone number 66 (8.9) 1 (1.2)
 Duplicate email 2 (0.2) 0 (0)
Redundant items
 Age check 33 (4.5) 2 (2.4)
Speed: ≤ 20 minutes 448 (60.6) 9 (11.0)
Open-ended responses
 Duplicated responses (entire text) 290 (39.2) 7 (8.5)
 Duplicated responses (first 25 characters) 351 (47.5) 18 (22.0)
 AI generated response 554 (75.0) 24 (29.3)

Note: N=Sample size for each fraud category; Category with no fraud flags/ indicators omitted

Among the 837 completed survey entries, 333 (40%) had an AI-generated open-ended response and fast completion time (≤ 20 minutes) and 234 (28%) survey entries were flagged for all three of these indicators. When examining 739 fraudulent entries, 365 (49%) met the criteria for both consecutive survey entry within 1 minute (indicator #1 and #2 from Table 1) and providing an AI generated open-ended response. A similar pattern was observed among the 82 entries classified as suspicious. Among the 82 suspicious entries, 49 entries (60%) had flagged open-ended responses (i.e., met criteria #9, #10, or #11 from Table 1) and 26% of the 49 entries flagged were also flagged for consecutive survey entry of ≤ 1 minute.

4.6 |. Comparison of bots vs nots

We examined differences in sociodemographic characteristics and key study variables between survey entries classified as fraudulent, suspicious, and potentially valid survey entries (i.e., no indicators of bot activity).

4.6.1 |. Participant characteristics

Table 4 presents the sociodemographic and HIV prevention characteristics for participants who completed more than 20% of the survey (N=837) and comparisons of these characteristics across the three fraud categories (i.e., fraudulent, suspicious, and potentially valid). The median age of participants was 28. Educational attainment was equally distributed across different levels of education. Specifically, 28% had completed a high school diploma or GED, 24% had earned an associate’s degree, 29% held a bachelor’s degree, and 19% had achieved a graduate degree or higher. Over half of the participants were born outside of the U.S. (54%). Most participants were in a relationship (72%) and identified as gay (76%). There were statistically significant differences in education between the three fraud categories (p<.0001). Notably, the fraudulent group had a higher proportion of participants reporting their highest level of education attainment as high school/GED (30%) relative to the suspicious (15%) or no flags (19%) groups. Thus, the fraudulent group had a smaller proportion with post-secondary education. A posteriori comparisons of three fraud categories indicated the education attainment distribution for the fraudulent group was significantly different from the suspicious group (p<.0001), but the other comparisons were not statistically significant (all p>0.05). The three categories did not significantly differ on any of the other sociodemographic characteristics (all p>0.05).

TABLE 4.

Sociodemographic characteristics by entry classification.

N Total
(N=837)
Fraudulent
(N=739)
Suspicious
(N=82)
No Flags
(N=16)
p-value
Age in years, Mdn (25th-75th) 834 28.0 (25.0-31.0) 29.0 (29.0-31.0) 28.0 (26.0-30.0) 25.0 (25.0-30.0) 0.2940
Education, n (%) 837 <0.0001
 High School Diploma/GED 236 (28.2) 221 (29.9) 12 (14.6) 3 (18.8)
 Associate’s degree 198 (23.7) 186 (25.2) 10 (12.2) 2 (12.5)
 Bachelor’s degree 242 (28.9) 203 (27.5) 32 (39.0) 7 (43.8)
 Graduate Degree or higher 161 (19.2) 129 (17.5) 28 (34.2) 4 (25.0)
Nativity, foreign-born, n (%) 837 449 (53.6) 396 (53.6) 42 (51.2) 11 (68.8) 0.4354
Married/partnered, n (%) 837 605 (72.3) 537 (72.7) 58 (70.7) 10 (62.5) 0.6323
Gay sexual identity, n (%) 833 629 (75.5) 552 (74.9) 62 (76.5) 15 (100) 0.0588
HIV testing in last 6 months, n (%) 830 652 (78.6) 587 (80.1) 52 (64.2) 13 (81.5) 0.0051
STI treatment in last 3 months, n (%) 831 246 (29.6) 223 (30.4) 17 (21.0) 6 (37.5) 0.1672
HIV counseling in lifetime, n (%) 828 633 (76.4) 561 (76.4) 59 (72.8) 13 (81.3) 0.6990
PrEP use in lifetime, n (%) 833 486 (58.3) 431 (58.6) 42 (51.9) 13 (81.5) 0.0865
PrEP use last 3 months, n (%) 486 270 (55.6) 235 (54.9) 27 (64.3) 8 (61.5) 0.4637

Note: Abbreviations: N = Data Available; Mdn (25th-75th) = Median (25th-75th percentile); column median (25th, 75th) and column number and percent (%) reported

4.6.2 |. HIV prevention characteristics

Most participants reported testing for HIV in the last six months (79%), and nearly one-third had received treatment for a sexually transmitted infection in the last 3 months (30%). Most participants had received HIV prevention counseling from a healthcare worker (i.e., provider, nurse, sex counselor) in their lifetime (76%). More than half of participants reported PrEP use in their lifetime (58%). Of the 486 participants who had a history of PrEP use, more than half reported using PrEP in the last three months (56%). There was a statistically significant fraud category difference in the proportion of participants who had been tested for HIV in the last 6 months (p=0.0051). Specifically, the fraudulent group (80%) had a significantly higher rate of HIV testing in the last 6 months relative to the suspicious group (64%, p=.0029), but the other comparisons were not statistically significant (all p>0.05).

4.7 |. Participant verification

All participants were contacted to ensure that valid participants had the opportunity to verify their participation. Verification was based on their ability to recall general details about the survey, such as that it collected information on sexual health and experiences with discrimination. Figure 3 contains the email script sent to participants.

FIGURE 3.

FIGURE 3

Participant Verification Email.

Of the email addresses provided, five were invalid. We received a total of 106 email responses. The vast majority of these were highly suspicious, as they arrived in rapid succession during overnight hours (i.e., between 00:00 and 05:00) and contained brief, generic messages, often duplicated across multiple email addresses, such as “Yes, I confirm” or “I will be happy to participate”. We followed up with these individuals to clarify that a brief verification call to confirm their participation. Only one individual completed the follow-up verification call. However, they declined to turn on their camera and, more importantly, could not provide any details to verify their participation in the survey. Avoiding contact with researchers has been noted as a common indicator of fraudulent participation in online studies, as fraudulent actors often evade direct verification attempts to avoid detection (Cascalheira, 2023; Lawrence et al., 2023). Combined with the suspicious nature of the email responses, this lack of engagement strongly reinforces the conclusion that these entries were fraudulent.

4.8 |. Insights from open-ended response analyses

The open-ended responses analyses offer important insights into survey bot behavior since the growing availability of generative AI models (Table 5). In previous studies, open-ended responses completed by bots were non-sensical and therefore easy to determine that they were not from human respondents (Cascalheira, 2023; Simone, 2019). While some responses contained one- to four-word nonsensical responses such as “gay” and “pilot”, the majority of responses were sensical and used uncharacteristically accurate punctuation, spelling, and syntax. Upon further inspection, there were a number of survey entries with identical open-ended responses across one or more questions, but an unexpected finding was the number of survey entries that began with the same words but were not exact matches to another record. For instance, 46 survey entries began their response to one of the open-ended questions with “Immigrants may see the American Dream as…”, followed by slight changes in the remainder of the sentence which often conveyed the same meaning and sentiment, which was suggestive of bot interference.

TABLE 5.

Open-ended Response Analyses (N=814).

Question 1: What does the American dream mean to you? [Mean response length: 102.4 characters]
Short, nonsensical responses: “gay”, “pilot”, “none”
Overly accurate responses: “For me, the American Dream means having the opportunity to work hard and achieve my goals, regardless of where I come from or my background. It’s about having the freedom to pursue my passions, build a successful career, and create a better life for myself and my loved ones. The American Dream represents the belief that with determination and equal opportunities, anyone can succeed and fulfill their aspirations in this land of possibilities.”
Similar responses: “To me, the American Dream means endless possibilities and opportunities. As a Latino American, I know that my ancestors came to this country in search of a better life. The American Dream made me believe that if I work hard, I have a chance to achieve my dreams.”

“To me, the American Dream means the pursuit of freedom and opportunity. As a Latino immigrant, I came to the United States in search of a better life and opportunity, which is what the American Dream represents.”

Question 2: Do you feel that the American Dream is accessible to everyone? Why or why not? [Mean response length: 91.9 characters]
Short, nonsensical responses: “may”, “can”, “the ability”
Overly accurate responses: “I believe that the American Dream is accessible to everyone who is willing to put in the effort and determination to pursue their goals. While there may be obstacles and challenges along the way, the fundamental promise of the American Dream is that anyone, regardless of background or circumstances, can succeed through hard work and perseverance. With access to education, opportunities, and a supportive community, individuals can overcome barriers and realize their dreams, making the American Dr” [response maxed out on allowable characters]
Similar responses: Yes, the American Dream is achievable by everyone if they have access to mentorship and support networks that can guide them toward success.”

Yes, the American Dream is achievable by everyone if they have access to a supportive community and resources for personal and professional development.”

Question 3: Do you feel that the American Dream is accessible to everyone? Why or why not? [Mean response length: 105.5 characters]
Short, nonsensical responses: “2222222”, “incomprehension”
Overly accurate responses: I think the American Dream can hold a similar meaning for both immigrants and those born in the U.S., but there may be some nuances and personal perspectives based on their unique experiences. Immigrants often come to the U.S. seeking better opportunities and a chance to build a brighter future, just like those born here. However, immigrants may also have additional layers to their dreams, such as finding stability, embracing their cultural heritage, and creating a sense of belonging in their ne” [response maxed out on allowable characters]
Similar responses: “For immigrants, the ‘American Dream’ may be more symbolic and represent their efforts and struggles, while for those born in the United States, the concept may be more abstract and vague.”

“For immigrants, realizing the ‘American Dream’ may require experiencing more challenges and difficulties, while people born in the United States may be able to enjoy such a dream more easily.”

Note: N=814 represents participants who completed the final survey section, which included the open-ended questions.

The use of Copyleaks software to scan open-ended responses, implemented after a manual review of entries, was illuminating, as 68% of entries were flagged as AI-generated. This finding demonstrates that bot programmers are using generative AI techniques to produce responses that superficially mimic human input. However, these responses retained identifiable automated patterns. For example, this section was completed at an unusually fast speed for human participants. The mean duration of time spent on this section, which included three open-ended responses and 18 matrix-style questions section, was 111 seconds. These findings highlight the need for adapting detection strategies to address the evolving sophistication of bot responses, and underscore importance of employing multiple bot detection strategies.

5 |. DISCUSSION

This paper presents methods for detecting bot interference in an online study of Latine SMM. Of the 837 survey entries with at least 20% of the items completed, the application of a fraud detection algorithm revealed that 88% had clear evidence of bot interference, 10% were highly suspicious of bot activity, and only 2% showed no indications of such interference. Efforts to verify participants through direct contact were unsuccessful, further suggesting that nearly all responses were likely generated by bots. Notably, this is the first study, to our knowledge, to employ an AI detection tool to identify AI-generated responses in open-ended survey questions. The tool proved highly effective, revealing that a substantial portion of these responses were AI-generated. By emphasizing empirical bot detection indicators, this study significantly advances efforts to identify and mitigate survey fraud in online research. To support continued progress in this area, we outline key recommendations for addressing bot interference in online studies (see Table 6).

TABLE 6.

Recommendations and Future Directions

Recommendation Rationale Implementation Tip Example from this study
Use dynamic, multi-tool detection protocols Bot programming is rapidly evolving, especially with AI. Static methods become outdated quickly. • Integrate a range of detection tools and regularly assess tool performance.

• Update protocols based on emerging fraud trends.
Added Copyleaks AI detection after human review found sensical and duplicated open-ended responses.
Avoid overly strict criteria Sequential bot detection methods may exclude valid responses and produce false positives. • Use a holistic approach that considers multiple forms of evidence. Apply ≥2 strong, validated indicators or ≥3 when relying on subjective or unvalidated indicators.

• Consider behaviors that may be common among marginalized groups, such as VPN use and shared devices, when implementing bot detection to prevent false positives.
Sequential bot detection and fraud detection algorithm yielded similar results; however, fraud detection algorithm allowed for consideration of multiple sources of evidence to support bot activity.
Verify participants when feasible Verifying identities can deter bots and fraudsters. • Consider brief phone calls or other accessible verification to reduce survey fraud.

• Weigh potential barriers against benefits, particularly when recruiting marginalized populations, and implement verification steps when necessary to protect against bots and individual fraud.
No participants completed video verification.
Address ethical concerns When selecting bot detection strategies, researchers must balance the risks and benefits to minimize both false positives (misclassifying valid participants as bots) and false negatives (allowing bot responses to appear legitimate). • Collaborate with IRBs to promote transparency and accountability.

• Develop bot detection study protocols that detail bot detection strategies to be used, criteria for classifying fraudulent records, standardized email and phone scripts for participant verification.

• Identify predefined actions for handling fraudulent survey entries, such as verification procedures and withholding compensation.
A bot detection protocol was created and approved by IRB prior to beginning the study.
Prioritize equitable study participation Strategies to prevent and mitigate bot interference (e.g., removing study incentives) might inadvertently prevent researchers from engaging genuine participants in their study • Assess and adjust incentive structures to deter bots without creating barriers for underrepresented participants.

• Remove incentives from advertisements to deter bots.

• Use alternative verification methods, such as utility bills, when participants may not have access to government-issued ID.
Advertising incentives in the social media ad may have made this study a target for survey fraud.
Prepare for time and resource demands of bot detection Bot interference can dramatically increase study costs and time. • Prioritize robust bot prevention and detection strategies to protect limited budgets and resources, especially for doctoral and early career researchers.

• Utilize pilot studies to test survey vulnerabilities and test effectiveness of bot detection protocols.
Bot detection and data cleaning strategies in this study took 3 months.

Although no participants were verified, compensating all 1,147 submissions or the 837 completed surveys would have cost approximately $45,880- $33,480.

Our findings are consistent with previous studies suggesting that bot interference has become pervasive in online-based research (Ballard et al., 2019; Storozuk et al., 2020; Wang et al., 2023; Zhang et al., 2022). For instance, Pozzar and colleagues (2020) found that 94.5% of responses in an online survey were fraudulent, 5.5% were suspicious, and none were classified as legitimate. A common factor among studies with high bot interference is the exclusive use of social media for advertising, along with the public availability of survey links in these advertisements. In contrast, researchers using more varied recruitment methods and restricting survey link distribution reported significantly lower levels of fraudulent responses, approximately 35% (Ballard et al., 2019; Irish & Saba, 2023). These findings underscore the importance of implementing diverse recruitment strategies and robust bot detection measures to safeguard data integrity in online research.

Despite clear evidence of widespread bot interference, the ability of bots to imitate human response patterns poses a complex challenge for detection efforts (Irish & Saba, 2023). We observed minimal differences in sociodemographic characteristics and HIV prevention outcomes across fraud categories, This suggests form-filling bots may produce responses that approximate normal distributions (Buchanan & Scofield, 2018; Irish & Saba, 2023). These findings underscore that relying solely on response patterns is insufficient for detecting fraudulent entries, as bots are capable of producing highly plausible data (Irish & Saba, 2023). Notably, findings from our open-ended questions further support this concern. The open-ended responses in were varied in content, mimicking the natural diversity expected from human participants. Our findings indicate that bots are able to identify contextual cues in surveys and adjust their responses accordingly, making them harder to detect. There is a need for future research to better understand how bots influence response patterns across a broader range of survey measures, to inform the development of more sophisticated detection methods capable of distinguishing human responses from bot-generated responses.

Historically marginalized communities may be more susceptible to harm from by bot interference due to the increasing reliance of online research as a strategy for recruiting these populations in health-related research (Dalessandro, 2018; Myers et al., 2022; Ramo & Prochaska, 2012). Although we cannot definitively attribute the high level of bot activity in our study to its focus on Latine SMM, our results are consistent with other researchers targeting historically marginalized populations who have reported substantial fraudulent activity (Ballard et al., 2019; Bybee et al., 2022; Griffin et al., 2022). Studies show that bots often report characteristics of marginalized populations, such as living with HIV or identifying as racial and ethnic minorities (Grey et al., 2015). In this study, over 50% of participants identified as immigrants, and 58% reported PrEP use, exceeding national estimates for Latine SMM which range from 24%-34% (Pitasi et al., 2021; Trujillo et al., 2019). Given the hostile political environment toward Latine and LGBTQ populations and the known use of social media bots to manipulate political narratives (Rossetti & Zaman, 2023), our findings are concerning. They suggest that bots could be used to commit survey fraud as well as to intentionally distort research findings and influence public discourse. In the future, researchers should prioritize detecting coordinated bot activity targeting research on marginalized groups and develop proactive strategies to protect the validity of findings that inform health research and public policy.

The high level of bot interference identified in this study underscores the need for rigorous bot detection strategies in online health research. Given the pervasive nature of bot interference, a valid critique of previous online research studies that fail to report bot detection or human verification methods is the potential compromise of data integrity. findings from such studies may be unreliable, leading to inaccurate conclusions that misrepresent population health needs and inform ineffective interventions. Implementing standardized bot detection protocols and transparency in reporting methodologies is essential to preserving the validity of online health research and ensuring that data-driven policies and interventions are based on accurate, high-quality evidence.

5.1 |. Limitations

There are several limitations to this study. First, there were a number of bot detection strategies that we did not implement. For example, hidden questions may be especially useful for identifying bots because they are not visible to human participants (Pinzón et al., 2023). Although our fraud detection algorithm was designed to minimize false positives, we acknowledge the possibility that some valid responses may have been inadvertently classified as fraudulent, given the high overall rate of flagged entries. Our study did not collect IP addresses from participants due to institutional restrictions; however, the collection of IP addresses may have provided additional insight into fraudster behaviors. To strengthen fraud prevention while maintaining participant anonymity, future studies might consider requesting IRB approval to collect hashed IP addresses, which are coded versions of the original IP address that cannot be traced back to the user but allow detection of duplicate entries. Hashed IPs offer a practical middle ground by enabling detection of duplicate entries without storing identifiable information. Additionally, we did not employ statistical analytics, such as mahalanobis distances or persontotal correlations, to examine response patterns (Irish & Saba, 2023). These methods could have further refined our detection of bot activity and improved the accuracy of identifying fraudulent responses. Despite these limitations, there are important implications for survey research we discuss through our recommendations (Table 6).

6 |. CONCLUSION

The presence of bots in online research will likely remain an ongoing challenge for social scientists. Since no single strategy can entirely prevent bots from infiltrating surveys or identify bot interference, it is essential to employ a multi-pronged approach to mitigate bot interference at every stage of the research process. By pre-registering bot detection protocols and transparently reporting detection strategies in research papers, social scientists can enhance data integrity, foster methodological transparency, and contribute to the development of more effective strategies for mitigating bot interference in online research. While our paper advances the understanding of bot detection in online research, significant questions remain regarding the optimal strategies to counteract this evolving threat. Continued research is needed to develop innovative solutions that protect the integrity of online survey data.

REFERENCES

  1. Andrade EA, Stoukides G, Santoro AF, Karasz A, Arnsten J, & Patel VV (2023). Individual and health system factors for uptake of pre-exposure prophylaxis among young Black and Latino gay men. Journal of General Internal Medicine. 10.1007/s11606-023-08274-5 [DOI] [PMC free article] [PubMed] [Google Scholar]
  2. Ballard AM, Cardwell T, & Young AM (2019). Fraud detection protocol for web-based research among men who have sex with men: Development and descriptive evaluation. JMIR Public Health and Surveillance, 5(1), e12344. 10.2196/12344 [DOI] [PMC free article] [PubMed] [Google Scholar]
  3. Bauermeister JA, Zimmerman MA, Johns MM, Glowacki P, Stoddard S, & Volz E (2012). Innovative recruitment using online networks: Lessons learned from an online study of alcohol and other drug use utilizing a web-based, respondent-driven sampling (webRDS) strategy. Journal of Studies on Alcohol and Drugs, 73(5), 834–838. [DOI] [PMC free article] [PubMed] [Google Scholar]
  4. Brainard J, Lane K, Watts L, & Bunn D (2022). The wasps are clever: Keeping out and finding bot answers in internet surveys used for health research (2022030243). Preprints. 10.20944/preprints202203.0243.v1 [DOI] [Google Scholar]
  5. Buchanan EM, & Scofield JE (2018). Methods to detect low quality data and its implication for psychological research. Behavior Research Methods, 50(6), 2586–2596. 10.3758/s13428-018-1035-6 [DOI] [PubMed] [Google Scholar]
  6. Bybee S, Cloyes K, Ellington L, Baucom B, Supiano K, & Mooney K (2022). Bots and nots: Safeguarding online survey research with underrepresented and diverse populations. Psychology and Sexuality, 13(4), 901–911. 10.1080/19419899.2021.1936617 [DOI] [PMC free article] [PubMed] [Google Scholar]
  7. Cascalheira CJ (2023). Yes stormtrooper, these are the droids you’re looking for: A method paper evaluating bot detection strategies in online psychological research. PsyArXiv. 10.31234/osf.io/gtp6z [DOI] [PMC free article] [PubMed] [Google Scholar]
  8. Christensen T, Riis AH, Hatch EE, Wise LA, Nielsen MG, Rothman KJ, Sørensen HT, & Mikkelsen EM (2017). Costs and efficiency of online and offline recruitment methods: A web-based cohort study. Journal of Medical Internet Research, 19(3), e6716. 10.2196/jmir.6716 [DOI] [PMC free article] [PubMed] [Google Scholar]
  9. Cloud Research. (2018, August 10). Concerns about bots on Mechanical Turk: Problems and solutions. CloudResearch. https://www.cloudresearch.com/resources/blog/concerns-about-bots-on-mechanical-turk-problems-and-solutions/ [Google Scholar]
  10. Copyleaks. (2024). Copyleaks AI detector [Computer software]. https://copyleaks.com/
  11. Dalessandro C (2018). Recruitment tools for reaching millennials: The digital difference. International Journal of Qualitative Methods, 17(1), 1609406918774446. 10.1177/1609406918774446 [DOI] [Google Scholar]
  12. Dewitt J, Capistrant B, Kohli N, Rosser BRS, Mitteldorf D, Merengwa E, & West W (2018). Addressing participant validity in a small internet health survey (The Restore Study): Protocol and recommendations for survey response validation. JMIR Research Protocols, 7(4), e96. 10.2196/resprot.7655 [DOI] [PMC free article] [PubMed] [Google Scholar]
  13. Fitch C, Haberer JE, Serrano PA, Muñoz A, French AL, & Hosek SG (2022). Individual and structural-level correlates of pre-exposure prophylaxis (prep) lifetime and current use in a nationwide sample of young sexual and gender minorities. AIDS and Behavior. 10.1007/s10461-022-03656-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  14. Garcia M, & Saw G (2019). Socioeconomic disparities associated with awareness, access, and usage of pre-exposure prophylaxis among Latino MSM ages 21–30 in San Antonio, TX. Journal of HIV/AIDS & Social Services, 18(2), 206–211. 10.1080/15381501.2019.1607795 [DOI] [PMC free article] [PubMed] [Google Scholar]
  15. Google. (2025). reCAPTCHA website security and fraud protection. Google Cloud. https://cloud.google.com/security/products/recaptcha [Google Scholar]
  16. Grey JA, Konstan J, Iantaffi A, Wilkerson JM, Galos D, & Simon Rosser BR (2015). An updated protocol to detect invalid entries in an online survey of men who have sex with men (MSM): How do valid and invalid submissions compare? AIDS and Behavior, 19(10), 1928–1937. 10.1007/s10461-015-1033-y [DOI] [PMC free article] [PubMed] [Google Scholar]
  17. Griffin M, Martino RJ, LoSchiavo C, Comer-Carruthers C, Krause KD, Stults CB, & Halkitis PN (2022). Ensuring survey research data integrity in the era of internet bots. Quality & Quantity, 56(4), 2841–2852. 10.1007/s11135-021-01252-1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  18. Harris PA, Taylor R, Minor BL, Elliott V, Fernandez M, O’Neal L, McLeod L, Delacqua G, Delacqua F, Kirby J, & Duda SN (2019). The REDCap consortium: Building an international community of software platform partners. Journal of Biomedical Informatics, 95, 103208. 10.1016/j.jbi.2019.103208 [DOI] [PMC free article] [PubMed] [Google Scholar]
  19. Himelein-Wachowiak M, Giorgi S, Devoto A, Rahman M, Ungar L, Schwartz HA, Epstein DH, Leggio L, & Curtis B (2021). Bots and misinformation spread on social media: Implications for COVID-19. Journal of Medical Internet Research, 23(5), e26933. 10.2196/26933 [DOI] [PMC free article] [PubMed] [Google Scholar]
  20. Irish K, & Saba J (2023). Bots are the new fraud: A post-hoc exploration of statistical methods to identify bot-generated responses in a corrupt data set. Personality and Individual Differences, 213, 112289. 10.1016/j.paid.2023.112289 [DOI] [Google Scholar]
  21. Keppler N (2023, February 15). Academics say bots keep targeting their research on LGBTQ health. Vice. https://www.vice.com/en/article/bvm9k8/survey-bot-fraud-lgbtq-health
  22. King-Nyberg B, Thomson EF, Morris-Reade J, Borgen R, & Taylor C (2023). The Bot Toolbox: An accidental case study on how to eliminate bots from your online survey. Journal for Social Thought, 7(1), Article 1. https://ojs.lib.uwo.ca/index.php/jst/article/view/14331 [Google Scholar]
  23. Lawrence PR, Osborne MC, Sharma D, Spratling R, & Calamaro CJ (2023). Methodological challenge: Addressing bots in online research. Journal of Pediatric Health Care, 37(3), 328–332. 10.1016/j.pedhc.2022.12.006 [DOI] [PubMed] [Google Scholar]
  24. Myers KJ, Jaffe T, Kanda DA, Pankratz VS, Tawfik B, Wu E, McClain ME, Mishra SI, Kano M, Madhivanan P, & Adsul P (2022). Reaching the “Hard-to-Reach” sexual and gender diverse communities for population-based research in cancer prevention and control: Methods for online survey data collection and management. Frontiers in Oncology, 12. 10.3389/fonc.2022.841951 [DOI] [PMC free article] [PubMed] [Google Scholar]
  25. Pasternak O (2019, January 10). Market research fraud—Distributed survey farms exposed. Persona.Ly - The Weekly Retention. https://persona.ly/blog/2019/01/market-research-fraud-distributed-survey-farms-exposed/
  26. Pew Research Center. (2018). Bots in the Twittersphere. Pew Research Center. https://www.pewresearch.org/internet/2018/04/09/bots-in-the-twittersphere/ [Google Scholar]
  27. Pew Research Center. (2020, February 18). Assessing the risks to online polls from bogus respondents. Pew Research Center. https://www.pewresearch.org/methods/2020/02/18/assessing-the-risks-to-online-polls-from-bogus-respondents/ [Google Scholar]
  28. Pinzón N, Koundinya V, Galt R, Dowling W, Boukloh M, Taku-Forchu NC, Schohr T, Roche L, Ikendi S, Cooper MH, Parker LE, & Pathak TB (2023). AI-powered fraud and the erosion of online survey integrity: An analysis of 31 fraud detection strategies. OSF. 10.31235/osf.io/95tka [DOI] [PMC free article] [PubMed] [Google Scholar]
  29. Pitasi MA, Beer L, Cha S, Lyons SJ, Hernandez AL, Prejean J, Valleroy LA, Crim SM, Trujillo L, Hardman D, Painter EM, Petty J, Mermin JH, Daskalakis DC, & Hall HI (2021). Vital signs: HIV infection, diagnosis, treatment, and prevention among gay, bisexual, and other men who have sex with men—United States, 2010–2019. Morbidity and Mortality Weekly Report, 70(48), 1669–1675. 10.15585/mmwr.mm7048e1 [DOI] [PMC free article] [PubMed] [Google Scholar]
  30. Pozzar R, Hammer MJ, Underhill-Blazey M, Wright AA, Tulsky JA, Hong F, Gundersen DA, & Berry DL (2020). Threats of bots and other bad actors to data quality following research participant recruitment through social media: Cross-sectional questionnaire. Journal of Medical Internet Research, 22(10), e23021. 10.2196/23021 [DOI] [PMC free article] [PubMed] [Google Scholar]
  31. Ramo DE, & Prochaska JJ (2012). Broad reach and targeted recruitment using Facebook for an online survey of young adult substance use. Journal of Medical Internet Research, 14(1), e1878. 10.2196/jmir.1878 [DOI] [PMC free article] [PubMed] [Google Scholar]
  32. Rossetti M, & Zaman T (2023). Bots, disinformation, and the first impeachment of U.S. President Donald Trump. PLOS ONE, 18(5), e0283971. 10.1371/journal.pone.0283971 [DOI] [PMC free article] [PubMed] [Google Scholar]
  33. SAS Institute Inc. (2015). Statistical analysis software 9.4 (Version 5th) [Computer software]. [Google Scholar]
  34. Simone M (2019, November 21). Bots started sabotaging my online research. I fought back. STAT. https://www.statnews.com/2019/11/21/bots-started-sabotaging-my-online-research-i-fought-back/
  35. Storozuk A, Ashley M, Delage V, & Maloney E (2020). Got bots? Practical recommendations to protect online survey data from bot attacks. The Quantitative Methods for Psychology, 16, 472–481. 10.20982/tqmp.16.5.p472 [DOI] [Google Scholar]
  36. Teitcher JEF, Bockting WO, Bauermeister JA, Hoefer CJ, Miner MH, & Klitzman RL (2015). Got bots? Practical recommendations to protect online survey data from bot attacks. The Journal of Law, Medicine & Ethics : A Journal of the American Society of Law, Medicine & Ethics, 43(1), 116–133. 10.1111/jlme.12200 [DOI] [PMC free article] [PubMed] [Google Scholar]
  37. Topp NW, & Pawloski B (2002). Online data collection. Journal of Science Education and Technology, 11(2), 173–178. 10.1023/A:1014669514367 [DOI] [Google Scholar]
  38. Trujillo L, Chapin-Bardales J, German EJ, Kanny D, & Wejnert C (2019). Trends sexual risk behaviors among Hispanic/Latino men who have sex with men—19 urban areas, 2011-2017. MMWR. Morbidity and Mortality Weekly Report, 68(40), 873–879. 10.15585/mmwr.mm6840a2 [DOI] [PMC free article] [PubMed] [Google Scholar]
  39. Wang J, Calderon G, Hager ER, Edwards LV, Berry AA, Liu Y, Dinh J, Summers AC, Connor KA, Collins ME, Prichett L, Marshall BR, & Johnson SB (2023). Identifying and preventing fraudulent responses in online public health surveys: Lessons learned during the COVID-19 pandemic. PLOS Global Public Health, 3(8), e0001452. 10.1371/journal.pgph.0001452 [DOI] [PMC free article] [PubMed] [Google Scholar]
  40. Zhang Z, Zhu S, Mink J, Xiong A, Song L, & Wang G (2022). Beyond bot detection: Combating fraudulent online survey takers. Proceedings of the ACM Web Conference 2022, 699–709. 10.1145/3485447.3512230 [DOI] [Google Scholar]

RESOURCES