Introduction
Self-administered web surveys have become increasingly common (Couper 2017; DeLeeuw 2018; Groves 2011; Prandner et al. 2023). Among these surveys, a large proportion of social research today relies on online panels from crowdsourcing platforms such as Amazon’s Mechanical Turk or Prolific. Online panels now account for a substantial share of data published in peer-reviewed empirical work in the social and behavioral sciences (Aguinis et al. 2021; Chandler et al. 2019; Paolacci and Chandler 2014). While these surveys offer promising results, concerns remain about their validity and reliability. This paper aims to improve response quality by examining the implementation of mandatory breaks within the online questionnaire.
Responses collected from online panels often raise data quality concerns, such as respondent inattentiveness, dishonesty, and survey fatigue (Chandler et al. 2020; Cornesse and Blom 2023). One major threat to data quality is a high respondent cognitive load (or cognitive burden). High cognitive load can lead to measurement errors, item non-response, or even break-offs (Dillman et al. 2014; Ward et al. 2017). Survey length is a key contributing factor affecting respondents’ cognitive burden (Emery et al. 2022). While shorter questionnaires are preferable, not all surveys can be shortened. This study aims to understand whether a mandatory break embedded within a survey can improve data quality, even in conditions where the survey is already at an ideal length of between 10 to 15 minutes (Revilla and Höhne 2020). To our knowledge, no prior experimental study has directly compared mandatory mid-survey breaks against no breaks on data quality outcomes.
To address concerns about low-quality data in online surveys, researchers have developed and implemented a range of assessment and monitoring methods to improve data quality, including attention-check questions and completion-time filtering. However, these approaches are reactive rather than proactive. A more proactive approach is to gamify surveys (Aubert et al. 2023), but not all surveys can be gamified.
For longer surveys, some researchers propose a “split questionnaire” design. Split questionnaire design can take the form of matrix sampling, in which different subgroups receive different sub-questionnaires, or delayed completion surveys, in which respondents complete one portion of the survey at a later time. However, these methods have pitfalls: matrix sampling requires imputation of missing data, and the delayed-completion method risks high drop-off and lower response rates (Adigüzel and Wedel 2008; Axenfeld et al. 2022; Peytchev and Peytcheva 2017; Raghunathan and Grizzle 1995). As an alternative approach to maintaining higher data quality, we explore introducing a mandatory break within the surveys. Unlike split-questionnaire designs, a mandatory break does not risk attrition due to delayed participation. Evidence from psychological literature suggests that introducing breaks into tasks can improve response quality. A short break within a task can encourage individuals to process more deliberately (Bilancini et al. 2024; Shirasuna et al. 2025). Meta-analytic evidence finds that short rest breaks produce small but statistically significant benefits for fatigue reduction and task performance, particularly for less cognitively demanding tasks (Albulescu et al. 2022). Our approach is designed to reduce cognitive burden and improve response quality with the goal of achieving similar benefits to the performance of respondents.
However, a mandatory break may also be viewed as a task interruption, which might lower data quality. Previous research has found that task interruptions impair performance because they require additional effort to navigate the task (Baethge et al. 2015; Yuan and Zhong 2024). If a mandatory break is perceived as a stressor or an interruption, rather than reducing fatigue and encouraging deliberation, introducing a mandatory break may decrease data quality, rather than improve it.
Given the limited prior research on the implementation of mandatory breaks embedded in surveys, we hypothesize that the effect can be positive or negative. While prior research suggests that a break within a cognitively burdensome questionnaire may increase data quality, other research mentioned above suggests that a mandatory break may have a negative effect. Therefore, we propose two competing hypotheses about the effect of mandatory breaks in surveys on all respondents:
H1: Alleviating effect. A mandatory break improves data quality by refreshing cognitive load.
H2: Worsening effect. A mandatory break decreases data quality by disrupting respondents’ engagement.
Methods
Our survey experiment examined the effect of a mandatory break on respondents in the United States, focusing on attitudes toward climate change policies. The questionnaire, comprising 71 questions, measured support for 22 climate policy items. The study employed a two-group experimental design. In the control group, participants completed the questionnaire without the mandatory break. In the experimental group, participants completed the first half of the questionnaire, including 11 key support item questions, before encountering a mandatory 60-second break. The climate policy items used standard closed-ended, Likert-type response scales—the dominant question format in attitudinal and public-opinion research—and the instrument fell within the 10-to-15-minute range that respondents and researchers typically regard as a reasonable survey length (Revilla and Höhne 2020).
The 22 climate policy items were administered as a parallel-structure battery, with respondents rating each policy on the same response scale. This battery format is widely adopted in attitudinal, public opinion, market research, and organizational surveys, and is the format in which concerns about respondent fatigue and straightlining are most often raised. This makes such instruments a natural test case for a mid-survey break intervention. In these respects, the questionnaire is representative of the moderate-length, closed-ended surveys that are routinely fielded on online panels, which supports the broader applicability of our findings beyond the specific topic studied here.
A total of 2,373 respondents were recruited via the Prolific platform in December of 2024. A convenience sample resembling U.S. demographics was created based on Prolific’s Census sample matching feature. Respondents were compensated $1.76 for a 10-minute survey. Additionally, half of the respondents assigned to the mandatory-break experimental condition received a $0.17 bonus, bringing the average hourly participation rate to about $10.02. This payment rate aligns with Prolific’s fair pay guidelines.
A growing concern for surveys fielded on online panels is the possibility that responses are produced or assisted by large language models (LLMs) rather than genuine human participants (Veselovsky et al. 2025). We took several steps to mitigate this risk. At the platform level, Prolific restricts participation to an identity-verified pool and applies pre-screening and ongoing behavioral monitoring designed to detect automated accounts and fraudulent activity, and independent comparisons have generally rated their data quality favorably relative to other panels (Douglas et al. 2023; Peer et al. 2021). At the study level, beyond the data-quality measures that constitute our outcomes, we screened on completion time and excluded responses (n = 260) flagged as timed out, erroneous, or otherwise invalid, resulting in a study sample of 2,113. Because the survey relied on closed-ended policy items rather than open-text responses, it offered limited opportunity for the kind of LLM-assisted free-text answering that is most difficult to detect. We nonetheless encourage future researchers using this design to incorporate platform authenticity checks, which have become more widely available since our data were collected.
Responses between the experimental and control groups were compared to determine whether the mandatory break produced a statistically significant difference between those who had a break and those who did not. During this break, a screen displaying a natural scenery image appears (see Appendix). A countdown timer counted down from 60 seconds.
Table 1 indicates the demographic composition of the full analytic sample. Starting from the 2,113 respondents retained after study-level screening, we further excluded 12 cases with item non-response on the main variables of interest. As a result, a total of 2,101 responses were analyzed in this study. Examining key sociodemographic characteristics, our sample includes 1,013 respondents identified as male (48.22%) and 1,068 as female (50.83%). Twenty respondents identified as either non-binary or a different term (0.95%). In terms of race, 69.68% identified as white, 14.95% Black, 5.47% Asian, 5.62% Latino, 3.24% multi-racial, 0.48% Native American, and 0.57% as other. Overall, the sample demographics resemble those of the U.S. population, according to the 2020 U.S. Census (U.S. Census Bureau 2021a; 2021b).
Dependent Variable
Previous studies have employed multiple measures to assess the quality of responses from online panels. Similarly, we build on Peer et al. (2021), who compared the data quality of different online panel respondents across four criteria: attention, comprehension, honesty, and reliability. Additionally, we examine repetitive response patterns (or “straightlining”) and the accuracy of sociodemographic data, as described above. In this study, we assess the data quality of responses based on the following five criteria:
-
Internal Reliability: Based on (1) averaged 11-item measures, (2) standard deviation across the 11 items, and (3) Cronbach’s alpha of response consistency between experimental groups.
-
Straightlining: Measurement of whether respondents excessively repeat their answers (e.g., same scale point) in a row.
-
Self-reported Honesty: Measured via a 0–100 slider, asking, “Did you participate properly in this survey, reading carefully, and giving your actual opinion?”
-
Attention Check: Based on whether the respondents passed or failed the attention check that was shown after the break.
-
Sociodemographic Data Inconsistency: Based on the comparison of the demographic data (such as age and gender) collected at the end of the questionnaire and the demographic data provided by Prolific.
Results
Internal Reliability
The internal reliability is inferred based on three measures: the averages of each of the 11 policy items after the mandatory break, the standard deviations, and Cronbach’s alpha of those items. Table 2 shows the differences in data quality between the experimental and control groups. The 11-item averages were 3.90 and 3.92 (on a 5-point scale) for the control and experimental groups, respectively. A t-test comparison showed no statistically significant differences (p = 0.51).
We also examined the standard deviation of each participant’s responses. Individual-specific standard deviation tells us the extent to which each participant’s responses are “spread” across the 11 items. The lack of spread can indicate that the individual responded with less variability, which may be a potential indicator of inattentiveness. The standard deviations of the 11-item responses were 0.91 and 0.92 for the control and experimental groups, respectively. No significant differences were found based on the t-test (p = 0.41).
Third, we compared Cronbach’s alpha coefficients of the 11-item responses between the control and experimental groups. Cronbach’s alpha is a statistic that denotes the internal consistency of those 11 items. The alphas were 0.86 and 0.83, for control and experimental conditions. Based on the Feldt test, the 95% confidence intervals of the two groups’ coefficients overlap, indicating no statistically significant difference. Based on three measures, we found no statistically significant difference in how respondents “behave” in terms of their responses, regardless of whether they were given a mandatory break. These results fail to provide support for Hypotheses 1 and 2; we find no evidence that the mandatory break affects overall data quality.
Straightlining
For 11-item questions after the break, the “straightlining” measure ranges from 1 (no repetition) to 11 (entirely consistent).[1] Results from the t-test show that there is no statistically significant difference (p = 0.11) in the frequency of repeated same responses, regardless of control (M = 5.75, SD = 2.61) or experimental group (M = 5.57, SD = 2.46).
In addition to a t-test, a Poisson regression was also analyzed. Poisson regression is a generalized linear model that predicts the expected “count” of the outcome, rather than normally distributed or binary outcomes (Hoffmann 2016, 136). Considering that the straightlining measure is in a discrete “count” form, a Poisson regression is considered appropriate. The results show that the experimental group is associated with a 3.10% decrease in the expected number of repeated answers. However, the coefficient is not statistically significant (p = 0.09). In other words, based on both t-test and Poisson regression results, no evidence is found to support that a break can decrease the frequency of careless, repetitive responses. Therefore, the findings from the straightlining measure also show no support for either hypothesis.
Self-Reported Honesty
The experimental group has a mean self-reported honesty score of 98.09 (on a scale from 0 to 100) and a standard deviation of 6.97. For the control group, the honesty score had a mean of 98.36 and a standard deviation of 5.89. No statistical significance was observed in the difference between those who had and did not have the mandatory break (p = 0.35), supporting neither Hypothesis 1 nor 2.
Attention Check
The attention check was shown to the respondent after policy item #14, which was three questions after the mandatory break was given to the experimental group (after policy item #11). Results show that those who received the break (the experimental group) had a higher rate of failing the attention check (5.35%) than the control group (4.65%). However, the proportion test shows that there is no statistically significant difference in the failure rates between the two groups (p = 0.46), providing no support for either Hypothesis 1 or 2.
Sociodemographic Data Inconsistency
We assessed response accuracy by comparing the experimental and control groups, focusing on the inconsistency in self-reported sociodemographic data between participants’ Prolific profiles and their responses in our survey. The three sociodemographic variables that we examined that had the most complete data supplied by Prolific were the respondents’ gender, age, and race and ethnicity. A discrepancy in any of these three variables for each respondent was recoded into a dichotomous variable.
Minor differences emerged between the experimental and control groups in the accuracy of their self-reported demographic data. 5.44% of the experimental group showed inconsistency between Prolific and survey responses, whereas 4.74% of the control group showed inconsistency. The proportions test shows that there is no statistically significant difference between the experimental and control groups (p = 0.47), thus providing no support for either hypothesis.
All five measures of data quality show that there is no statistically significant difference between the control group and the experimental group. At the full sample level, the mandatory breaks did not affect the data quality among the respondents.
Conclusion
This study sought to determine whether a mandatory break could enhance data quality in web surveys administered to online panels. Building on previous research that used four measures of data quality (Peer et al. 2021), we used five measures to assess the effectiveness of this intervention. Overall, our results show that the mandatory break provided no consistent benefit (or loss) to data quality within the context and configuration of our survey.
For practitioners designing surveys of comparable length and content, the practical implication is straightforward: adding a one-minute mandatory break is unlikely to be a worthwhile use of resources, since it raises costs and respondent burden without improving any of the data-quality indicators we examined. Effort is better directed toward established practices to increase data quality, such as by keeping the questionnaire concise, by not including the mandatory break, and by using the bonus money to increase the pay rate. Other conditions and configurations of an experiment might yield different results (e.g., different survey subjects, varying survey lengths, different break lengths, different target populations).
In addition, we speculate that demographic factors may moderate the effect of mandatory breaks. Prior work suggests that gender variations exist in task quality outcomes under various conditions that alter the effort or attention a task demands, including time pressure (Bilancini et al. 2024), noise (Abbasi et al. 2022), and psychosocial stress such as induced competition (Cahlíková et al. 2020). If mandatory breaks similarly reshape the demands of the survey response, their effects may likewise differ across respondents. Including gender, demographic characteristics (e.g., race and ethnicity, age, educational attainment) may therefore influence how respondents respond to mandatory breaks. A systematic examination of such moderators was beyond the scope of the present study, but it is a promising direction that we invite future researchers to pursue.
Although our overall findings show null or negative results, future studies can test whether offering optional breaks, varying break duration, or informing participants in advance may alter our findings. Previous literature shows that the effect of introducing breaks in tasks can depend on other factors, such as the break timing, duration, frequency, or break content (for review, Albulescu et al. 2022), and survey design more broadly requires a tailored rather than one-size-fits-all approach (Dillman et al. 2014).
As web-based data collection expands, improving instrument design remains a central objective. Rather than implementing mandatory breaks indiscriminately, researchers should test alternative break strategies across conditions while remaining attentive to group-level sensitivities. We suspect mandatory breaks may still prove beneficial in certain survey contexts.
Corresponding author contact information
Seon Yup Lee
seonyupl@k-state.edu
828 Mid-Campus Drive, Manhattan, KS 66506
