Introduction
Skin color is a key element of how racial identity is perceived by self and others and may be the basis for discrimination that affects a wide range of health and wellbeing outcomes (Caraballo-Cueto and Godreau 2021; Dixon and Telles 2017; Monk 2021; Perreira and Telles 2014; Telles 2014). Although visual skin color scales have improved standardization and reliability in recent years, concerns remain about how respondents interpret these measures and whether interviewer-administered assessments may create discomfort or inconsistent reporting, particularly in culturally distinct contexts. These concerns are especially relevant for surveys conducted outside the continental United States, where local systems of racial classification and color terminology may differ substantially from U.S. racial frameworks.
These concerns extend beyond technical issues of reliability and standardization. Respondents may be sensitive to questions involving skin color and interviewer classification because doing so reproduces a social dynamic of being evaluated on the basis of one’s race and appearance (Enders and Thornton 2022; Hannon and DeFina 2016). Researchers have therefore questioned whether respondents may perceive such measures as uncomfortable, inappropriate, or stigmatizing—particularly in populations with histories of racial discrimination and color-based inequality (Hannon and DeFina 2016; Lor et al. 2017; Telles 2014).
Cognitive interviewing provides a useful approach for evaluating how respondents interpret and answer sensitive sociodemographic questions before large-scale survey implementation. Beyond assessing comprehension, cognitive interviewing can identify sources of hesitation, ambiguity, or interpretive variation that may affect administration and data quality (Beatty and Willis 2007; Willis 2005). This approach is particularly important for measures related to race and skin color, where local meanings and social histories may shape how respondents understand survey items.
This study evaluates the Spanish-language Yadon–Ostfeld Skin Color Scale (Y–O Scale) during pretesting for the Puerto Rico Panel Study of Income Dynamics (PR-PSID). The Y–O Scale is a 10-point visual measure designed to improve consistency in skin color measurement through standardized tonal gradation. Although the scale has been used in U.S.-based studies, it has not been evaluated in Puerto Rico, where locally grounded color terminology coexists with federally mandated U.S. racial classification systems.
Using cognitive interviews with 21 adults in Puerto Rico, we examine three questions relevant to survey practice: (1) whether participants understood and could readily use the scale, (2) whether the item generated discomfort or interpretive difficulty, and (3) what implementation considerations emerged for interviewer-administered skin color measurement. Findings contribute practical evidence regarding the acceptability, reliability, and administration of skin color measures in culturally distinct survey contexts.
Background
Skin color in Puerto Rico reflects a long history of colonialism, racial mixture, and colorism—the privileging of lighter skin tones in social and economic life—that continues to shape how color is perceived and experienced today (Caraballo-Cueto and Godreau 2021; Dixon and Telles 2017; Telles 2014). Local descriptors such as blanco, trigueño, and moreno reflect fluid social categories that combine skin tone, hair texture, and ancestry—and do not map neatly onto U.S. federal racial classifications. This divergence has longstanding consequences for survey measurement: in the 2020 Census, 75 percent of Puerto Ricans identified as “Some Other Race,” either alone or in combination (Figueroa-Lazu et al. 2022). Recent initiatives, including Puerto Rico Law No. 24 of 2021, have further highlighted the need for improving measures to assess racial inequality in Puerto Rico (Díaz-Torres et al. 2024; Puerto Rico 2021).
The Yadon–Ostfeld Skin Color Scale (Y–O Scale) is a standardized 10-point instrument that uses calibrated hand images to represent evenly spaced skin tones along a light-to-dark gradient (Ostfeld and Yadon 2022; 2023; see Figure 1). Developed using spectrophotometer-validated color calibration, the scale addresses limitations of earlier skin color palettes, including uneven tonal gradation and limited reliability (Gordon et al. 2022; Hannon and DeFina 2016). It has since been applied in research on racial ideology, Latino identity, discrimination, and law enforcement (Ostfeld and Yadon 2022; Pew Research Center 2021; Talbert and Matos 2024). Despite its growing use, the scale was developed in English for U.S. survey contexts and has not been evaluated in Puerto Rico. Administering it in Spanish in this distinct racial and linguistic setting makes an assessment of its cultural fit particularly important. This study addresses that gap.
Method
Study Design and Context
This study was conducted as part of the Puerto Rico Panel Study of Income Dynamics (PR-PSID), a longitudinal survey examining demographic, economic, and social conditions in Puerto Rico. We used cognitive interviews with concurrent probing—including both scripted and spontaneous probes—to evaluate the performance of the Y–O Scale in an interviewer-administered format. The broader cognitive interview protocol included modules on skin color, Spanish–English language use, and prior U.S. migration experience; this analysis focuses exclusively on the skin color measure.
Sample and Recruitment
We conducted 21 cognitive interviews with Spanish-speaking adults residing in Puerto Rico between November 20 and December 5, 2022. Volunteers were recruited island-wide through flyers, in-person outreach, and social media and email invitations. Those who completed the online screening survey constituted the recruitment sample (n = 170; see Table 1). The survey, hosted by Brown University, collected sociodemographic information including self-reported skin color (on a 1–6 scale), marital status, U.S. migration experience, and English-language skills (see Appendix A for the skin color screening item). Eligibility was limited to adults aged 21 or older currently residing in Puerto Rico. From the clean sample of eligible volunteers (n = 112; see Table 1), we selected participants to maximize variation across these characteristics, with the goal of reaching a target sample of 20–24 participants. To minimize no-shows, we prioritized participants residing near the two interview sites in Cayey or San Juan. Of the 44 selected participants, 24 were contacted and scheduled, of whom 21 completed the interview. Digital recruitment and professional networks likely drew volunteers with higher educational attainment than the broader Puerto Rican population: most participants held college or graduate degrees; ages ranged from 21 to 70 or older; and 81 percent had lived in the United States for at least three months (see Table 2).
Procedures
Before beginning the rating tasks, the interviewer introduced the scale by reading a standardized explanation clarifying that the item referred to visible skin color and asking participants to select the number that most closely matched their skin tone (see Appendix A for the full introduction and probes). Interviews were conducted in private rooms at the University of Puerto Rico–Cayey and Hispania/Ipsos facilities in San Juan. Each session lasted 60–90 minutes, with the skin color module administered first (approximately 2–6 minutes). Participants received $50 compensation. The skin color module comprised three sequential rating tasks using a laminated 10-point Y–O Scale: the interviewer first rated the participant’s skin color independently and did not disclose this rating to the participant before or during the self-assessment, ensuring that the two ratings were made independently; the participant then self-rated; and partnered participants rated their spouse or partner. After selecting a number, participants described their chosen color using their own terms. Scripted and spontaneous probes were used to explore comprehension, hesitation, and interpretations of the item (see Appendix A). The first author—a Puerto Rican who self-identifies as White within the local racial context—conducted all interviews; her positionality may have influenced rapport and how participants articulated skin color experiences (Enders and Thornton 2022).
Analysis
Interview transcripts were coded using deductive and inductive methods within a matrix framework to identify issues related to comprehension, response processes, identity, colorism, and item interpretation. Coding decisions were reviewed with the research team to support analytic consistency. We generated descriptive statistics and calculated an intraclass correlation coefficient (ICC) to assess agreement between interviewer- and self-rated skin color (Koo and Li 2016). Although the sample size was small, ICC measures inter-rater agreement rather than population representativeness and is therefore appropriate for evaluating consistency between paired ratings; we report the 95% confidence interval to reflect the precision of the estimate given the small sample (Bonett 2002). Spanish-language quotations were translated using DeepL and reviewed against the original.
Findings
Did participants readily understand and use the Y–O Scale?
Participants generally found the Y–O Scale easy to use and understood the task without difficulty, consistent with the standardized introduction the interviewer delivered at the start of the rating tasks. Most described the visual gradient as intuitive and selected a response category readily. Several participants physically compared their hand to the scale, and one likened the process to selecting a makeup shade. The phrase color de piel (“skin color”) was familiar and not perceived as offensive or inappropriate. Mean ratings were identical across self-, proxy-, and interviewer assessments (4.6), and self- and interviewer ratings showed similar variability (see Table 3).
Some participants hesitated between adjacent categories, particularly when skin tone fell between two values or varied due to sun exposure. Others noted that locally meaningful descriptors such as trigueña or café con leche did not correspond neatly to a single numeric value. One participant described the numeric format as unfamiliar: “It’s not that I’m confused, but the question requires something that is not natural for me—to place myself as just a number.” These reflections indicate that participants understood the task but sometimes experienced the numeric format as a limiting way to describe their skin color—a feature of the measure rather than a comprehension problem.
For practice: Interviewers should be trained to normalize hesitation between adjacent tones as an expected feature of the task, not a sign of confusion. Scripted language encouraging participants to “select the closest visual match” helps maintain standardization while reducing uncertainty. Interviewers should also accept—without redirection—locally meaningful descriptors that participants offer alongside their numeric choice; these enrich the data and reflect the social salience of skin color rather than any difficulty with the item.
Were interviewer- and self-ratings consistent?
Interviewer and participant self-ratings showed high agreement (ICC[2,1] = 0.84, 95% CI = 0.64–0.93, p < .001; n = 21), indicating strong consistency in how the scale was applied across raters. Following Gordon et al. (2022), who use ICC to assess agreement between interviewer and participant self-ratings of skin color, this estimate falls within the range characterized as excellent inter-rater reliability (ICC > 0.75; Cicchetti 1994). This is particularly relevant given documented concerns about inconsistent classification in interviewer-administered skin color measures (Gordon et al. 2022; Hannon and DeFina 2016). The sample size is modest, and the confidence interval should be interpreted with that limitation in mind — smaller samples yield less precise ICC estimates (Bonett 2002). Gordon et al.'s (2022) in-person validation study included 46 participants and likewise called for replication in larger samples. This finding provides initial evidence of strong alignment and will be replicated once PR-PSID survey data collection is finalized.
Proxy reporting for spouses or partners was generally straightforward. One participant noted that a reddish undertone (colorao) was not represented on the scale—a minor limitation in tonal coverage that did not impede task completion but points to an area for further scale development.
For practice: These initial inter-rater reliability findings support the use of the Y–O Scale in interviewer-administered survey contexts, pending replication in larger samples. Survey teams should invest in structured interviewer training that standardizes how the scale is presented and how hesitation is managed. Consistent color calibration across print and digital formats is also essential: participants attended closely to subtle tonal differences, and variation in color rendering across formats could introduce measurement error.
Did the item generate discomfort or interpretive difficulty?
The Y–O Scale did not generate discomfort or refusals. Participants engaged willingly with the task and, when probed, described three main interpretations of the item’s purpose: capturing self- and other-perceptions of skin color (n = 8), addressing color-based inequality and prejudice (n = 8), and documenting racial diversity for research or policy purposes (n = 4). These interpretations reflect the social salience of skin color in Puerto Rico rather than difficulty with the measure itself, as participants drew spontaneously on culturally grounded and meaningful local constructs—trigueña, cremita, café con leche, jincha—when explaining their choices (see Appendix B for a full list of descriptors used).
A subset of participants with prior U.S. migration experience (n = 6 of 17) associated the item with U.S. Census racial classification systems, describing frustration with federally imposed racial categories. One described such items as un signo de molestia—a source of irritation with political overtones. Importantly, these associations emerged during probing, not during task completion: even participants who referenced U.S. racial frameworks completed the rating task consistently and without difficulty. In contrast, participants without U.S. migration experience (n = 4) drew on Caribbean color vocabularies and did not invoke federal classification systems. These patterns are consistent with the framework developed by Goerman et al. (2025), who show that cultural effects in survey measurement arise not only in translation but also in how participants situate items within broader social and classificatory systems. Table 4 summarizes cultural matches and interpretive mismatches identified across the interviews, distinguishing between measurement processes that functioned well and interpretive contexts that varied across participants. Interpretive variation reflects how participants situate the measure within their own racial and migration histories—not a problem with the scale’s design.
For practice: A brief introductory statement—for example, “This question asks about the color of your skin as it appears visually, not about your racial identity or background”—can reduce associations with U.S. Census frameworks, particularly for participants with prior U.S. migration experience. This is a low-cost, high-value protocol refinement. Question-by-question (QxQ) interviewer guides should anticipate that participants may connect the item to broader experiences of racial classification and should equip interviewers to acknowledge such responses neutrally before returning to the rating task.
Conclusion
This study provides empirical evidence addressing longstanding concerns about the feasibility and acceptability of interviewer-administered skin color measures (Hannon and DeFina 2016; Lor et al. 2017). Using cognitive interviews conducted during PR-PSID pretesting in Puerto Rico, we found that participants readily understood and completed the Yadon–Ostfeld Skin Color Scale and that interviewer- and self-ratings showed strong agreement. The sensitivity concerns that have prompted caution about skin color measurement did not materialize: no participants refused the task, discomfort was not observed, and interpretive variation—where it occurred—did not interfere with reliable scale use.
These findings demonstrate the value of cognitive interviewing as a tool for evaluating sensitive sociodemographic measures across culturally distinct settings, distinguishing between interpretive variation that reflects local social meanings and problems that would affect measurement performance. The findings also support further investigation into the implementation of the Y–O Scale in Puerto Rico and in similar Caribbean and U.S. Latino survey contexts. Practical refinements identified here, including structured interviewer training, brief introductory language, QxQ guidance, and consistent color calibration across formats, are low-burden steps that can strengthen data quality in future survey implementations. The sample was relatively small and skewed toward higher educational attainment, which may have made the task easier to complete and the reliability estimates more favorable than would be found in a broader sample. Future studies should include broader variation in educational background to assess whether findings hold across a broader range of educational backgrounds.
Acknowledgments
We thank the Institute of Interdisciplinary Research at the University of Puerto Rico–Cayey, Hispania/Ipsos in San Juan, and the Population Studies and Training Center at Brown University for institutional and logistical support. We are grateful to Orlando Maldonado Meléndez for assistance with recruitment and data collection, and to all participants for their time and insights.
Ethical Considerations
The Brown University Institutional Review Board approved the study (Protocol 2022003440) on November 1, 2022. The University of Puerto Rico–Cayey also reviewed and supported the study.
Consent to Participate
Participants reviewed study information and indicated consent during the online pre-screening survey. At the interview, the interviewer reiterated the study purpose, procedures, and participants’ rights and obtained verbal consent, which was audio-recorded.
Declaration of Conflicting Interest
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding Statement
This research was supported by the Russell Sage Foundation (Award #2111-35014, “Demographic Representation and Income Stratification in Puerto Rico”; Amount: $174,995; Percentage: 90%), which supplements the main project supported by the Eunice Kennedy Shriver National Institute of Child Health and Human Development (NICHD; Award #R01HD106069, “Puerto Rico Panel Study of Income Dynamics”; Amount: $3,333,315; Percentage: 5%). Both awards were granted to Elizabeth Fussell (Brown University) and Narayan Sastry (University of Michigan). Infrastructure support for the project was provided by the Population Studies and Training Center at Brown University, funded by NICHD (Award #P2CHD041020; Amount: $235,565; Percentage: 5%). The content is solely the responsibility of the authors and does not represent the official views of the National Institutes of Health.
Corresponding author contact information
Theresa Thompson-Colón, Independent Survey Methodologist, Montréal, QC, Canada, tthompson.colon@gmail.com.