1. Introduction
Survey researchers are increasingly asked to address questions that are conceptually complex and empirically dynamic. Projects often begin with broad theoretical or policy motivations but quickly require multiple forms of evidence to develop and test those ideas. Designing a research process that coherently integrates literature, qualitative insights, and quantitative evidence can therefore be as important as the findings themselves. In this article, we present a methodological roadmap for building that connection, moving systematically from a literature review to qualitative exploration and then to quantitative validation. This roadmap is intended as a transferable approach that other researchers can adapt to their own substantive questions.
The approach developed here consists of two complementary components that together expand the methodological toolbox for survey research. The first focuses on translating themes derived from literature and qualitative interviews into measurable survey constructs. This stage demonstrates how qualitative insights can inform instrument design, operationalize abstract concepts, and strengthen content validity before data collection. The second component addresses a persistent challenge in survey research: drawing credible and externally valid conclusions from small convenience samples. By benchmarking survey results against an existing representative reference sample, we show when and how carefully structured statistical analysis can improve generalizability even when probability-based data collection is not feasible.
Although we illustrate this approach using an applied study of organizational change in the federal government, the contribution of this paper is primarily methodological. Throughout the article, we outline transferable procedural steps, data integration strategies, and analytic checks for research settings where surveys intersect with qualitative inquiry or sampling options are limited.
2. Research Design
We employed a three-phase mixed methods design intended to connect conceptual construct development with empirical validation. The goal was to illustrate how researchers can move from identifying themes in existing knowledge to developing and testing measurable constructs using qualitative and quantitative techniques.
First, we conducted an extensive literature review to conceptualize the focal construct and identify dimensions that potentially define it, using relevant keywords. Because such constructs are often multidimensional and interpreted differently across disciplines, we reviewed both academic and gray literature. This step enabled us to synthesize existing perspectives, identify conceptual gaps, and outline an initial set of subconstructs and relationships that then guided our interview guide and our thinking about a framework to measure the primary latent construct.
Second, we refined and operationalized the framework derived from theory. Through semi-structured interviews with subject matter experts (SMEs), we explored how the construct manifests in real-world settings, how practitioners understand its dimensions, and what contextual factors shape its interpretation. These qualitative insights allowed us to revise, expand, and structure the framework in a way grounded in both lived experience and theoretical guidance.
Third, we conducted a quantitative validation study to empirically assess the framework. Developing a structured web survey from the refined dimensions and indicators, we collected data from individuals who match our target population. The purpose of this phase was to examine the reliability and validity of the proposed framework and to determine whether the theorized dimensions coherently capture the underlying latent construct.
Together, these three components: a theory-driven literature review, qualitative interview for framework development, and a web survey for quantitative validation, form a coherent mixed-methods strategy that enables the rigorous conceptualization, construction, and empirical assessment of a latent construct that would otherwise be difficult to measure directly.
In this study, we applied this mixed-methods approach to examine how artificial intelligence (AI) is understood, governed, and integrated within U.S. federal agencies. We first synthesized existing academic and gray literature on AI adoption in federal agencies to identify the core dimensions (Fluency, Adoption, Adaptation) relevant to assessing its organizational impact. We then refined these dimensions through qualitative interviews with SMEs involved in federal AI initiatives. Finally, we developed and administered a web-based survey to validate the resulting framework and evaluate perceptions around operational and organizational effects of AI across the federal population.
3. Qualitative Phase: Literature Review and Expert Interviews
In the qualitative phase, we focused on clarifying key concepts and developing a framework that we could later test with survey data. Our goal in this phase was not to produce generalizable estimates, but to move in a structured way from existing work to concrete dimensions and candidate measures that would define our latent construct of interest.
3.1. Literature Review
We began with a targeted review of academic and gray[1] literature around adoption of AI in the public sectors and landscape of AI usage in the government using a standard set of relevant keywords as search terms and iteratively adding more terms as they emerged since we were interested in measuring impact of AI on federal agencies. We used Google Scholar, Scopus, Policy Commons and SSRN along with conventional web-based search engines (Google and Bing) to trace the literature. We recruited graduate assistants to code the key findings from the identified literature into a corpus, that we later synthesized to identify the core themes around our primary topic of interest – impact of AI adoption in public and private sectors. We narrowed down to four core themes that can shape the idea of ‘impact’ in our case. See Appendix for details. The resulting narrative review highlighted the importance of measuring the impact of AI adoption on government agencies more holistically as opposed to focusing on segregated aspects of AI adoption. Thus, we developed a theoretical framework that connects these identified themes overall.
3.2. Expert Interviews
Next, we conducted semi-structured interviews with SMEs to refine and operationalize the themes we saw in the literature. We interviewed Chief Information Officers (CIO) in federal government positions who had direct experience with the development, implementation, or oversight of these tools in federal or related settings. We recruited participants using snowball sampling, starting from contacts in our professional networks and asking interviewees to recommend additional experts. Interviews took place between Fall of 2022 and Summer of 2023, were conducted by two of the authors, lasted about minutes and were held via videoconferencing. We used an interview guide focusing on specific projects within participants’ agencies that involved AI-based solutions as case studies. The guide elicited detailed information on their roles in planning and implementation, the challenges encountered, and their overall perceptions of the use of AI within their agencies (full guide in Appendix).
We audio-recorded all interviews with participants’ consent and then transcribed them. We analyzed the transcripts using a thematic approach. We first created a set of deductive codes based on the concepts from the literature, and we then added inductive codes that emerged as we read the interviews. Two members of the research team worked together on the codebook, double-coded an initial set of transcripts, and resolved differences through independent inputs from a third member of the research team. We then grouped related codes into broader themes that helped us see how the literature-based concepts played out in practice and where experts’ views suggested adjustments or additions.
We used the combined insights from literature and the interviews to define a small set of core dimensions that describe our latent construct. We then arranged these dimensions into a provisional framework that specifies how they relate to one another. For each dimension, we identified subdimensions and concrete indicators that could be asked about in a survey[2], and we drafted items that reflected the language experts used in the interviews. The specific details regarding the definitions of the dimensions, and the conceptualized framework in our example can be found in the Appendix.
Our qualitative sample was small: we interviewed experts due to resource constraints. This limits what we can say about how common any specific view is in the wider expert community. Studies on thematic saturation (Hennink and Kaiser 2022; Hennink et al. 2017) suggest that samples of this size are often sufficient to identify many of the main themes in a relatively focused domain, but they are less suited to fully exploring all nuances or rare perspectives. For this reason, we treated these interviews as an exploratory step that helps us define and organize key dimensions, rather than as a standalone source for generalizable conclusions.
The interviews still played an important role in our design. They allowed us to connect ideas from literature to the language and experiences of practitioners, and they helped us turn broad concepts into concrete dimensions and candidate indicators. We then used the survey to examine how well the framework holds up when empirically measured with survey data.
4. Quantitative Phase: Web Survey
In the quantitative phase, we used a web survey to assess the framework that we developed from literature and interviews. We designed the survey instrument[3] to mirror the structure of the framework. For each core dimension, we created multiple items that reflected the subdimensions identified in the qualitative phase and used wording that was close to the language experts used in the interviews. We also included items to capture demographic characteristics of the respondents.
The target population for the survey was federal employees in the US (including consultants and contractors to the government). We recruited the participants through professional events[4], listservs[5], and individual networks. The survey was conducted online in November 2024 using Qualtrics and remained open for three months. In total, respondents participated, out of which we received completes. Due to the small number of consultants and contractors, we excluded them from further analysis, restricting the target population to current and former federal employees as of November 2024 Our sample had the two following features:
-
There was no sampling frame, i.e., the selection probability for the individuals in the target population was unknown.
-
The selection probabilities for certain individuals in the population could be zero.
These precluded us from traditional design-based population inference. Hence, we combined our survey with a reference probability survey, Federal Employees Viewpoint Survey (FEVS)[6] that is large enough and is designed to represent our target population (Valliant 2020). This illustrates one way survey researchers can leverage auxiliary information to strengthen conclusions from small or nonprobability samples (Chen et al. 2020; Kim et al. 2021; Schonlau et al. 2007). Table 1 shows the demographic distribution in FEVS and
FEVS includes the same core demographic variables as our survey, but it does not include our key outcomes of interest. Following Kim et al. 2021, we fit an ordinal regression model in our survey data that predicted the outcome(s) from the shared demographic variables to account for the ordered response scales. Under standard assumptions of positivity and transportability—that everyone in the population has non-zero probability of being included in FEVS and that the model fitted on our survey also holds in FEVS—we then used this model to predict the outcomes for all FEVS respondents. This “mass imputation” step created a synthetic outcome for each unit in the large representative probability sample for further inference. Refer to Appendix for the mathematical details.
However, our nonprobability survey is very small (n = 20) relative to the reference probability survey, which limits how far we can push population level inference based on mass imputation. With so few cases, the outcome model is noisy, and any misspecification can directly bias the projected estimates. Valid inference also depends on strong, hard to verify assumptions about self-selection and model transportability, so we treat our integrated estimates as illustrative outputs with clearly stated assumptions rather than definitive population statistics. Thus, the main contribution is to show, in a concrete example, how mass imputation can be implemented and what kinds of assumptions and limitations survey researchers should consider when combining a small nonprobability sample with a large probability survey.
5. Validating Conceptualized Framework
Based on the FEVS data with mass-imputed outcomes of our interest we conducted a confirmatory factor analysis (CFA) that sought to validate the theoretical framework we developed at the qualitative phase of this study. We treated the main dimensions in the framework as latent constructs, each measured by a set of survey questions. First, we specified separate measurement models for each dimension and estimated factor loadings to see how strongly each item relates to its intended factor. Then we combined these into a structural model that included the directional relationships among the latent constructs, reflecting the structure suggested by the qualitative work (in our case, from fluency to adoption to adaptation).
We compared two competing CFA models. We started with the framework developed in the earlier phases and then made theory-driven modifications, such as adding covariances between item residuals when the earlier qualitative findings suggested that two items are closely related. We used standard fit indices to decide our final model, as shown below:
In our final CFA model, the observed indicators loaded meaningfully onto the three latent dimensions—Fluency, Adoption, and Adaptation—supporting the proposed measurement structure. We can interpret this based on the factor loading estimates and latent covariances obtained from the summary of the CFA model. Fluency showed generally strong loadings, indicating a well-defined construct. Adoption displayed a more heterogeneous pattern, reflecting variability in how different adoption-related indicators contribute to the construct. Adaptation was characterized by strong overall loadings, capturing organizational responses to AI implementation, while also reflecting perceived shifts in strategic focus. Finally, Fluency, Adoption, and Adaptation are all strongly associated, indicating that the three latent dimensions are closely related, though the directions of these associations vary.
6. Implications
This study shows how a mixed methods design can help survey researchers move from broad ideas to tested measures in a structured way. By starting with a focused literature review, using a small set of expert interviews to refine and name key dimensions, and then converting those dimensions into survey items, we demonstrate a practical path from concepts to an empirically tested framework.
Our approach illustrates an exploratory sequential design in which qualitative work does not stand alone but feeds directly into survey development and validation. The interviews help us identify language, subdimensions, and relationships that matter in practice; the survey later lets us test whether these hold up empirically. Using confirmatory factor analysis, we show how one can assess whether items derived from qualitative insights form coherent scales and represent the intended latent constructs.
A second contribution is our use of mass imputation to link a small nonprobability survey with a large reference probability survey, illustrating how this strategy can attenuate selection bias; however, one should be cautious with the sample size of the receiver survey before making valid population-level inferences. Taken together, these steps, i.e., literature to interviews to framework development, framework to survey items, and nonprobability survey data linked to a probability survey for further analyses, offer a set of concrete procedures that other survey researchers can adapt to design and strengthen survey-based research in settings with constrained samples, or emerging constructs that have not yet been well defined or measured.
Author Contribution
Ujjayini Das and Srijeeta Mitra contributed equally to the planning and execution of this research. They share joint first authorship on this article. Tejwansh S. Anand as Principal Investigator and Bertrand A. Stoffel as Co-Principal Investigator supported at different stages of the work.
Conflict of Interest
The authors declare no conflict of interest.
Statement on Funding
This work was supported through the Deloitte Initiative for Artificial Intelligence and Learning (DIAL) at the University of Maryland for the academic year 2023–2024.
Acknowledgments
We extend our gratitude to the research team at Robert H. Smith School of Business for contributing to background research, with a special mention to Ishita Joshi, for her assistance in initial qualitative coding.
Corresponding author contact information
Ujjayini Das
Email: ujstat@umd.edu
Address: 1218 Lefrak Hall, 7251 Preinkart Drive, College Park, MD 20742
Srijeeta Mitra
Email: smitra98@umd.edu
Address: 1218 Lefrak Hall, 7251 Preinkart Drive, College Park, MD 20742

