Introduction
Determining what is a good sample size for qualitative research is one of the most debated yet critical decisions a researcher makes during the design phase. Unlike quantitative studies, where statistical power calculations dictate a specific number of participants to ensure generalizability, qualitative research operates on a fundamentally different logic: the pursuit of depth, richness, and theoretical saturation rather than statistical representation. Even so, there is no single magic number, no universal formula, and no standard table to consult; instead, the "right" sample size is contingent upon the research question, the chosen methodology, the heterogeneity of the population, and the resources available. Understanding this nuance is essential for producing credible, trustworthy findings that withstand academic scrutiny and provide genuine insight into the human experience.
Detailed Explanation
At the heart of qualitative sample size determination lies the concept of data saturation (often used interchangeably with theoretical saturation in grounded theory). In practice, saturation occurs when additional data collection no longer yields new themes, insights, or codes relevant to the research question. This is a qualitative judgment call made during the research process, not a number calculated before it begins. Here's the thing — in practical terms, you have reached a good sample size when the interviews, focus groups, or observations start to feel repetitive—when you can accurately predict what the next participant will say because the thematic landscape has been fully mapped. Because of this, qualitative researchers typically employ purposive sampling—selecting participants deliberately based on their ability to illuminate the phenomenon under study—rather than random probability sampling.
This is where a lot of people lose the thread.
The required sample size varies drastically depending on the qualitative tradition employed. Because of that, a phenomenological study seeking to understand the "lived experience" of a specific event (e. On top of that, g. , surviving a natural disaster) often requires a very small, homogeneous sample—typically 5 to 10 participants—because the goal is deep idiographic understanding of a shared essence. Even so, conversely, grounded theory studies, which aim to generate a broad theoretical framework, often require larger samples (20–60+ participants) to compare incidents across diverse groups and refine theoretical categories. Consider this: Ethnography usually involves a single "case" (a community or organization) but requires prolonged engagement with dozens of informants over months or years. Practically speaking, Case study research might focus on just one or a few cases, but the depth of data per case is immense. That's why, defining a "good" sample size is impossible without first anchoring it in the specific methodological approach.
Step-by-Step Concept Breakdown
To determine an appropriate sample size for your specific project, follow this conceptual workflow rather than looking for a fixed integer:
1. Define the Scope and Specificity of the Research Question
Broad, exploratory questions (e.g., "How do teachers experience curriculum change?") generally require larger, more diverse samples to capture variation. Narrow, specific questions (e.g., "How do novice physics teachers experience the first week of implementing a specific inquiry-based module?") require smaller, highly homogeneous samples. The narrower the scope, the fewer participants needed to reach saturation.
2. Assess Population Heterogeneity
If your population is highly heterogeneous (e.g., "healthcare professionals" including doctors, nurses, admins, and technicians across urban and rural settings), you need a larger sample to capture the diversity of perspectives. If the population is homogeneous (e.g., "first-time mothers aged 25–30 who delivered via C-section at a specific hospital"), saturation is reached much faster because experiences overlap significantly Practical, not theoretical..
3. Select the Qualitative Methodology
Align your sample size expectations with your methodology:
- Phenomenology / Narrative: 5–15 participants.
- Grounded Theory: 20–60 participants (iterative sampling).
- Ethnography: 1 cultural scene / 20–50+ informants.
- Case Study: 1–4 cases (with deep data per case).
- Content / Thematic Analysis: 10–30 interviews (often cited as a sweet spot for thematic saturation).
4. Evaluate Data Quality and Depth
A "good" sample size is not just about quantity but quality. Ten rich, hour-long interviews with articulate, reflective participants who have deep experience with the phenomenon are infinitely more valuable than thirty superficial, 15-minute interviews with disengaged participants. If your data collection method yields "thick description" (detailed context, emotion, non-verbal cues), you need fewer participants. If data is thin (e.g., short email responses), you need more.
5. Plan for Iterative Analysis
Qualitative sampling is sequential, not simultaneous. You do not recruit 20 people at once. You recruit 3–5, transcribe, analyze, identify gaps, and then recruit the next batch based on theoretical needs (theoretical sampling). This iterative loop is the mechanism for determining when to stop. Your sample size is finalized only when you declare saturation.
Real Examples
Consider a study exploring patient experiences of telehealth during the pandemic. Practically speaking, * Scenario A (Phenomenology): The researcher wants to understand the essence of the telehealth encounter for elderly patients with chronic conditions. That's why they recruit 8 participants (homogeneous: all over 70, all managing diabetes, all using the same platform). After 6 interviews, themes stabilize (technological anxiety, loss of physical touch, convenience). Interviews 7 and 8 confirm these themes with no new codes. Plus, saturation reached. Sample size = 8. Here's the thing — this is a "good" sample size for this design. * Scenario B (Grounded Theory): The researcher wants to develop a theory of digital health literacy adoption across the lifespan. Now, they need maximum variation: teens, working parents, elderly, urban, rural, high/low income. Also, they conduct 35 interviews. That said, saturation for the core category "trust calibration" occurs around interview 25, but they continue to 35 to test the theory against negative cases (people who refused telehealth entirely). Sample size = 35. This is a "good" sample size for this design Easy to understand, harder to ignore. Which is the point..
Now consider a focus group study on workplace culture. A single focus group (6–10 people) is rarely enough. Because of that, a good design might involve 4–6 focus groups (totaling 30–50 participants) stratified by department or tenure to capture group dynamics and dissenting views. Here, the "unit of analysis" is the group interaction, not the individual, changing the math entirely.
Scientific or Theoretical Perspective
The theoretical underpinning for qualitative sample size rests on information power (Malterud et al.On top of that, , 2016), a framework that challenges the saturation metaphor. Information power argues that the more information the sample holds relevant to the study, the fewer participants are needed. It identifies five determinants:
- Study Aim: Broad aim = lower information power per participant = larger sample needed. Consider this: narrow aim = higher information power = smaller sample. 2. Sample Specificity: Dense specificity (participants hold precise, deep knowledge) = higher power = smaller sample.
- So Established Theory: If strong theory exists (deductive), less data needed. If exploring unknown territory (inductive), more data needed.
- Quality of Dialogue: Strong rapport, open dialogue = high power = smaller sample. And 5. Analysis Strategy: Cross-case analysis (comparing cases) requires more cases than in-depth case analysis.
This framework provides a scientific rationale for justifying sample size in ethics proposals and peer review, moving the conversation away from "Is 10 enough?" to "Does this sample possess sufficient information power for this specific aim?" It validates small samples if the other four dimensions are high, offering a dependable defense against reviewers demanding quantitative-style power calculations.
Common Mistakes or Misunderstandings
Mistake 1: Confusing Qualitative "Saturation" with Quantitative "Representativeness"
Researchers often worry their
Researchers often worry their sample is not representative of the broader population, yet in qualitative inquiry representativeness is not the primary criterion for adequacy. Consider this: the central concern is whether the data furnish sufficient informational power to answer the research question, which depends on the dimensions outlined by Malterud and colleagues. This means a small, purposefully selected sample can be more appropriate than a larger convenience sample if the former meets the five information‑power criteria—particularly when the aim is narrowly focused, the participants possess deep, relevant knowledge, and a strong theoretical framework guides the inquiry.
Worth pausing on this one.
Mistake 1: Confusing Qualitative “Saturation” with Quantitative “Representativeness”
Saturation is a point at which new data no longer yield novel insights, not an indicator that the sample reflects the full spectrum of a population. But researchers sometimes treat saturation as a proxy for statistical generalizability, leading them to either stop data collection prematurely (if saturation appears early) or to collect excessively many interviews in an attempt to achieve a “representative” sample. In practice, saturation should be judged against the study’s specific objectives: if the aim is to explore a poorly understood phenomenon, saturation may require more participants; if the aim is to illustrate a well‑defined concept, fewer interviews may be sufficient. The key is to align the saturation point with the information‑power determinants, not with an arbitrary numeric threshold Most people skip this — try not to..
Mistake 2: Relying Solely on Saturation as the Sample‑Size Decision Rule
Building on the previous point, an over‑reliance on saturation can mask other critical considerations. Here's a good example: a researcher may stop after 12 interviews because no new themes emerge, yet if the sample lacks variation in a crucial demographic (e.g., only middle‑class participants when studying health disparities across income levels), the findings will be biased. Information power reminds us to evaluate the breadth of variation, the depth of expertise, and the richness of dialogue, not merely the point at which thematic novelty wanes It's one of those things that adds up..
Mistake 3: Ignoring the Role of Negative or Deviant Cases
A common oversight is to collect only data that confirm the emerging theory, thereby limiting the study’s ability to test its boundaries. Now, including participants who deviate from the predominant pattern—such as those who refuse digital health tools despite high digital literacy—enhances the robustness of the analysis. Failing to do so can lead to over‑generalized conclusions and weaken the credibility of the theory being developed Worth keeping that in mind..
Mistake 4: Using Convenience Sampling Without Articulating Its Rationale
Convenience sampling (e., recruiting participants from a single clinic or a single workplace) is often justified by logistical constraints, but when presented without explicit justification it can be perceived as methodological weakness. Practically speaking, g. g., they are information‑rich, can articulate the phenomenon in depth, and collectively cover the variation needed for the study’s aims). That's why to defend a convenience sample, the researcher must demonstrate that the selected participants possess the requisite information power (e. Transparent documentation of recruitment strategies and sampling decisions is essential for auditability and reviewer confidence The details matter here. Surprisingly effective..
Mistake 5: Failing to Update the Sample Size Justification in Response to Emergent Insights
Qualitative studies are iterative; new insights may reveal gaps in the initial sampling strategy. Practically speaking, for example, early interviews might indicate that a particular subgroup is under‑represented, prompting the researcher to add additional participants to achieve balanced variation. If the original sample‑size rationale is not revisited and revised, the final study may still suffer from insufficient information power.
Conclusion
The determination of an appropriate sample size in qualitative research transcends simplistic rules of thumb or the pursuit of statistical representativeness. And rooted in the information‑power framework, sample size must be evaluated against the study’s aims, the specificity and depth of the participant pool, the strength of the existing theoretical context, the quality of interpersonal engagement, and the chosen analytical strategy. In practice, by consciously addressing these dimensions—and by avoiding common pitfalls such as conflating saturation with representativeness, over‑relying on saturation thresholds, neglecting deviant cases, presenting convenience samples without justification, and ignoring iterative sampling adjustments—researchers can substantiate their methodological choices with a rigorous, theory‑driven rationale. This approach not only strengthens the scholarly integrity of qualitative investigations but also enhances the trustworthiness and impact of the insights they generate Not complicated — just consistent. Less friction, more output..