September 12, 2026
How to Calculate Sample Size: Easy, Essential Formulas You Can Use
How many respondents do I need is one of the first questions graduate students ask at StatAce. Panels scrutinize this issue during proposal defense. It relates to sample size calculation during the defense.
Too small a sample leads to weak statistical power. Additionally, too large a sample wastes time and resources you cannot spare. This article reviews the formulas researchers actually use, when each applies, and how to avoid defense gaps. The discussion includes practical steps for planning.
Before You Calculate Anything: Answer These Questions
- Do you know your population size? If yes, and it’s finite and identifiable (e.g., all nursing students enrolled in a specific college this semester), you can use a finite population formula like Slovin’s or Cochran’s. If your population is unknown or effectively infinite (e.g., “all social media users in the Philippines”), you’ll need a different approach.
- What kind of study are you running? A descriptive survey, a comparison between groups, a correlation, or an experimental design each has different sample size logic.
- What margin of error and confidence level are you willing to accept? Most social science theses use a 95 percent confidence level and a 5 percent margin of error, but this should be a deliberate choice, not a default you never examine.
Formula 1: Slovin’s Formula (Most Common in Philippine Theses)
However, Slovin’s formula is the one most Filipino undergraduate and some graduate theses default to for sample size calculation.
It is simple and only requires knowing your population size.
n = N / (1 + N e²)
Where:
- n = required sample size
- N = total population size
- e = margin of error (expressed as a decimal, e.g., 0.05 for 5 percent)
Worked example: Suppose your population is 850 students enrolled in a business program, and you’re using a 5 percent margin of error.
n = 850 / (1 + 850 × 0.05²) n = 850 / (1 + 850 × 0.0025) n = 850 / (1 + 2.125) n = 850 / 3.125 n = 272 (rounded up)
You would need 272 respondents.
A caution about Slovin’s formula: it’s popular because it’s easy. It’s not always the most defensible choice. Moreover, it assumes a homogeneous population. It doesn’t account for expected variability in your variable of interest. Additionally, many statisticians and thesis panels increasingly prefer Cochran’s formula. They also favor software-based power analysis for graduate-level work. If your adviser or panel expects a more rigorous justification, be ready to explain why you chose Slovin’s over alternatives. Or use one of the methods below instead.
Formula 2: Cochran’s Formula (For Larger or Unknown Populations)
Cochran’s formula is more statistically grounded and is often expected in graduate-level (especially doctoral) research.
Step 1, for an unknown or very large population:
n₀ = (Z² × p × q) / e²
Where:
- Z = the Z-score for your desired confidence level (1.96 for 95 percent confidence)
- p = estimated proportion of the population with the attribute of interest (use 0.5 if unknown, since this maximizes the required sample size and is the most conservative choice)
- q = 1 − p
- e = margin of error (decimal form)
Worked example with 95 percent confidence, p = 0.5, and a 5 percent margin of error:
n₀ = (1.96² × 0.5 × 0.5) / 0.05² n₀ = (3.8416 × 0.25) / 0.0025 n₀ = 0.9604 / 0.0025 n₀ = 384.16, rounded up to 385
Step 2, if you have a known finite population, apply a correction:
n = n₀ / (1 + (n₀ − 1) / N)
If your population were 5,000, for instance:
n = 385 / (1 + (384) / 5000) n = 385 / (1 + 0.0768) n = 385 / 1.0768 n = 357.6, rounded up to 358
Formula 3: Comparing Two Group Means (Experimental or Quasi-Experimental Designs)
If your study compares two groups (e.g., control vs. experimental), a simple population-based formula isn’t appropriate. You need a power analysis based on effect size, significance level, and desired statistical power.
The general logic, most easily run through software like G*Power rather than by hand, requires you to specify:
- Effect size (Cohen’s d): the expected magnitude of difference between groups. Small (0.2), medium (0.5), or large (0.8) are common conventions, though your effect size should ideally come from prior literature or a pilot study, not guesswork.
- Alpha level: typically 0.05
- Power (1 − beta): typically 0.80, meaning an 80 percent chance of detecting a true effect if one exists
For a medium effect size (d = 0.5), alpha of 0.05, and power of 0.80 in an independent samples t-test, G*Power typically recommends around 64 participants per group, or 128 total. This number changes substantially with effect size, so running the actual calculation in software rather than relying on rules of thumb is strongly recommended for thesis-level work.
Formula 4: Correlation and Regression Studies
For studies testing a correlation or a regression model, sample size depends on the number of predictors and the expected effect size, not simply on population size.
A commonly cited rule of thumb for multiple regression is:
N ≥ 50 + 8m (for testing the overall model)
N ≥ 104 + m (for testing individual predictors)
Where m is the number of predictors.
These are rough guidelines from Tabachnick and Fidell, not hard requirements, and a formal power analysis in G*Power (using the “Linear multiple regression” test family) will give you a more defensible number for your specific design.
Formula 5: Structural Equation Modeling (SEM)
SEM and AMOS-based studies typically require larger samples than simple regression because of the number of parameters being estimated. Common guidelines include:
- A minimum of 200 respondents for models of moderate complexity
- A ratio of at least 10 to 20 respondents per estimated parameter
- Some methodologists recommend 5 to 10 respondents per observed indicator/item as an absolute minimum, with 15 to 20 per indicator preferred for more stable estimates
If your SEM model has many latent constructs and indicators, this can push your required sample well above what a simple Slovin’s or Cochran’s calculation would suggest, which is a common surprise for students who calculated their sample size before finalizing their measurement model.
Common Sample Size Mistakes We See in Proposals
Using Slovin’s formula for an experimental or SEM study. Slovin’s formula was designed for simple descriptive surveys with a known finite population. It is not appropriate for comparing group means or testing a full structural model.
Choosing p = 0.5 without knowing this is the conservative default. This isn’t a mistake in itself, it’s usually the right choice, but panels will ask why you chose it, so know the reasoning (it maximizes the required sample size when you have no prior estimate).
Calculating sample size after data collection, to justify a sample you already have. This is a common defense red flag. Your sample size calculation should come from your research design, not be reverse-engineered to match however many responses you managed to collect.
Ignoring expected attrition or non-response rate. If you’re distributing surveys and expect only a 70 percent response rate, you need to inflate your target sample accordingly (target sample divided by expected response rate) rather than using your calculated minimum as your distribution target.
Applying rules of thumb (like 10 respondents per item) without checking if they still hold for your specific model. These rules are convenient starting points, not universal laws. A structural model with many cross-loadings or a regression model with a small expected effect size may need considerably more.
The Bottom Line
No single formula works for every study. Descriptive surveys with a known population usually call for Slovin’s or Cochran’s formula. Group comparisons call for power analysis based on effect size. Regression and SEM studies scale with the number of predictors or parameters, not population size. The formula you choose should match your research design, and you should be able to explain that choice to your panel in plain language, not just cite a number from a formula you plugged numbers into.
If you’re not sure which sample size approach fits your specific design, or you’d like help running a formal power analysis in G*Power before you finalize your proposal, that’s exactly the kind of question StatAce exists to help with.
Not sure how many respondents your study actually needs? Reach out to StatAce, we help Philippine graduate students and researchers calculate defensible sample sizes before data collection begins.
Leave a Reply