Sample Size Calculation for a T-Test: How Many Participants Do You Need?

For a two-group t-test with a medium effect size (d = 0.5), 80% power, and an alpha of .05, you need roughly 64 participants per group — about 128 total. If your effect is small (d = 0.2) that number balloons to around 394 per group, and if it's large (d = 0.8) it drops to about 26 per group. Guessing the number wrong is one of the most common reasons a thesis gets flagged in review: too few participants and your study is underpowered (you miss real effects), too many and you've wasted time and money.

Key Takeaways

  • Sample size for a t-test depends on three inputs: your expected effect size (Cohen's d), your desired statistical power (usually .80), and your significance level (usually α = .05).
  • A good rule of thumb: independent-samples t-test needs ≈64 per group for a medium effect (d = 0.5), 80% power, two-tailed α = .05.
  • Smaller expected effects require dramatically larger samples — halving the effect size roughly quadruples the sample you need.
  • A paired-samples (within-subjects) t-test needs fewer participants than an independent-samples design for the same effect, because it removes between-person variability.
  • Always calculate sample size before collecting data (a priori power analysis), not after — post-hoc power analysis is widely criticised and rarely accepted by reviewers.

What is sample size calculation for a t-test?

Sample size calculation for a t-test is the process of working out how many participants you need to reliably detect a difference between two means — before you collect any data. It answers the question every supervisor asks: "How did you decide on your n?"

The calculation balances four quantities, and if you fix three, the fourth is determined:

  • Effect size (Cohen's d) — how big the difference you expect is, in standard-deviation units.
  • Alpha (α) — your false-positive risk, conventionally .05.
  • Power (1 − β) — your chance of detecting a real effect, conventionally .80.
  • Sample size (n) — what you're solving for.

Because these four are linked, a power analysis is really just algebra once you've committed to the other three values.

Which numbers do I plug in?

For most behavioural and social-science studies, two of the three inputs are fixed by convention:

  • Alpha = .05 (two-tailed unless you have a strong directional hypothesis).
  • Power = .80 — meaning an 80% chance of detecting the effect if it truly exists. Some fields now prefer .90.

The number you actually have to think about is effect size. Cohen's benchmarks are the usual starting point:

Effect size Cohen's d Interpretation
Small 0.2 Subtle difference, hard to see without large samples
Medium 0.5 Noticeable difference, "visible to the naked eye"
Large 0.8 Obvious, substantial difference

The best source for your d is a pilot study or a previously published effect in your area. Only fall back on Cohen's conventions when you genuinely have no prior information — and when you do, justify your choice explicitly in your methods section.

How do I calculate the sample size? (Worked example)

Imagine you're testing whether a new 6-week mindfulness programme lowers anxiety scores compared with a waitlist control — an independent-samples t-test.

Step 1 — Set your parameters. You review the literature and find similar interventions produce a medium effect, so you choose d = 0.5. You use α = .05 (two-tailed) and power = .80.

Step 2 — Run the calculation. The formula for each group in a two-sample t-test is:

n per group ≈ 2 × ((z₁₋α/₂ + z₁₋β) / d)²

Plugging in z₁₋α/₂ = 1.96 and z₁₋β = 0.84:

n ≈ 2 × ((1.96 + 0.84) / 0.5)² = 2 × (5.6)² = 2 × 31.36 ≈ 63 per group.

Software that uses the exact t distribution (rather than the normal approximation) rounds this up to 64 per group, giving a total of 128 participants.

Step 3 — Add a buffer for dropout. If you expect 15% attrition, recruit 64 ÷ 0.85 ≈ 76 per group so you still have 64 at analysis.

How you'd report it (APA 7):

An a priori power analysis indicated that a sample of 128 participants (64 per group) was required to detect a medium effect (d = 0.50) with 80% power at α = .05 (two-tailed).

How does the design change the number?

A paired-samples t-test (measuring the same people twice, e.g. before and after) is far more efficient. Because each participant acts as their own control, you remove between-person variability.

For the same d = 0.5, 80% power and α = .05, a paired design needs roughly 34 participants total — not per group, total — compared with 128 for the independent design. That efficiency is a major reason within-subjects designs are popular when they're feasible. If you're deciding between designs, see our related guide on independent vs paired t-tests.

Scenario (d = 0.5, power = .80, α = .05) Sample needed
Independent-samples t-test ≈64 per group (128 total)
Paired-samples t-test ≈34 total
One-sample t-test ≈34 total
Independent, small effect (d = 0.2) ≈394 per group
Independent, large effect (d = 0.8) ≈26 per group

Why not just calculate power after collecting data?

Post-hoc (observed) power analysis — plugging your observed effect size back into a power calculation after the study — is strongly discouraged. Because observed power is mathematically tied to your p-value, it tells you nothing new: a non-significant result will always show low "power." Reviewers in psychology and medicine increasingly reject it. The credible approach is an a priori power analysis, done and documented before you touch your data.

What if I don't know my effect size?

This is the hardest part for most students. Three honest options:

  1. Borrow an effect size from the closest published study and cite it.
  2. Run a pilot to estimate the standard deviations and mean difference, then compute d.
  3. Use the smallest effect size of interest (SESOI) — the smallest difference that would actually matter in practice — and power for that. This is the most defensible approach when the literature is thin.

Whatever you choose, state it plainly in your methods. A sentence like "we powered for the smallest clinically meaningful change of 5 points on the anxiety scale" reads far better than "we used a medium effect."

Running it without the formula

Doing this by hand works, but it's easy to mix up one-tailed and two-tailed z-values or forget the dropout buffer. StatRyx runs an a priori power analysis for you: pick your test (independent, paired, or one-sample), enter your effect size, power, and alpha, and it returns the exact n per group plus a ready-to-paste APA sentence. When your data is collected, StatRyx also picks and runs the correct t-test automatically — and checks the assumptions (normality, equal variances) before it does, switching you to a Welch or Mann–Whitney alternative if needed.

Stop calculating this by hand — run it free in StatRyx → Try StatRyx

Frequently Asked Questions

How many participants do I need for a t-test?

For an independent-samples t-test detecting a medium effect (d = 0.5) with 80% power and α =

Stop calculating this by hand. Upload your dataset and StatRyx's AI runs the correct test and returns copy-paste-ready APA 7 output in seconds — no SPSS license, no syntax.

Run your data through StatRyx free →
← All posts