Effect size is a number that tells you how big a difference or relationship actually is — not just whether it exists — so it answers the question a p-value can't: "Does this result actually matter?" If you've ever reported a "significant" finding and had a supervisor ask "but how big is the effect?", this is the number they wanted. Effect size matters because a result can be statistically significant and practically trivial at the same time, especially in large samples.
Key Takeaways
- Effect size measures the magnitude of a result — how large a difference between groups is, or how strong a relationship between two variables is.
- A p-value tells you whether an effect is likely real; effect size tells you how big it is. You need both to interpret a finding properly.
- Common effect sizes include Cohen's d (mean differences), eta squared (η²) (ANOVA), and r (correlations and associations).
- APA 7 requires reporting effect sizes alongside significance tests — a p-value on its own is now considered incomplete.
- With a large enough sample, almost anything becomes statistically significant, which is exactly why effect size is essential for judging real-world importance.
What is effect size in simple terms?
Effect size is the size of the thing you found. If you compare two teaching methods and one group scores higher, the effect size tells you how much higher — in a standardised way you can compare across studies. Think of statistical significance as answering "is there a signal?" and effect size as answering "how loud is it?"
Here's the intuition. Imagine a new sleep app that increases average sleep by 2 minutes a night. With 50,000 users, that 2-minute difference could be highly statistically significant (p < .001) — but 2 minutes is practically meaningless. The p-value is impressed; the effect size is not. That gap between "statistically real" and "actually meaningful" is the whole reason effect size exists.
Why does effect size matter more than the p-value alone?
A p-value is heavily influenced by sample size, so on its own it can make a trivial finding look impressive. Collect enough data and even a microscopic difference will cross p < .05. Effect size is (mostly) independent of sample size, so it stays honest about how large the effect really is.
This is why APA 7 and most journals now require an effect size for every inferential test. Reporting t(48) = 2.10, p = .041 tells a reader the difference is unlikely to be chance — but it says nothing about whether the difference is a rounding error or a life-changing improvement. Adding d = 0.61 tells them it's a moderate-to-large, meaningful difference. If you're still unsure how p-values and significance thresholds work, our guide on what a p-value actually means pairs naturally with this one.
The three effect sizes you'll actually use
Different tests use different effect size measures. These three cover the vast majority of student and research work.
| Effect size | Used with | Small | Medium | Large |
|---|---|---|---|---|
| Cohen's d | t-tests (mean differences) | 0.20 | 0.50 | 0.80 |
| Eta squared (η²) | ANOVA | .01 | .06 | .14 |
| r | Correlation / association | .10 | .30 | .50 |
These cutoffs (from Cohen, 1988) are rules of thumb, not laws. A "small" effect can be hugely important in medicine (a small reduction in mortality) and a "large" effect can be uninteresting in a lab curiosity. Always interpret the number against your field.
How do you calculate and interpret Cohen's d? (Worked example)
Cohen's d expresses the difference between two group means in units of standard deviation. The formula is the difference between the means divided by the pooled standard deviation.
Say you test a mindfulness intervention on exam anxiety in 60 students (30 in each group):
- Intervention group: M = 42.0, SD = 8.0
- Control group: M = 47.5, SD = 8.4
Step 1 — Difference between means: 47.5 − 42.0 = 5.5 points.
Step 2 — Pooled standard deviation (roughly the average of the two SDs here): ≈ 8.2.
Step 3 — Divide: 5.5 ÷ 8.2 = 0.67.
So d = 0.67. Using Cohen's benchmarks, that's a medium-to-large effect — the intervention group scored about two-thirds of a standard deviation lower on anxiety. That single number is far more informative than "the difference was significant." An independent-samples t-test on this data might return t(58) = 2.59, p = .012, d = 0.67 — significance and magnitude together.
What does eta squared (η²) tell you in ANOVA?
Eta squared tells you what proportion of the variance in your outcome is explained by your grouping variable. It ranges from 0 to 1, and you read it as a percentage.
For example, in a one-way ANOVA comparing three study techniques across 45 students, you might find F(2, 42) = 4.31, p = .019, η² = .17. Here η² = .17 means 17% of the variation in test scores is explained by which study technique students used — a large effect by Cohen's standard. The F and p say the groups differ; the η² says the difference is substantial, not trivial. (Many researchers report partial eta squared, η²p, in factorial designs — a related measure for isolating one factor's contribution.)
How do you report effect size in APA 7 format?
APA 7 asks you to report the effect size immediately after the test statistic and p-value, using the same italics conventions. A few correct examples:
- t-test: t(58) = 2.59, p = .012, d = 0.67
- ANOVA: F(2, 42) = 4.31, p = .019, η² = .17
- Correlation: r(43) = .38, p = .011
Note the APA rules: italicise the test statistic and p, drop the leading zero on values that can't exceed 1 (p = .012, not 0.012), and keep the leading zero on Cohen's d because d can be greater than 1. Where your journal expects it, add a confidence interval, e.g. 95% CI [0.14, 1.19] for the d above.
Getting effect size right without the manual work
The mechanics — pooling standard deviations, matching the right effect size to the right test, formatting it to APA spec — is exactly the part people get wrong or skip. This is where StatRyx helps: you upload your data, and StatRyx picks the correct test, calculates the matching effect size, and writes the full APA 7 result line for you — Cohen's d, η², or r included, with confidence intervals. It's built for researchers who need the statistics done correctly without becoming a statistician first. If you're deciding between a t-test and its nonparametric cousin, see our guide on Mann-Whitney vs the t-test, since the reported effect size differs between them too.
Stop calculating this by hand — run it free in StatRyx → Try StatRyx
Frequently Asked Questions
What is the difference between statistical significance and effect size?
Statistical significance (the p-value) tells you whether an effect is likely to be real rather than chance. Effect size tells you how large that effect is. A result can be significant but tiny, or large but non-significant in a small sample — which is why you should always report and interpret both together.
Is a bigger effect size always better?
Not necessarily — it depends on what you're studying. A "large" effect size is only impressive if the effect is meaningful in context, and a "small" effect can be extremely important in fields like medicine where even a modest improvement