A t-test is a statistical test that tells you whether the difference between two averages is big enough to be real, or whether it's just random chance — and you use it when you're comparing the means of two groups (or one group measured twice) on a numeric outcome. If you're staring at two sets of numbers wondering "are these actually different, or does it just look that way?", the t-test is almost certainly the tool you're reaching for.
Key Takeaways
- A t-test compares two means and tells you whether their difference is statistically significant (unlikely to be due to chance).
- There are three types: one-sample (compare a group to a known value), independent-samples (compare two separate groups), and paired-samples (compare the same group before and after).
- Your data needs to be numeric and roughly normally distributed; if you have three or more groups, use ANOVA instead, and if your data is heavily skewed, use the Mann-Whitney U test.
- A result is "significant" when p < .05, meaning there's less than a 5% chance the difference happened by luck alone.
- Always report an effect size (Cohen's d) alongside the p value — statistical significance tells you if there's a difference, effect size tells you how big it is.
What is a t-test in simple terms?
A t-test answers one question: is the gap between two averages larger than you'd expect from random noise? Imagine two groups of students — one taught with a new method, one with the old — and the new group scores a few points higher on average. A t-test checks whether that few-point gap is a genuine effect or something that could easily have happened by chance.
The "t" refers to the t-statistic, a single number that captures how far apart your two means are relative to how much the scores naturally wobble around. A big difference between means plus low variability inside each group gives you a large t-value, which signals a real difference. A small difference buried in noisy data gives you a small t-value, which signals "probably nothing."
When should I use a t-test?
Use a t-test when all three of these are true: you're comparing exactly two averages, your outcome variable is numeric (like scores, weights, reaction times), and your data is roughly normally distributed.
Here's how to pick the right flavour:
- One-sample t-test — You have one group and want to compare its average to a known or expected value. Example: Is the average IQ in your sample different from the population average of 100?
- Independent-samples t-test — You have two separate groups of different people and want to compare them. Example: Do men and women differ in average sleep hours?
- Paired-samples t-test — You measure the same people twice and compare the two measurements. Example: Do patients' anxiety scores change from before to after therapy?
If you're unsure which test your design calls for, StatRyx reads your variables and automatically selects the correct t-test (or flags when a different test fits better) — no manual decision tree required.
When should I NOT use a t-test?
Skip the t-test when you have more than two groups, when your data is badly skewed, or when your outcome isn't numeric. Running multiple t-tests across three or more groups inflates your false-positive rate — use a one-way ANOVA instead. If your numeric data is heavily non-normal (common with small samples or Likert-style ratings), the non-parametric alternative is the Mann-Whitney U test for independent groups or the Wilcoxon signed-rank test for paired data. If you're weighing those options, see our guide on Mann-Whitney vs the t-test.
A worked example with real numbers
Suppose you run a study on 60 participants to test whether a mindfulness app reduces stress. You randomly assign 30 people to use the app for four weeks and 30 to a waitlist control, then measure stress on a 0–50 scale. Because you have two separate groups, you run an independent-samples t-test.
Your results come out like this:
- App group mean: M = 18.4, SD = 6.1
- Control group mean: M = 23.7, SD = 5.8
- Test result: t(58) = 3.45, p = .001, d = 0.89
Here's what each number means:
- t(58) = 3.45 — The t-statistic is 3.45, and 58 is the degrees of freedom (roughly your total sample size minus the number of groups: 60 − 2 = 58). A t-value this large means the group gap is well beyond what noise would produce.
- p = .001 — There's only a 0.1% chance you'd see a difference this big if the app truly did nothing. Since .001 is below the .05 threshold, the result is statistically significant.
- d = 0.89 — Cohen's d measures effect size. A value of 0.89 is a large effect (0.2 = small, 0.5 = medium, 0.8 = large), so the app didn't just produce a statistically detectable difference — it produced a practically meaningful one.
Plain-English conclusion: The mindfulness app group reported significantly lower stress than the control group, and the size of that difference was large.
How do I report a t-test in APA 7 format?
APA 7 wants the test statistic, degrees of freedom, p value, and an effect size, written in a specific style. A clean write-up of the example above looks like this:
An independent-samples t-test showed that the app group reported significantly lower stress (M = 18.4, SD = 6.1) than the control group (M = 23.7, SD = 5.8), t(58) = 3.45, p = .001, d = 0.89, 95% CI [2.23, 8.37].
Note the APA conventions: the test statistic t, the p value, and M/SD are all italicised, the leading zero is dropped from the p value (.001, not 0.001), and a confidence interval for the mean difference is included. Getting these details right is where many manuscripts get sent back — StatRyx generates this exact APA 7 sentence for you, including the CI and effect size, so you can paste it straight into your results section.
T-test types at a glance
| Test | Use when | Example question | Non-parametric alternative |
|---|---|---|---|
| One-sample t-test | Comparing one group to a known value | Is our sample IQ different from 100? | Wilcoxon signed-rank |
| Independent-samples t-test | Comparing two separate groups | Do two teaching methods differ? | Mann-Whitney U |
| Paired-samples t-test | Comparing the same group twice | Did scores change after training? | Wilcoxon signed-rank |
| One-way ANOVA (not a t-test) | Comparing three or more groups | Do three diets differ in weight loss? | Kruskal-Wallis |
What counts as a "significant" result?
A t-test result is significant when the p value is below .05, the conventional threshold in psychology, medicine, and the social sciences. A p of .03 means there's a 3% probability of seeing your result by chance if there were truly no difference — low enough to reject that "no difference" assumption. But significance alone isn't the whole story: with a large enough sample, even a trivial difference can be "significant," which is why APA requires you to report Cohen's d so readers know whether the effect actually matters.
Stop calculating this by hand — run it free in StatRyx → Try StatRyx
Frequently Asked Questions
What is the difference between a t-test and ANOVA?
A t-test compares the means of exactly two groups, while an ANOVA compares the means of three or more groups. Using multiple t-tests instead of one ANOVA inflates your chance of a false positive, so once you have a third group, switch to ANOVA.