A t-test is a statistical test that checks whether the average (mean) of one group is meaningfully different from another group's average — or from a fixed number — rather than different just by chance. If you're staring at two columns of numbers wondering "are these groups actually different, or did I get lucky?", the t-test is almost certainly the tool you're reaching for — and it's one of the most common reasons researchers open their data in frustration.
Key Takeaways
- A t-test compares means between two groups (or one group against a known value) to see if the difference is statistically significant.
- There are three types: one-sample, independent-samples (two separate groups), and paired-samples (the same people measured twice).
- Use a t-test when your outcome is numeric, roughly normally distributed, and you're comparing no more than two groups.
- A p-value below .05 is the conventional threshold for calling a difference "statistically significant."
- Always report an effect size (Cohen's d) alongside the p-value, because significance alone doesn't tell you how big the difference is.
What Does a T-Test Actually Do?
A t-test asks a simple question: is the gap between two averages big enough that it's unlikely to be a fluke? Imagine two groups of students — one that used a new study app and one that didn't — and their exam scores average 78 and 72. A t-test weighs that 6-point difference against how spread out the scores are and how many people you tested. A big difference with tight, consistent scores signals a real effect; a small difference with wildly scattered scores probably doesn't.
The output is a t-statistic (how many "standard errors" apart the means are) and a p-value (the probability of seeing a difference this large if there were truly no difference). The bigger the t-statistic and the smaller the p-value, the more confident you can be that the groups genuinely differ.
When Should I Use a T-Test?
Use a t-test when all of the following are true: your outcome variable is numeric (test scores, reaction times, blood pressure), you're comparing two groups or fewer, and your data is roughly normally distributed. If you have three or more groups to compare, you need an ANOVA instead, not a string of t-tests. If your outcome is a category (passed/failed, yes/no), you need a chi-square test.
Here's a quick rule of thumb: a t-test is for comparing two averages; anything beyond two groups, or non-numeric outcomes, needs a different test. If your data is heavily skewed or ordinal, the non-parametric cousin of the t-test — the Mann-Whitney U test — is the safer choice.
The Three Types of T-Test
Choosing the right type of t-test trips up more researchers than the test itself. Here's how they differ.
| Type of t-test | When to use it | Example |
|---|---|---|
| One-sample | Compare one group's mean to a known or hypothesised value | Is the average IQ in your sample different from the population average of 100? |
| Independent-samples | Compare the means of two separate, unrelated groups | Do men and women differ in average sleep hours? |
| Paired-samples | Compare two measurements from the same people | Do patients' anxiety scores change before vs. after therapy? |
The single most important distinction: if the same participants appear in both groups, you need a paired-samples t-test; if the two groups contain different people, you need an independent-samples t-test. Picking the wrong one inflates or deflates your p-value and can lead to the wrong conclusion.
A Worked Example With Real Numbers
Say you're testing whether a mindfulness program reduces stress. You recruit 40 participants and randomly assign 20 to the mindfulness group and 20 to a control group. After eight weeks, you measure stress on a 0–50 scale.
- Mindfulness group: M = 22.4, SD = 5.1
- Control group: M = 27.8, SD = 5.6
Because these are two different sets of people, you run an independent-samples t-test. The result:
t(38) = 3.19, p = .003, d = 1.01
Here's what each piece means, in plain English:
- t(38) = 3.19 — the two means are about 3.19 standard errors apart. The "38" is the degrees of freedom (your total sample size of 40 minus 2).
- p = .003 — there's only a 0.3% chance you'd see a difference this large if mindfulness had no real effect. Since .003 is well below .05, the difference is statistically significant.
- d = 1.01 — Cohen's d, the effect size. A d of 1.01 is a large effect (0.2 = small, 0.5 = medium, 0.8 = large). The mindfulness group scored more than a full standard deviation lower in stress.
So you'd conclude: the mindfulness program significantly reduced stress, and the effect was large. Notice that the p-value tells you the difference is real, while the effect size tells you it's big — you need both.
What Counts as "Significant"?
A result is conventionally called statistically significant when the p-value is below .05, meaning there's less than a 5% chance the difference is due to random noise. But significance is not the whole story: with a huge sample, even a trivial difference can be "significant." That's why APA 7 requires you to report an effect size (like Cohen's d) and, increasingly, a 95% confidence interval for the difference between means. A significant p-value with a tiny effect size is often not worth writing home about.
How Do I Report a T-Test in APA 7 Format?
APA 7 has a precise format, and getting the italics and leading zeros right matters. The template is:
t(df) = [value], p = [value], d = [value]
A full write-up of the example above would read:
An independent-samples t-test showed that participants in the mindfulness group reported significantly lower stress (M = 22.4, SD = 5.1) than those in the control group (M = 27.8, SD = 5.6), t(38) = 3.19, p = .003, d = 1.01, 95% CI [2.00, 8.80].
Note the APA rules: the test statistic t, the p, and M/SD are italicised; p-values drop the leading zero (.003, not 0.003); and you report the confidence interval in square brackets. These details are exactly where manual write-ups go wrong — and where a tool that formats the output for you saves hours.
Running Your T-Test Without the Headache
You don't need to memorise formulas or wrestle with SPSS menus to do this correctly. StatRyx automatically checks your data, picks the correct type of t-test, verifies the assumptions, and returns a ready-to-paste APA 7 write-up — including the effect size and confidence interval most people forget. If your data turns out to be skewed and a t-test isn't appropriate, StatRyx flags it and suggests the right alternative (often the Mann-Whitney U test) instead of silently giving you a misleading result. If you're still deciding which test fits your data, our guide on choosing between the Mann-Whitney U and the t-test walks through the decision.
For context on cost: a standalone IBM SPSS Statistics subscription runs around \$99 per month for a single user — a steep price for running a handful of t-tests during a thesis.
Stop calculating this by hand — run it free in StatRyx → Try StatRyx
Frequently Asked Questions
What is the difference between a t-test and an ANOVA?
A t-test compares the means of two groups, while an ANOVA compares the means of three or more groups at once. Running multiple t-tests instead of a single ANOV