A t-test is a statistical test that tells you whether the average of one group is meaningfully different from another average — and you use it when you're comparing the means of one or two groups on a single numeric outcome. If you're staring at two columns of numbers (say, test scores for a treatment group and a control group) and wondering whether the gap between their averages is real or just random noise, the t-test is almost always the tool you're reaching for.
Key Takeaways
- A t-test compares means. It answers one question: is the difference between two averages larger than you'd expect by chance?
- Use a t-test when your outcome is numeric (continuous) and you're comparing one or two groups — not three or more (that's ANOVA).
- There are three versions: one-sample, independent-samples, and paired-samples — the right one depends on how many groups you have and whether they're related.
- Significance is usually judged at p < .05, and you should always report an effect size (Cohen's d) alongside it.
- StatRyx picks the correct t-test variant for your data and produces the APA 7 write-up automatically, so you don't have to guess which version applies.
What is a t-test, in plain terms?
A t-test checks whether two averages are far enough apart to be convincing. Imagine two groups of students: one studied with flashcards, the other with a textbook. The flashcard group scored 78 on average; the textbook group scored 72. Is that 6-point gap a real effect, or could it have happened by luck if you re-ran the study?
The t-test weighs the size of the difference against how spread out the scores are. A big gap with tightly clustered scores is convincing. The same gap with wildly scattered scores is not. The output — the t statistic and its p value — formalizes that gut feeling into a number you can report.
When should I use a t-test?
Use a t-test when all three of these are true:
- Your outcome variable is continuous — things measured on a scale, like test scores, reaction times, blood pressure, or income.
- You're comparing one or two groups — not three or more.
- You care about the mean (average), not proportions or categories.
If you have three or more groups, you need a one-way ANOVA instead. If your outcome is a category (passed/failed, yes/no), you need a chi-square test. And if your data is badly skewed or ordinal (like Likert ratings), a Mann-Whitney U test is often the safer choice — if you're torn between those two, see our guide on Mann-Whitney vs the t-test.
What are the three types of t-tests?
Picking the wrong variant is the single most common t-test mistake we see. Here's how to choose.
| T-test type | Use it when… | Example |
|---|---|---|
| One-sample | You compare one group's mean to a known or hypothesized value | Is the average IQ of your sample different from the population norm of 100? |
| Independent-samples | You compare the means of two separate, unrelated groups | Do men and women differ in average sleep hours? |
| Paired-samples | You compare two measurements from the same people | Do patients' anxiety scores drop from before to after therapy? |
Independent vs paired: the key distinction
The deciding question is simple: are the two sets of numbers coming from the same people (or matched pairs), or from two different groups?
- Same people measured twice → paired-samples t-test (e.g., pre-test and post-test on the same students).
- Two different sets of people → independent-samples t-test (e.g., experimental group vs control group).
Getting this wrong inflates or deflates your p value, so it's worth pausing on. StatRyx detects whether your two columns are related or independent and selects the matching test for you.
A worked example: does a study method improve scores?
Let's run an independent-samples t-test with real numbers.
The study: 40 students were randomly assigned to study with flashcards (n = 20) or a textbook (n = 20). We then compared exam scores.
- Flashcard group: M = 78.0, SD = 8.2
- Textbook group: M = 72.0, SD = 9.1
Running the independent-samples t-test gives:
t(38) = 2.19, p = .035, d = 0.69
Here's what each number means:
- 38 is the degrees of freedom (total sample size minus the number of groups: 40 − 2 = 38).
- 2.19 is the t statistic — the size of the difference relative to the spread. Bigger absolute values mean a more convincing gap.
- p = .035 means there's a 3.5% probability of seeing a difference this large if the two methods were actually equal. Because .035 is below the conventional .05 threshold, we call the difference statistically significant.
- d = 0.69 is Cohen's d, the effect size — a medium-to-large effect. By convention, d = 0.2 is small, 0.5 is medium, and 0.8 is large.
Plain-language conclusion: flashcard students scored significantly higher than textbook students, and the effect was substantial in practical terms — not just statistically detectable.
What counts as "significant" — and why effect size matters
A result is conventionally statistically significant when p < .05. But significance alone can mislead. With a very large sample, even a trivial 1-point difference can be "significant." That's why APA 7 requires you to report an effect size (Cohen's d for t-tests) and ideally a confidence interval for the mean difference.
In the example above, the full APA-style report would be:
An independent-samples t-test showed that flashcard users scored significantly higher (M = 78.0, SD = 8.2) than textbook users (M = 72.0, SD = 9.1), t(38) = 2.19, p = .035, d = 0.69, 95% CI [0.45, 11.55].
Notice the details APA cares about: t and p are italicized, there's no leading zero on the p value (.035, not 0.035), and the effect size and CI are both reported.
What assumptions does a t-test make?
A t-test gives trustworthy results only when your data roughly meets these conditions:
- Normality: the outcome is approximately normally distributed within each group (matters less with larger samples, roughly n > 30 per group).
- Independence: observations don't influence each other (violated if, say, you measure the same person multiple times and use an independent test).
- Homogeneity of variance: the two groups have similar spread. If they don't, use Welch's t-test, which corrects for unequal variances and is a safe default.
If assumptions are badly violated — heavy skew, tiny samples, ordinal data — switch to the non-parametric Mann-Whitney U test (for independent groups) or Wilcoxon signed-rank test (for paired data). StatRyx checks these assumptions automatically and flags when a non-parametric alternative is more appropriate.
How to run a t-test without doing the math by hand
You can compute a t-test in SPSS (a license runs roughly $100+ per month for individuals), in R (free but code-heavy), or in free desktop apps like JASP and jamovi. Each requires you to already know which variant you need and how to interpret the output.
StatRyx takes a different approach: you upload your data, and it identifies whether you need a one-sample, independent, or paired t-test, checks the assumptions, runs the correct calculation (including Welch's correction when needed), and hands back a copy-paste APA 7 results sentence with the effect size filled in.
Stop calculating this by hand — run it free in StatRyx → Try StatRyx