Reliability vs Validity in Research: A Plain-Language Guide With Examples

Reliability is whether a measure gives consistent results every time, while validity is whether it actually measures what it claims to measure — a test can be reliable without being valid, but it cannot be valid without being reliable. If you're writing a methods section and keep mixing these two up, you're not alone: it's one of the most common points of confusion in thesis and dissertation research, and reviewers notice when you get it wrong.

Key Takeaways

  • Reliability = consistency. A reliable measure produces the same result under the same conditions, repeatedly.
  • Validity = accuracy. A valid measure captures the actual construct you intend to study, not something else.
  • Reliability is necessary but not sufficient for validity. A bathroom scale that is 5 kg heavy is perfectly reliable (consistent) but invalid (wrong).
  • Cronbach's alpha above .70 is the most widely cited threshold for acceptable internal-consistency reliability in social science research.
  • You report both. APA 7 methods sections expect reliability coefficients (e.g., α) and evidence of validity for every instrument you use.

What Is the Difference Between Reliability and Validity?

Reliability and validity are the two pillars of measurement quality in research. Reliability asks "Would I get the same answer if I measured again?" Validity asks "Am I measuring the right thing at all?" They are related but separate — and confusing them weakens your study's credibility.

Picture a target. Reliability is how tightly your shots cluster together; validity is how close that cluster sits to the bullseye. You can land five shots in a tight group (high reliability) in the top-left corner, far from the center (low validity). You can also have shots scattered all over — low reliability — which makes hitting the bullseye on average meaningless. Only when shots are both tight and centered do you have a measure that is reliable and valid.

Why Can a Measure Be Reliable but Not Valid?

A measure can be reliable but not valid because consistency and accuracy are independent properties. Consider a kitchen scale that always reads exactly 200 grams too high. Weigh the same apple ten times and you get the same number every time — perfect reliability. But every reading is wrong, so the scale is invalid as a measure of true weight.

In behavioral research this happens constantly. A questionnaire that measures test anxiety might produce rock-solid, repeatable scores — yet if those items secretly tap general perfectionism instead, the scale is reliable but not a valid measure of anxiety. This is why reliability is the floor, not the ceiling: you need it before validity is even possible, but it never guarantees validity on its own.

Types of Reliability

Reliability comes in a few flavors, and your methods section should name the one you actually tested:

  • Internal consistency — Do the items on a scale hang together? Measured with Cronbach's alpha (α) or McDonald's omega (ω). The convention is α ≥ .70 is acceptable, ≥ .80 is good, and ≥ .90 is excellent (though very high α can also signal redundant items).
  • Test-retest reliability — Does the same person score similarly if you measure them twice, a week or two apart? Reported as a correlation (r).
  • Inter-rater reliability — Do two independent coders or raters agree? Measured with Cohen's kappa (κ) for categorical ratings or an intraclass correlation coefficient (ICC) for continuous ones.
  • Parallel-forms reliability — Do two versions of the same test (e.g., Form A and Form B of an exam) produce comparable scores?

Types of Validity

Validity is broader, because there are many ways a measure can go wrong:

  • Content validity — Do the items cover the full scope of the construct? Usually judged by expert reviewers.
  • Construct validity — Does the measure behave the way theory predicts? Split into convergent validity (correlates with related measures) and discriminant validity (does not correlate with unrelated ones).
  • Criterion validity — Does the measure predict an outcome it should? Includes concurrent (now) and predictive (later) validity.
  • Face validity — Does it simply look like it measures the right thing? The weakest form, but useful for participant buy-in.

A Worked Example With Real Numbers

Say you're developing a 10-item Academic Burnout Scale and you collect data from 120 graduate students.

Step 1 — Check internal-consistency reliability. You run Cronbach's alpha and get α = .88. That's above the .80 "good" threshold, so the ten items are measuring something consistently.

Step 2 — Check test-retest reliability. Forty of those students retake the scale two weeks later. The correlation between time 1 and time 2 is r(38) = .81, p < .001. Scores are stable over time.

Step 3 — Check convergent validity. You correlate burnout scores with an established exhaustion measure and find r(118) = .64, p < .001, 95% CI [.52, .74] — a strong positive relationship in the expected direction. Good convergent evidence.

Step 4 — Check discriminant validity. You correlate burnout with shoe size and get r(118) = .04, p = .66 — essentially zero, as it should be. Burnout isn't bleeding into something unrelated.

Taken together: your scale is reliable (α = .88, retest r = .81) and shows validity evidence (converges with exhaustion, stays independent of shoe size). That's a defensible instrument. Running these coefficients by hand is tedious and error-prone — StatRyx computes Cronbach's alpha, test-retest correlations, and validity correlations and hands you the APA-formatted output in one pass.

Reliability vs Validity: Key Differences

Feature Reliability Validity
Core question Is it consistent? Is it accurate / measuring the right thing?
Analogy Tight cluster of shots Shots centered on the bullseye
Typical statistic Cronbach's α, κ, ICC, test-retest r Convergent/discriminant r, factor loadings, criterion r
Common threshold α ≥ .70 No single cutoff — built from a body of evidence
Dependency Can exist without validity Requires reliability to exist
What it protects against Random measurement error Systematic error / measuring the wrong construct

How Do You Report Reliability and Validity in APA 7?

In an APA 7 methods section, report the reliability coefficient for each instrument inline — for example: "The Academic Burnout Scale showed strong internal consistency (α = .88)." For validity evidence, report the relevant correlations with test statistic, degrees of freedom, exact p value, and a confidence interval where appropriate: "Burnout scores correlated with exhaustion, r(118) = .64, p < .001, 95% CI [.52, .74]."

Note the APA conventions that trip people up: italicize test statistics (r, p, α), drop the leading zero on values that cannot exceed 1 (write p < .001, not 0.001), and report exact p values unless they fall below .001.

Which Measure Do You Actually Need to Run?

If you're adapting or creating a scale, you almost always want internal consistency (Cronbach's alpha) at minimum, plus at least one piece of validity evidence (usually convergent). If your data involves raters coding behavior, you need inter-rater reliability (kappa or ICC) instead. When you're choosing between correlation-based tests to establish validity, our guide on corre

Stop calculating this by hand. Upload your dataset and StatRyx's AI runs the correct test and returns copy-paste-ready APA 7 output in seconds — no SPSS license, no syntax.

Run your data through StatRyx free →
← All posts