September 28, 2026
Reliability vs. Validity: Two Concepts Students Always Mix Up
Research Methodology, Statistics and Data Analysis, Thesis & Dissertation Writing
If there’s one question that reliably makes a thesis defense uncomfortable, it’s this: “How did you establish the validity and reliability of your instrument?” Students often answer with a single number, usually Cronbach’s alpha, and assume the question is closed. It isn’t. Reliability and validity are two different properties of a measurement, and having one does not guarantee the other. Panels know this, and they notice when a student treats them as interchangeable.
This article breaks down the difference in plain language, shows how each is assessed, and points out the mistakes we see most often in Philippine thesis proposals and defenses.
The Core Difference in One Sentence
Reliability is about consistency: does your instrument give stable, repeatable results? Validity is about accuracy: does your instrument actually measure what it claims to measure?
The Bathroom Scale Analogy
Imagine a bathroom scale that always reads 5 kilograms too heavy. Every time you step on it, it shows the same wrong number. That scale is highly reliable (consistent) but not valid (accurate).
Now imagine a scale that gives you a different reading every time you step on it, sometimes too high, sometimes too low, averaging out roughly right. That scale isn’t reliable, and because the readings bounce around so much, you can’t trust any single one of them to be valid either.
The key takeaway: a measure can be reliable without being valid, but it cannot be valid without being reliable. Reliability is a necessary condition for validity, not a sufficient one.
Reliability: Is Your Instrument Consistent?
Reliability comes in several forms, and which one you report depends on your study design.
Internal consistency reliability. Do the items in your scale hang together as measures of the same construct? This is what Cronbach’s alpha tests, and it’s the most commonly reported form of Reliability in survey-based theses. A commonly used guideline treats 0.70 or higher as acceptable, though the right threshold depends on the purpose of the instrument and how many items it has.
Test-retest Reliability. If you give the same instrument to the same people twice, separated by a reasonable time gap, do you get similar scores? This is measured with a correlation between the two administrations and is most useful for stable traits rather than fluctuating moods or situations.
Inter-rater Reliability. If two or more observers or coders are scoring the same thing (interview responses, essays, behavioral observations), do they agree? Cohen’s kappa or the intraclass correlation coefficient (ICC) are the usual statistics here, and this form matters a great deal for qualitative coding and observational studies.
Parallel-forms reliability. If you create two different versions of the same test, do they produce equivalent results? This is less common in thesis-level work but appears in educational testing.
Validity: Is Your Instrument Measuring the Right Thing?
Validity is a broader and more layered concept than Reliability, and a single statistic rarely establishes it. The main types you should know:
Content validity. Do the items adequately cover the full domain of the construct you’re trying to measure? This is typically established through expert review, where a panel of subject-matter experts rates each item for relevance, and it’s often quantified using a content validity index (CVI). In most Philippine thesis contexts, this is the validity evidence panels expect to see first, before any pilot testing.
Construct validity. Does the instrument actually measure the theoretical construct it’s supposed to? This is usually examined through factor analysis (exploratory or confirmatory), which checks whether your items cluster into the dimensions your theory predicts. Within construct validity, two sub-types often come up:
- Convergent validity: items that should measure the same construct do correlate strongly with each other (often assessed with average variance extracted, or AVE, in SEM-based studies).
- Discriminant validity: items or constructs that should be distinct from each other are, in fact, empirically distinct (commonly assessed using the Fornell-Larcker criterion or the HTMT ratio).
Criterion validity. Does your instrument correlate with an established external measure or predict a relevant outcome? If you’ve developed a new scale for job satisfaction, for instance, does it correlate as expected with a well-established satisfaction scale (concurrent validity), or does it predict later turnover (predictive validity)?
Face validity. On its surface, does the instrument look like it measures what it claims to? This is the weakest form of validity evidence, because it’s a superficial judgment rather than an empirical test, and panels generally won’t accept it as a substitute for the other types.
Side-by-Side Comparison
| Reliability | Validity | |
|---|---|---|
| Core question | Is it consistent? | Is it accurate? |
| What it tells you | Whether results are stable and repeatable | Whether the instrument measures the intended construct |
| Common evidence | Cronbach’s alpha, test-retest correlation, inter-rater agreement | Expert review (CVI), factor analysis, AVE, correlation with criterion measures |
| Can exist without the other? | Yes, a measure can be reliable but invalid | No, a valid measure must also be reliable |
| Typically established | After pilot testing, with data | Through expert review before data collection, then confirmed with data |
Common Mistakes We See
Reporting only Cronbach’s alpha and calling it “validity and reliability.” Alpha addresses one form of Reliability. It says nothing about whether your items measure the right construct. If your panel asks about validity and you only have an alpha value, you’ve answered a different question.
Assuming an adapted or borrowed instrument is automatically valid and reliable in your context. A scale validated on American or European respondents may not behave the same way with Filipino respondents, particularly if it was translated or if cultural context affects how items are interpreted. Even a well-established instrument should be re-checked for Reliability with your own sample, and ideally re-examined for validity in your setting.
Skipping content validation by experts. Many students go straight to pilot testing without having qualified experts review the items first. Expert review is a low-cost, high-value step, and panels expect to see documentation of it, including who the experts were and how their ratings were summarized.
Confusing a high alpha with a good scale. A very high alpha (above roughly 0.95) can actually signal redundancy, meaning several items are asking essentially the same question in slightly different words. More items and more redundancy inflate alpha without adding any real measurement value.
Treating validity as a yes/no property. Validity is a matter of accumulated evidence, not a single test you pass or fail. Strong studies present several types of validity evidence that support each other, rather than relying on one.
Forgetting to check dimensionality before computing alpha. Cronbach’s alpha assumes your items measure a single underlying construct. If your scale actually has multiple subscales, running alpha across all items together can give a misleading number. Check the factor structure first, then compute alpha for each subscale separately.
A Quick Checklist for Your Thesis
Before your defense, make sure you can point to:
- Expert review for content validity, with documentation (who reviewed, what changes were made)
- A pilot test with reliability results (Cronbach’s alpha for each subscale)
- Factor analysis or another form of construct validity evidence appropriate to your design
- If using SEM, convergent and discriminant validity evidence (AVE, composite Reliability, HTMT, or Fornell-Larcker)
- A clear explanation, in your own words, of why each type of evidence supports your instrument, not just the numbers themselves
The Bottom Line
Reliability tells you whether your instrument gives consistent results. Validity tells you whether those results mean what you think they mean. A strong thesis needs both, supported by more than one type of evidence, and the ability to explain the difference clearly when a panelist asks. Students who can do that, rather than pointing at a single statistic, tend to come across as genuinely understanding their own measurement, which is exactly what defense panels are trying to assess.
If you’re not sure whether your instrument has enough validity and reliability evidence to hold up at defense, or you’d like help running the factor analysis or reliability checks, that’s exactly the kind of question StatAce exists to help with.
Need help validating your questionnaire or checking your reliability results before your defense? Reach out to StatAce; we help Philippine graduate students and researchers build instruments that hold up under panel scrutiny.
Leave a Reply