September 24, 2026
10 Common Statistical Mistakes in Thesis Research
Academic Writing, Research and Statistics, Statistics and Data Analysis
Statistical analysis is one of the most challenging parts of writing a thesis. Even researchers who have carefully designed their studies and collected quality data can make mistakes when selecting statistical tests, analyzing results, or interpreting findings.
These mistakes can lead to inaccurate conclusions, questionable research findings, and difficult questions during a thesis defense.
The good news? Most statistical mistakes are preventable. Whether you’re using SPSS, Jamovi, Stata, or another statistical software package, understanding these common errors can help you produce more credible and defensible research.
Here are ten statistical mistakes thesis writers should avoid.
1. Choosing the wrong statistical test
One of the most common mistakes is selecting a statistical test simply because it is familiar or frequently used in previous studies.
For example, a researcher might use Pearson correlation to examine the relationship between two variables without first considering the measurement scales, distribution, or nature of the relationship.
How to avoid it: Choose a statistical test based on your research objectives, hypotheses, measurement scales, study design, and relevant statistical assumptions.
| Research objective | Possible statistical test |
|---|---|
| Compare two independent group means | Independent-samples t-test |
| Compare three or more independent group means | One-way ANOVA |
| Examine a linear relationship between two quantitative variables | Pearson correlation |
| Examine a monotonic relationship using ranks | Spearman correlation |
| Predict a continuous outcome | Linear regression |
| Examine associations between categorical variables | Chi-square test |
These are general examples. The final choice depends on your data and research design.
2. Ignoring statistical assumptions
Statistical tests have assumptions that researchers should evaluate before interpreting their results.
For example, linear regression requires researchers to consider linearity, independence of errors, and the behavior of residuals. Depending on the intended inference, additional assumptions may apply.
Ignoring serious violations can produce misleading estimates, standard errors, confidence intervals, or significance tests.
How to avoid it: Examine your data using appropriate diagnostic procedures, including residual plots, distributional assessments, and tests for unequal variances where relevant. When assumptions are not adequately met, consider suitable robust or alternative methods.
3. Misinterpreting the p-value
Many thesis writers believe that a p-value below 0.05 proves their hypothesis is correct. Others assume that a p-value above 0.05 means there is absolutely no relationship between the variables.
Both interpretations are incorrect.
A p-value describes how incompatible the observed results, or more extreme results, would be with the specified statistical model and null hypothesis. It does not measure the probability that the research hypothesis is true.
How to avoid it: Interpret p-values alongside effect sizes, confidence intervals, study design, and relevant prior evidence.
For example, a correlation of r=0.15 with p=0.02 may be statistically significant, but the observed relationship is relatively weak. Its practical importance depends on the research context.
4. Assuming correlation means causation
Imagine that your study finds a statistically significant relationship between financial literacy and saving behavior among college students. Can you conclude that financial literacy causes students to save more money? Not necessarily.
Other variables, such as family income, financial attitudes, or parental influence, may contribute to the observed relationship.
How to avoid it: Use language appropriate to your research design. For a correlational study, describe variables as being associated with or related to one another rather than claiming that one variable causes the other.
5. Using Cronbach’s alpha incorrectly
Cronbach’s alpha is frequently used in thesis research to assess the internal consistency of questionnaire items. However, a high alpha does not automatically establish that an instrument is valid or that all items measure one construct.
Researchers also sometimes remove questionnaire items solely to increase alpha without considering their theoretical importance.
How to avoid it: Examine item-total correlations, the theoretical meaning of the items, and evidence of dimensionality. Use exploratory or confirmatory factor analysis when appropriate.
Remember that reliability and validity are different concepts. A questionnaire can produce internally consistent scores without measuring the intended construct adequately.
6. Using an inadequate or inappropriate sample size
A common mistake is selecting a sample size based solely on convenience or copying the sample size used in another thesis.
An inadequate sample can reduce statistical power, produce imprecise estimates, and make it difficult to identify meaningful relationships. However, a larger sample does not automatically correct sampling bias.
How to avoid it: Determine your sample size using an appropriate method, such as power analysis, considering your research design, expected effect size, significance level, and desired statistical power.
For example, a study using structural equation modeling (SEM) may have different sample-size requirements from one using a simple correlation analysis.
7. Failing to check for missing data and outliers
Researchers sometimes proceed directly to statistical analysis without inspecting their datasets. Missing responses, incorrect data entries, duplicate records, and extreme observations can substantially influence statistical results.
For example, an incorrectly encoded monthly income of ₱500,000 instead of ₱50,000 could distort the mean and influence regression estimates.
How to avoid it: Screen your data before running your main statistical analyses. Identify missing values, investigate unusual observations, and document any data-cleaning decisions.
Do not automatically delete outliers. Some extreme observations represent genuine and important information.
8. Reporting statistical results without interpreting them
One of the most noticeable weaknesses in Chapter 4 is presenting tables generated by statistical software without explaining what the findings mean. A researcher might report a correlation coefficient, significance value, and regression coefficient without connecting the findings to the research objectives.
How to avoid it: After presenting each statistical table, explain the results in relation to the corresponding research question or hypothesis.
Example of an appropriate interpretation: A Pearson correlation analysis revealed a positive relationship between financial literacy and saving behavior, r=0.45, p=0.003. This suggests that respondents with higher financial literacy scores tended to report better saving behavior. However, the correlational research design does not establish causality. Note: Illustrative results only; not actual research data.
A good interpretation explains the direction, magnitude, statistical significance, and substantive meaning of the findings without making unsupported claims.
9. Conducting multiple statistical tests without considering false positives
Some researchers conduct numerous statistical tests until they obtain a statistically significant result.
However, running many hypothesis tests increases the likelihood of finding at least one statistically significant result by chance, even when all null hypotheses are true.
This practice can lead to misleading conclusions, particularly when researchers selectively report only significant findings.
How to avoid it: Specify your research hypotheses and primary analyses before examining the results. When conducting multiple comparisons, consider appropriate adjustments, such as Bonferroni correction or procedures that control the false discovery rate.
Report both statistically significant and non-significant findings when they are relevant to your research questions.
10. Drawing conclusions that go beyond the statistical evidence
The final mistake occurs when researchers make conclusions or recommendations that their statistical findings cannot adequately support. For example, a study might find a positive association between employee motivation and job performance. The researcher then recommends that increasing motivation will necessarily improve employee performance.
Although this recommendation may sound reasonable, the correlational findings alone cannot establish that the proposed intervention will produce the desired outcome.
How to avoid it: Ensure your conclusions align with your research design, statistical results, sampling method, and study limitations.
Clearly distinguish what your research demonstrated, what your findings suggest, and what requires further investigation.
Final Thoughts: Good Statistical Analysis Strengthens Your Thesis
Statistical software makes data analysis more accessible, but it cannot replace a researcher’s understanding of statistical principles. Whether you are preparing an undergraduate thesis, master’s thesis, doctoral dissertation, or journal article, choosing appropriate statistical methods and interpreting your results correctly are essential to producing credible research.
Remember that a statistically significant result is not necessarily an important finding, and a non-significant result does not automatically mean your research has failed. The goal of statistical analysis is not simply to obtain significant results. It is to generate meaningful evidence that addresses your research questions.
Need Help With Your Thesis Statistics?
Choosing the right statistical test, interpreting your results, or preparing Chapter 4 can be challenging. StatsAce provides research and statistical consulting services to help students, faculty, and researchers approach data analysis with confidence.
We assist with research methodology, questionnaire development, statistical analysis using SPSS and Jamovi, regression, factor analysis, SEM, and interpreting statistical results.
Leave a Reply