SIMULATION-BASED COMPARISON OF PARAMETRIC AND NONPARAMETRIC STATISTICAL METHODS UNDER DISTRIBUTIONAL VIOLATIONS
DOI:
https://doi.org/10.69980/1jznyc83Keywords:
Parametric methods, Nonparametric methods, Monte Carlo simulation, Distributional violations, Type I error, Statistical powerAbstract
Parametric statistical procedures are commonly used due to their interpretability and efficiency, but can be suboptimal when assumptions of normal data, homogeneity of variance, and resistance to extreme observations are not met. In this study, parametric and nonparametric statistical methods were contrasted with respect to specific violations of their underlying distributions in a Monte Carlo simulation paradigm with empirical validation. Student's independent-samples t-test, Welch's t-test, and the Mann–Whitney U test were tested under the following conditions: normal, skewed, heavy-tailed, heteroscedastic and contaminated distributions. The sample size, effect size and outlier contamination were systematically manipulated, and method performance was evaluated largely by Type I error rates and statistical power. The results indicated that both Student's and Welch's tests were very effective under conditions of approximate normality and were stronger as sample size increased. The Mann-Whitney U test, however, performed better when the distributions were highly skewed, had heavy tails and were contaminated, especially at small and moderate sample sizes. Empirical comparisons also suggested that method choice may affect the results of statistical significance when the data is not normally distributed and when variance is not equal. In general, the results presented here support a condition-dependent philosophy for choosing statistical tests: one should consider the shape of the distribution, sample size, the nature of the variance, and unusual behavior of the data points.
References
1. Aczel, B., Szaszi, B., Sarafoglou, A., Kekecs, Z., Kucharský, Š., Benjamin, D., ... & Wagenmakers, E. J. (2020). A consensus-based transparency checklist. Nature human behaviour, 4(1), 4-6.
2. Ali, M. M., Imon, R., Ali, I., & Yousof, H. M. (Eds.). (2025). Statistical Outliers and Related Topics. CRC Press.
3. Assis, V. R. (2022). Ferreira et al., 2021 - Data table (Version 1) [Data set]. Mendeley Data. https://doi.org/10.17632/r84m23jrzj.1
4. Coolen, F. P., & Himd, S. B. (2020). Nonparametric predictive inference bootstrap with application to reproducibility of the two-sample Kolmogorov–Smirnov test. Journal of Statistical Theory and Practice, 14(2), 26.
5. Coolen, F. P., & Marques, F. J. (2020). Nonparametric predictive inference for test reproducibility by sampling future data orderings. Journal of statistical theory and practice, 14(4), 62.
6. Curtis, D. (2024). Welch’st test is more sensitive to real world violations of distributional assumptions than student’st test but logistic regression is more robust than either. Statistical Papers, 65(6), 3981-3989.
7. Field, A. (2018). Discovering statistics using IBM SPSS statistics (Vol. 5). London: sage.
8. Gelman, A., Hill, J., & Vehtari, A. (2020). Regression and other stories. Analytical methods for social research.
9. James, G., Witten, D., Hastie, T., & Tibshirani, R. (2021). An introduction to statistical learning: with applications in R (Vol. 2). New York: springer.
10. Karch, J. D. (2021). Psychologists should use Brunner-Munzel’s instead of Mann-Whitney’s U test as the default nonparametric procedure. Advances in Methods and Practices in Psychological Science, 4(2), 2515245921999602.
11. Knief, U., & Forstmeier, W. (2021). Violating the normality assumption may be the lesser of two evils. Behavior research methods, 53(6), 2576-2590.
12. Korkmaz, S., & Demir, Y. (2023). Investigation of some univariate normality tests in terms of type-I errors and test power. Journal of Scientific Reports-A, (052), 376-395.
13. Lakens, D. (2022). Sample size justification. Collabra: psychology, 8(1), 33267.
14. Lakens, D., & DeBruine, L. M. (2021). Improving transparency, falsifiability, and rigor by making hypothesis tests machine-readable. Advances in Methods and Practices in Psychological Science, 4(2), 2515245920970949.
15. McElreath, R. (2016). Statistical rethinking: A Bayesian course with examples in R and Stan (Vol. 122). Boca Raton, FL: CRC press.
16. Noguchi, K., Konietschke, F., Marmolejo-Ramos, F., & Pauly, M. (2021). Permutation tests are robust and powerful at 0.5% and 5% significance levels. Behavior Research Methods, 53(6), 2712-2724.
17. Nosek, B. A., Hardwicke, T. E., Moshontz, H., Allard, A., Corker, K. S., Dreber, A., ... & Vazire, S. (2022). Replicability, robustness, and reproducibility in psychological science. Annual review of psychology, 73, 719-748.
18. Riesthuis, P. (2024). Simulation-based power analyses for the smallest effect size of interest: A confidence-interval approach for minimum-effect and equivalence testing. Advances in Methods and Practices in Psychological Science, 7(2), 25152459241240722.
19. Rousselet, G. A., & Wilcox, R. R. (2018). Reaction times and other skewed distributions: problems with the mean and the median. BioRxiv, 383935.
20. Rousselet, G., Pernet, C. R., & Wilcox, R. R. (2023). An introduction to the bootstrap: a versatile method to make inferences by using data-driven simulations. Meta-Psychology, 7.
21. Scheel, A. M., Tiokhin, L., Isager, P. M., & Lakens, D. (2021). Why hypothesis testers should spend less time testing hypotheses. Perspectives on Psychological Science, 16(4), 744-755.
22. Shatz, I. (2024). Assumption-checking rather than (just) testing: The importance of visualization and effect size in statistical diagnostics. Behavior Research Methods, 56(2), 826-845.
23. Simkus, A., Coolen-Maturi, T., Coolen, F. P., & Bendtsen, C. (2025). Statistical perspectives on reproducibility: definitions and challenges. Journal of Statistical Theory and Practice, 19(3), 40.
24. Wilcox, R. (2022a). One-way and two-way ANOVA: Inferences about a robust, heteroscedastic measure of effect size. Methodology, 18(1), 58-73.
25. Wilcox, R. R. (2022b). Two‐way ANOVA: Inferences about interactions based on robust measures of effect size. British Journal of Mathematical and Statistical Psychology, 75(1), 46-58.
26. Wilcox, R. R., & Rousselet, G. A. (2023). An updated guide to robust statistical methods in neuroscience. Current Protocols, 3(3), e719.


