A BAYESIAN STATISTICAL FRAMEWORK FOR ASSESSING ASSOCIATIONS BETWEEN CATEGORICAL VARIABLES IN SMALL SAMPLES

Authors

  • Milee Brahmakshatriya Research Scholar, GLS University

DOI:

https://doi.org/10.69980/tmzf5210

Keywords:

Bayesian inference, categorical variables, small samples, Bayes factor, posterior probability

Abstract

Assessing associations between categorical variables can be challenging in small samples, particularly when contingency tables contain sparse or low-frequency cells. This study applied a Bayesian statistical framework to a publicly available secondary dataset containing 100 observations to evaluate associations between 11 categorical predictors and a binary risk-status outcome. Descriptive frequencies and contingency tables were first examined, while chi-square tests were used as reference analyses. Bayesian inference was then used to quantify the strength and uncertainty of associations through Bayes factors, posterior probabilities, posterior Cramer's V estimates, and 95% credible intervals. Of the 100 observations, 66 were classified as Risk and 34 as No Risk. Location range showed the strongest association with risk status, followed by Price, Quality, Finances, Warranty, Business, and Market. Communication, Performance, and Attitude demonstrated moderate Bayesian evidence, whereas Delivery showed little evidence of association. Posterior effect estimates confirmed that Location range had the largest association magnitude, while Delivery had the weakest. Frequentist and Bayesian findings were generally consistent; however, the Bayesian approach provided additional information on the relative strength of evidence and uncertainty surrounding the observed relationships. Overall, the findings demonstrate that Bayesian methods provide a practical and informative approach for categorical association analysis in small samples, particularly when sparse contingency cells limit reliance on conventional asymptotic procedures.

References

1. Bainter, S. A. (2017). Bayesian estimation for item factor analysis models with sparse categorical indicators. Multivariate behavioral research, 52(5), 593-615.

2. Barthelmes, V. M., Heo, Y., Fabi, V., & Corgnati, S. P. (2017). Exploration of the Bayesian Network framework for modelling window control behaviour. Building and Environment, 126, 318-330.

3. Cain, M. K., & Zhang, Z. (2019). Fit for a Bayesian: An evaluation of PPP and DIC for structural equation modeling. Structural Equation Modeling: A Multidisciplinary Journal, 26(1), 39-50.

4. Carpenter, J. R., & Smuk, M. (2021). Missing data: A statistical framework for practice. Biometrical Journal, 63(5), 915-947.

5. Chen, Y., Li, X., Liu, J., & Ying, Z. (2025). Item response theory—A statistical framework for educational and psychological measurement. Statistical Science, 40(2), 167-194.

6. Enders, C. K., Keller, B. T., & Levy, R. (2018). A fully conditional specification approach to multilevel imputation of categorical and continuous variables. Psychological methods, 23(2), 298.

7. Ferrari, F., & Dunson, D. B. (2021). Bayesian factor analysis for inference on interactions. Journal of the American Statistical Association, 116(535), 1521-1532.

8. Gamoura, S. (2025). Paper_IJPE_Repository_1_Dataset_Purchases_Original_and_Augmented (Version 2) [Data set]. Mendeley Data. https://doi.org/10.17632/24j2xp2xvy.2

9. Häse, F., Aldeghi, M., Hickman, R. J., Roch, L. M., & Aspuru-Guzik, A. (2021). Gryffin: An algorithm for Bayesian optimization of categorical variables informed by expert knowledge. Applied Physics Reviews, 8(3).

10. Ince, R. A., Giordano, B. L., Kayser, C., Rousselet, G. A., Gross, J., & Schyns, P. G. (2017). A statistical framework for neuroimaging data analysis based on mutual information estimated via a gaussian copula. Human brain mapping, 38(3), 1541-1573.

11. Jones, S., Johnstone, D., & Wilson, R. (2017). Predicting corporate bankruptcy: An evaluation of alternative statistical frameworks. Journal of Business Finance & Accounting, 44(1-2), 3-34.

12. Khakifirooz, M., Chien, C. F., & Chen, Y. J. (2018). Bayesian inference for mining semiconductor manufacturing big data for yield enhancement and smart production to empower industry 4.0. Applied Soft Computing, 68, 990-999.

13. Kvålseth, T. O. (2025). Association measures for nominal categorical variables. In International Encyclopedia of Statistical Science (pp. 103-107). Berlin, Heidelberg: Springer Berlin Heidelberg.

14. Li, C. H. (2021). Statistical estimation of structural equation models with a mixture of continuous and categorical observed variables. Behavior research methods, 53(5), 2191-2213.

15. Ma, Z., & Chen, G. (2018). Bayesian methods for dealing with missing data problems. Journal of the Korean Statistical Society, 47(3), 297-313.

16. Makowski, D., Ben-Shachar, M. S., Chen, S. A., & Lüdecke, D. (2019). Indices of effect existence and significance in the Bayesian framework. Frontiers in psychology, 10, 2767.

17. Ovaskainen, O., Tikhonov, G., Norberg, A., Guillaume Blanchet, F., Duan, L., Dunson, D., ... & Abrego, N. (2017). How to make more out of community data? A conceptual framework and its implementation as models and software. Ecology letters, 20(5), 561-576.

18. Qi, X., Wang, S., Fang, C., Jia, J., Lin, L., & Yuan, T. (2025). Machine learning and SHAP value interpretation for predicting comorbidity of cardiovascular disease and cancer with dietary antioxidants. Redox biology, 79, 103470.

19. Quintana, D. S., & Williams, D. R. (2018). Bayesian alternatives for common null-hypothesis significance tests in psychiatry: a non-technical guide using JASP. BMC psychiatry, 18(1), 178.

20. Rubin, A. F., Gelman, H., Lucas, N., Bajjalieh, S. M., Papenfuss, A. T., Speed, T. P., & Fowler, D. M. (2017). A statistical framework for analyzing deep mutational scanning data. Genome biology, 18(1), 150.

21. Wang, Y., Zhang, Y., Lu, Y., & Yu, X. (2020). A Comparative Assessment of Credit Risk Model Based on Machine Learning——a case study of bank loan data. Procedia Computer Science, 174, 141-149.

Downloads

Published

2026-04-26