ROBUST VARIABLE SELECTION IN HIGH-DIMENSIONAL REGRESSION USING ADAPTIVE PENALIZATION UNDER CONTAMINATED DATA

Authors

  • Dr. Edward J. Harrington Department of Statistical Learning and Computational Mathematics, Westford University, Manchester, United Kingdom
  • Prof. Katarina M. Novak Institute of Applied Statistics and Data Science, Central European Research University, Prague, Czech Republic
  • Dr. Samuel T. Laurent Department of Quantitative Methods and Statistical Computing, Laurentian Institute of Technology, Lyon, France

DOI:

https://doi.org/10.69980/aeb1he22

Keywords:

high-dimensional regression, robust variable selection, adaptive penalization, Elastic Net, Huber loss

Abstract

High-dimensional regression methods often have many correlated predictors, which makes them prone to overfitting, to having unstable estimates of the regression coefficients, and to making unreliable selections of the predictors, especially when data includes outliers or influential observations. In this study, adaptive penalization methods were studied for robust variable selection in progressively contaminated regression. A comparison of ridge regression, LASSO, Elastic Net, Adaptive LASSO, and Robust Adaptive LASSO was made using the prediction, sparsity and robustness criteria. The models were tested with controlled contamination of 0%, 5%, 10%, 15% and 20% to prove the stability of the models under increasingly harsh conditions. The predictive accuracy was measured by root mean squared error (RMSE), mean absolute error (MAE) and R2 values, and the sparsity of the model was evaluated based on the number of predictors selected. Under these conditions, the data was uncontaminated, and Ridge regression gave the smallest RMSE while Adaptive LASSO gave the sparsest model. But the more contaminated the data, the more stable Robust Adaptive LASSO was compared to the other estimators. Its RMSE showed only a 10.27% rise, and R2 showed an 8.56% decrease, when the contamination level was increased from 0% to 20%. The strong adaptive model only selected 14 predictors at the maximum level of contamination, whereas the LASSO and Elastic Net models selected 66 and 74 variables, respectively.

References

[1]. Amato, U., Antoniadis, A., De Feis, I., & Gijbels, I. (2021). Penalised robust estimators for sparse and high-dimensional linear models: U. Amato et al. Statistical Methods & Applications, 30(1), 1-48.

[2]. Bai, Y., Tian, M., Tang, M. L., & Lee, W. Y. (2021). Variable selection for ultra-high dimensional quantile regression with missing data and measurement error. Statistical Methods in Medical Research, 30(1), 129-150.

[3]. Biziaev, T., Kopciuk, K., & Chekouo, T. (2025). Using prior-data conflict to tune Bayesian regularized regression models. Statistics and Computing, 35(2), 53.

[4]. Cao, X., Gregory, K., & Wang, D. (2023). Inference for sparse linear regression based on the leave-one-covariate-out solution path. Communications in Statistics-Theory and Methods, 52(18), 6640-6657.

[5]. Chatla, S. B., & Mandal, A. (2026). Robust variable selection in high-dimensional nonparametric additive model: SB Chatla, A. Mandal. Annals of the Institute of Statistical Mathematics, 78(1), 115-140.

[6]. Fan, J., Guo, Y., & Jiang, B. (2022). Adaptive Huber regression on Markov-dependent data. Stochastic processes and their applications, 150, 802-818.

[7]. Filzmoser, P., & Nordhausen, K. (2021). Robust linear regression for high‐dimensional data: An overview. Wiley Interdisciplinary Reviews: Computational Statistics, 13(4), e1524.

[8]. Kepplinger, D. (2023). Robust variable selection and estimation via adaptive elastic net S-estimators for linear regression. Computational Statistics & Data Analysis, 183, 107730.

[9]. Kim, S., Turkoz, M., Jeong, M. K., & Elsayed, E. A. (2024). Monitoring of group-structured high-dimensional processes via sparse group LASSO. Annals of Operations Research, 340(2), 891-911.

[10]. Kurata, S., & Hirose, K. (2025). Robust and consistent model evaluation criteria in high-dimensional regression. Journal of Statistical Planning and Inference, 106358.

[11]. Machkour, J., Muma, M., Alt, B., & Zoubir, A. M. (2020). A robust adaptive Lasso estimator for the independent contamination model. Signal processing, 174, 107608.

[12]. Mai, T. T. (2026). Heavy Lasso: sparse penalized regression under heavy-tailed noise via data-augmented soft-thresholding. Statistics and Computing, 36(1), 29.

[13]. Mandal, A., & Ghosh, S. (2025). Robust variable selection criteria for the penalized regression. Journal of Multivariate Analysis, 105540.

[14]. Maurya, D., Barik, A., & Honorio, J. (2026, August). Provable guarantees for robust feature selection in sparse linear models in high-dimensions. In Forty-Second Annual Conference on Uncertainty in Artificial Intelligence.

[15]. Nan, R., Wang, J., Li, H., & Luo, Y. (2025). Robust Variable Selection via Bayesian LASSO-Composite Quantile Regression with Empirical Likelihood: A Hybrid Sampling Approach. Mathematics, 13(14), 2287.

[16]. Nasir, M. J. M., Khan, R. N., Nair, G., & Nur, D. (2025). Coordinate gradient descent algorithm in adaptive LASSO for pure ARCH and pure GARCH models. Computational Statistics, 40(7), 3527-3561.

[17]. Stokell, B. G., & Shah, R. D. (2022). High-dimensional regression with potential prior information on variable importance. Statistics and Computing, 32(3), 52.

[18]. Su, P., Tarr, G., & Muller, S. (2024). Robust variable selection under cellwise contamination. Journal of Statistical Computation and Simulation, 94(6), 1371-1387.

[19]. Sun, Q., Zhou, W. X., & Fan, J. (2020). Adaptive huber regression. Journal of the American Statistical Association, 115(529), 254-265.

[20]. Takahashi, A., & Nomura, S. (2024). Efficient path algorithms for clustered Lasso and OSCAR. Japanese Journal of Statistics and Data Science, 7(2), 967-998.

[21]. Tunguz, B. (2021). Superconductivty data data set [Data set]. Kaggle. https://www.kaggle.com/datasets/tunguz/superconductivty-data-data-set

[22]. Wang, F., Mukherjee, S., Richardson, S., & Hill, S. M. (2020). High-dimensional regression in practice: an empirical study of finite-sample prediction, variable selection and ranking: Wang et al. Statistics and computing, 30(3), 697-719.

[23]. Wang, P., Lu, J., Weng, J., & Mitra, S. (2025). Conditional sufficient variable selection with prior information: P. Wang et al. Computational Statistics, 40(5), 2519-2551.

[24]. Wang, X., Kong, L., Zhuang, X., & Wang, L. (2024). Variance estimation in high-dimensional linear regression via adaptive elastic-net. Journal of Industrial and Management Optimization, 20(2), 630-646.

[25]. Wang, Y., & Karunamuni, R. J. (2022). High-dimensional robust regression with Lq-loss functions. Computational Statistics & Data Analysis, 176, 107567.

[26]. Wu, Y., & Wang, L. (2020). A survey of tuning parameter selection for high-dimensional regression. Annual review of statistics and its application, 7(1), 209-226.

[27]. Xi, L. J., Guo, Z. Y., Yang, X. K., & Ping, Z. G. (2023). Application of LASSO and its extended method in variable selection of regression analysis. Zhonghua yu fang yi xue za zhi [Chinese journal of preventive medicine], 57(1), 107-111.

[28]. Ye, Y., Shao, Y., & Li, C. (2022). A sparse approach for high-dimensional data with heavy-tailed noise. Economic research-Ekonomska istraživanja, 35(1), 2764-2780.

[29]. Zhu, X., Qin, Y., & Wang, P. (2025). Sparsified simultaneous confidence intervals for high-dimensional linear models: X. Zhu et al. Metrika, 88(5), 709-733.

[30]. Żogała-Siudem, B., & Jaroszewicz, S. (2024). Variable screening for Lasso based on multidimensional indexing: B. Żogała-Siudem, S. Jaroszewicz. Data Mining and Knowledge Discovery, 38(1), 49-78.

Downloads

Published

2024-09-30