A general framework is proposed that integrates both optimistic and pessimistic optimization approaches in solving the regression problem to address outlier cleaning and robustification in a unified fashion and develops solution methods that can be applied to handle data sets of different scales.
Abstract
In this study, we propose a general framework that integrates both optimistic and pessimistic optimization approaches in solving the regression problem to address outlier cleaning and robustification in a unified fashion. Although data cleaning aims to down-weight the outliers, robustification renders the regression models to heavily rely on extreme data. The main objective of this framework is to construct a new optimization scheme capable of withstanding the influence of outliers without harming the robustness level, by combining these two rather contrasting concepts and operations. In addition to showing its generalization to a few well-known regression models, a set of structural properties of our framework is derived to ensure its statistical significance and to understand its computational demand. Then, we develop solution methods, including mixed integer formulations, alternating direction method of multipliers algorithms, and computation enhancement techniques, that can be applied to handle data sets of different scales. Numerical results on both synthetic and benchmark data sets from the University of California, Irvine (UCI) Machine Learning Repository verify the superiority of our new framework and demonstrate the unified strength to handle complex data sets.
History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning.
Funding: X. Qian received financial support from the U.S. National Science Foundation (NSF) [Grants SHF-2215573 and IIS-2212419].
Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2024.0884 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2024.0884 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ .
In regression modeling, collinearity among input variables, unevenness in the output observations, and outlier points can affect parameter estimation and reduce the optimality of the models. Several approaches exist to address these problems, including penalized and nonparametric models. However, each has its challenges and performs well only for a specific purpose, leaving the other problems unaddressed. In this paper, with a primary focus on fuzzy regression models, we propose a method based on a new linear uniform model that covers the functions of all these methods, such as reducing and controlling collinearity, coping with unevenness in a dataset, and outlier effects, as well as addressing the problems in their structures, such as the lack of closed form, the nonlinearity of the model parameter formula relation, and the single-purpose nature of the obtained models. Furthermore, when the normal distribution is assumed, it performs better than the best method for model fit, i.e., the least-squares method. In this paper, we demonstrate the optimal performance of the proposed method in addressing the aforementioned problems through various numerical and practical examples and compare it with other existing methods.
M. Kashani, M. Arashi, Mohammad Farshad et al.· International Journal of Unc...· 0 citations
Phase I analysis is essential to understand the variability of the process and determine its stability. Providing Phase II control charts with poor parameters’ estimates leads to a weak performance. In case of incomplete Phase I data, the problem of missing values must be dealt with before estimating the process parameters. Researchers commonly rely on the traditional Mean Substitution (MS) and/or the Stochastic Regression (SRG) imputation methods, whilst the number of studies exploiting machine learning algorithms in the SPC field is rapidly increasing. Accordingly, in this study, we consider two common and powerful machine learning‐based imputation methods; which are the k‐Nearest Neighbors (kNN) and Support Vector Regression (SVR). We compare their effect with the two traditional methods, the MS and the SRG methods, on the performance of the G‐chart designed to monitor the process variability. Our results show that kNN imputation either surpasses the performance of the traditional methods or provides a similar performance. An application of the G‐chart is also illustrated. We recommend the use of the kNN imputation while monitoring the process dispersion.
Dina A. Desoki, Nesma A. Saleh, A. Saad et al.· Quality and Reliability Engi...· 0 citations
The findings corroborate the concern that standard resampling methods often yield biased GE estimates in nonstandard settings, underscoring the importance of tailored GE estimation.
R. Hornung, Malte Nalenz, Lennart Schneider et al.· Statistical Science· 0 citations
As data accumulation continues to expand and information technologies evolve, machine learning methods have become widely adopted, making the effectiveness of learning algorithms crucial. Among the most popular machine learning models is Lasso regression, renowned for its feature selection capabilities and ability to address multicollinearity. This paper introduces novel algorithms for estimating Lasso regression parameters by reformulating the problem as an inverse single-point optimization task. Two algorithms are proposed: Lasso-I, which implements coordinate descent with L1 regularization, and Lasso-H, a hybrid approach that combines Lasso-I with wrapper techniques for feature selection using information criteria. The iterative algorithms involve calculating partial derivatives and selecting arguments for adjustment based on residual sum of squares or information criteria. Algorithm evaluation was performed using linear and logistic regression models across diverse datasets from KEEL and UCI repositories, alongside various metrics including the AIC, MSE, and
R
2
. The experimental results demonstrate that the algorithms effectively address parameter estimation problems, with Lasso-H achieving optimal AIC values in 90% of logistic regression cases. The proposed methods eliminate the need for explicit regularization parameter specification while maintaining robust feature selection capabilities and effective multicollinearity mitigation, demonstrating high accuracy and reliability across high-dimensional datasets.
It is argued that without additional, correctly specified, side information, any EO+ method can result in at most second-order improvements, and a negative conclusion is provided, namely ``no free lunch is possible", on the statistical power of EO+.
Real-world regression problems often involve noise, redundancy, multicollinearity, and nonlinear relationships that limit the effectiveness of classical models. This study investigates Type-1 Fuzzy Functions (T1FF) combined with several feature selection strategies, with particular emphasis on the integration of Lasso regression into the T1FF framework, which has not been directly examined in prior research. By incorporating Fuzzy C-Means-based membership degrees into the modelling process, T1FF provides a flexible way to capture uncertainty and nonlinear structure without relying on expert-defined fuzzy rules. The proposed framework was evaluated on six datasets, namely Boston, Auto, College, Steel Fatigue Strength, Fish Price, and Smart Pressure Control, using RMSE and MAPE as performance criteria. The results show that T1FF-based models generally outperform classical LM, Ridge, and Lasso models on most datasets, although the best-performing T1FF variant varied depending on dataset characteristics. In particular, the Lasso-based T1FF model yielded competitive results overall and achieved especially strong performance on the College and Fish Price datasets, while Full and Forward T1FF methods showed the most consistent MAPE-based ranking across datasets. Overall, the findings indicate that integrating feature selection and regularization methods into the T1FF framework provides a promising and flexible approach for regression modelling on complex real-world data.
M. Şahin, N. Tak· Karadeniz Fen Bilimleri Derg...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.