Automated Feature Engineering Techniques for Tabular Data
Abstract
AFE is an automated feature engineering system that has become a key enabler to scalable and high-performing machine learning systems with scalable systems that run on tabular data. The conventional feature engineering makes excessive use of domain, trial and error and iterative optimization, which are computed consuming time, error prone and hard to repeat. As data-driven applications in finance, healthcare, manufacturing, and e-commerce have been exponentially increasing, there is an increasing need in automated, systematic, and reliable methods of features construction. The overall objective of automated Feature Engineering methods is to generate, transform, select, and optimize features adventurously, from raw tabular data, with minimal human intervention, and at a higher predictive efficiency. The paper contains a complete detailed analysis of automated feature engineering approaches to tabular data with references to their theoretical principles, algorithmic approaches, and real-world examples. The paper expounds on rule construction based feature construction, statistical construction, deep learning based representation, evolutionary learning, and reinforcing learning methods, and end to end AutoML. An intricate literature review shows major achievements, comparisons, and unresolved issues. The suggested methodology defines the three branches of feature generation, selection and evaluation as a single automated pipeline via mathematical representation and algorithmic processes. The effectiveness of automated feature engineering is proved by experimental results that show that the method can enhance the accuracy and robustness of models as well as improve generalization with respect to the multiple benchmark datasets. Lastly, issues of limitations, interpretability, computational trade-offs, and research directions are discussed in the paper. The given publication meets the IEEE publication standards and offers well-organized, high-quality information to a researcher or an organization practitioner dealing with tabular data analytics.