Redundant, irrelevant, and noisy features make it very hard to analyse high-dimensional data, especially when the number of features is much larger than the number of samples. Conventional feature selection methods, such as filter, wrapper, and embedded methods, are unable to balance predictive accuracy, feature subset...