A softmax model is studied that decouples the concept point and the labeling scheme in multiple instance regression data and derives a parametric family of iterations in the noiseless limit, which generalizes a method known as the EM-DD algorithm.
Abstract
In multiple instance regression (MIR) data are organized into bags (collections of instances in feature space) and the goal is to learn a mapping that assigns labels to bags. A typical assumption is that there is a so-called concept point in feature space, the proximity to which dictates the bag label. Motivated by modern MIR architectures which are based on attention, we study a softmax model that decouples the concept point and the labeling scheme. The two are respectively determined by a query direction and a value direction value in feature space. This problem isolates a basic challenge of learning both the query and value vectors from bag-level supervision. From this model we derive a parametric family of iterations in the noiseless limit, which generalizes a method known as the EM-DD algorithm. We then derive concentration results for the MLE estimators of the query and value vectors obtained from a random selection of instances. Our result for the value vector shows that a single random initialization of the value vector already points in the correct direction on average, so that a polynomial (in the number of instances per bag and the feature dimension) number of bags is enough for the EM algorithm to converge in $O(1)$ steps with high probability. A key aspect of this analysis is the interplay between concentration of empirical covariance matrices and extremal statistics arising from the selection rule.
Evaluating the performance of Greedy K-center across a variety of metric spaces shows that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.
This paper presents a novel perspective on training methods that diverges from traditional fine-tuning approaches for deep learning-based natural language processing models by considering the weight of each instance. Specifically, we propose a novel Instance Weight Estimation Model (IWEM) that calculates the weight of...
Seunghyun Park, J. Park, Byung-Won On et al.· IEEE Access· 0 citations
This study innovatively analyzes the hyperparameters of tree-based linear methods and suggests an efficient and effective guideline that leads to consistent improvements across datasets, thereby strengthening tree-based linear methods as a stronger XMTC baseline.
This paper considers a particular type of distribution shift, label shift, and develops an e-icient inference procedure for general parameters characterizing the unlabeled target population, and proposes a progressive estimation strategy that unfolds in three stages: an initial heuristic guess, a consistent estimation,...
This paper introduces three modular components: an attention mechanism that weights the contribution of retrieved cases, a locality-aware regularizer that favors label-similar neighbors, and an optional case adaptation module that refines the retrieved estimate.
Xiaomeng Ye, Yu Wang, David B. Leake et al.· Proceedings of the Thirty-Fi...· 0 citations
This review consolidates the landscape of CP adaptations for MLL under a unified framework, examining the types of outputs and guarantees they provide, where label dependencies are incorporated, and how inference cost scales with the number of labels.