May 2026· arXiv.org· Vol abs/2605.15723· 0 citations· 64 references
Computer Science
TL;DR
Graph-Optimized Multimodal Alignment (GOMA), which lets a jointly trained model support single-modality and dual-attribute retrieval through a task-specific readout, achieves state-of-the-art performance on all 14 primary measures against 14 external methods.
Abstract
Multimodal retrieval uses images, text, and object relationships to answer different questions about the same collection. A query may seek an object's paired description, another object in the same category, or an object connected by an observed relationship. These goals rely on different notions of relevance. Paired matching requires object-specific distinctions, whereas cross-object retrieval benefits from relational agreement. Existing methods learn strong cross-modal correspondence or propagate information over a graph, but a shared output regularized toward neighbors can weaken identity distinctions. Moreover, gains from graph regularization can diminish after uniform graph propagation. We introduce Graph-Optimized Multimodal Alignment (GOMA), which assigns these roles to two connected embeddings. Each modality produces a content embedding directly supervised for paired identity and a semantic embedding jointly trained with cross-modal pairs and observed relationships. For complete records, both embeddings form an initial fused representation, semantic agreement sets positive weights on observed edges, and restart graph propagation reinjects this initial signal. This design lets a jointly trained model support single-modality and dual-attribute retrieval through a task-specific readout. Across six datasets and four tasks, GOMA achieves state-of-the-art performance on all 14 primary measures against 14 external methods. Controlled comparisons further show how separate supervision, graph regularization, and semantic-guided graph propagation shape the final representation and align the learned signal with each retrieval target.
This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.
P. Abrahamsson, O. Salo, Jussi Ronkainen et al.· arXiv.org· 727 citations· ⚡54
The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.
M. Pikkarainen, Jukka Haikara, O. Salo et al.· Empirical Software Engineeri...· 401 citations· ⚡48
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
The possibility of inferring high-dimensional data inference in a model that consists of a prior and an auxiliary differentiable constraint given some additional information is considered, thereby allowing a range of potential applications in adapting models to new domains and tasks.
Alexandros Graikos, Esmeralda S. Whitammer, N. Jojic et al.· Neural Information Processin...· 316 citations· ⚡15
Radiology AI is evolving beyond report generation. CARE-X explores a unified approach that combines flexible reasoning, calibrated predictions, and measurement-based tools for chest X-ray interpretation. The post Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement appeared first on Microsoft Research.