Skip to content
Preprint

HugSelect: An Explainable Multi-Criteria Decision-Support Framework for foundation-model selection

Aug 2026 · 0 citations · 51 references
Computer Science

TL;DR

It is argued that foundation-model selection should be treated as an explicit, auditable software-component selection task rather than as keyword search, popularity ranking, or opaque conversational advice, and HugSelect, an explainable decision-support framework for foundation-model selection is proposed.

Abstract

Foundation models are increasingly reused as software components, making model selection a critical software-engineering decision. Current model hubs primarily support discovery through popularity metrics, often neglecting functional capabilities, operational constraints, and community-perceived quality. We argue that foundation-model selection should be treated as an explicit, auditable software-component selection task rather than as keyword search, popularity ranking, or opaque conversational advice. This paper proposes HugSelect, an explainable decision-support framework for foundation-model selection. HugSelect builds a knowledge base of 71,274 models by combining repository metadata, extracted functional capabilities, and perceived quality attributes derived from community discussions into a unified pipeline. It ranks candidate models using a weighted additive model that exposes criterion-level score decompositions. We evaluated HugSelect through pipeline validation, comparative case studies against four commercial LLM-based recommendation systems (44 scenarios), fine-grained ablation, and an exploratory user study (n = 10). Extraction pipelines achieved an F1 score of 0.801 for functional features and an accuracy of 0.84 for quality-attribute mapping. HugSelect achieved a model-level Coverage@10 of 0.61 and family-level Coverage@10 of 0.91, showing recommendation quality comparable to that of the evaluated commercial systems, with no significant overall differences in ranking quality, while providing stable, traceable, and inspectable reasoning. Ablation confirmed that functional features were the main driver of retrieval accuracy, and preliminary user feedback suggests that the framework is useful and intuitive.

View source

Similar papers

Open access Sep 2026

Hybrirank: A Dual-Axis Algorithm Ranking Framework via Dataset-Based and Preference-Guided Rank Fusion

Algorithm selection remains a challenging task due to the absence of standardized tools that support informed decision-making for identifying suitable algorithms for AI and ML tasks. It plays a crucial role in designing and deploying intelligent systems. This paper presents Hybrirank, a hybrid framework that ranks algo...

Pavan Manikanta Raghava Kasturi, Vithya Ganesan, N. Kirubakaran · 0 citations
Open access Sep 2026

Grounding Techniques in LLM-Based Recommender Systems: A Systematic Literature Mapping

The integration of Large Language Models is transforming recommender systems, offering unprecedented capabilities for complex reasoning and natural language generation. However, their propensity to generate hallucinations (incorrect or invented information) compromises reliability and user trust, limiting their adoptio...

Andrés Felipe Solis Pino, N. Duque-Méndez, Pablo H. Ruiz et al. · 0 citations
#artificial intelligence Book Open access Aug 2026

The Utility of LLMs in Recommender Systems Explanation Evaluation

Explanations play a crucial role in creating trustworthy recommender systems (RS), yet choosing a good explanation method presents challenges. Many explanation methods exist, but little guidance exists on which is best for which setting. Existing explanation generation methods often produce abstract outputs that requir...

Kathrin Wardatzky, Oana Inel, Luca Rossetto et al. · 0 citations
Preprint Aug 2026

Comparing Domain-Model Similarity Metrics Against Human Expert Ratings

Domain models are a primary artefact in model-driven software engineering, where they capture the shared understanding between stakeholders and serve as the contractual basis for downstream software development. Automatic comparison of these semantic models has diverse application areas such as requirements engineering...

Vasiliy Seibert · 0 citations
Open access Sep 2026

WikiConstraintCalibration: Content-Aware Cost Preference Estimation for Constraint Threshold Selection in Wiki-Grounded LLM Agents

Wiki-grounded LLM agents enforce output constraints through safety rules and scope boundaries derived from structured knowledge bases, yet selecting an appropriate constraint scorer threshold $\theta$ remains challenging. Since outputs are blocked when $s(o) \geq \theta$, a low threshold may over-restrict safe response...

Bai-Ling Zhang · 0 citations
Open access Sep 2026

An Explainable Multi-Criteria Decision-Making Framework for Evaluating Malware Detection Models Across Heterogeneous Datasets

The diversity of today’s malware and the conflicting criteria for predictive performance, computational efficiency, dependability, and interpretability have made the choice of a suitable malware detection model more complicated. Existing research focuses predominantly on predictive performance, while the multidimension...

Husam Jasim Mohammed, Riyadh Rahef Nuiaa Alogaili, Mohanad S. Jabbar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.