This work proposes CADE (Contrastive Adaptive Debias Ensemble), a training-free, plug-and-play method that leverages modality-specific answer priors that yields significant gains on the proposed benchmark, which can foster the development of more fair and reliable AI systems for sustainable development.
Zihang Lin, Huaiyuan Qin, Mu Yang et al.· arXiv.org· 0 citations
This work identifies a critical gap: when unauthorized tools are visible in an agent's context, models select them in 48-68% of adversarial scenarios, even when explicitly instructed not to, and proposes a proxy-enforced attribute-based access control layer for MCP that filters tool registries at discovery time.
This work presents SkillSafetyBench, a runnable benchmark for evaluating skill-facing safety failures, and suggests that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.
This work proposes ABC: Any-Subset Autoregressive Models via Non-Markovian Diffusion Bridges in Continuous Time and Space, and derives SDE dynamics via changes-of-measure on path space, yielding another advantage: path-dependent conditioning on arbitrary subsets of the state history and/or future.
Gabriel Guo, Thanawat Sornwanee, L. Hao et al.· arXiv.org· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
G-Loss is presented, a graph-guided loss function that incorporates semi-supervised label propagation to use structural relationships within the embedding manifold to build a document-similarity graph that captures global semantic relationships.
LLMs have limitations when it comes to cultural coverage and competence, and in some cases, show specific cultural biases. Although prior studies have examined the cultural capabilities of LLMs, none have specifically investigated their regional preferences in generic culture-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ), with questions available in 24 languages. We evaluate LLMs by prompting them to answer questions from CROQ and provide a sample location. The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan in their answers. Moreover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs. Low-resource languages, on the other hand, show more inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training. Dataset available at https://huggingface.co/datasets/HiTZ/CROQ
Joseba Fernandez de Landa, Carla Pérez-Almendros, J. Camacho-Collados· arXiv.org· 1 citation
The benefits of integrating prior bias when considering locomotion tasks with simple hinge actuators are demonstrated and a Parameter Impact metric is introduced which showcases diminishing returns for MLPs but not for CPGs.
Kevin Godin-Dubois, Anil Yaman, Anna V. Kononova· arXiv.org· 0 citations
This work introduces CodeRQ-Bench, the first benchmark for evaluating LLM reasoning quality across three coding task categories: generation, summarization, and classification, and proposes VERA, a two-stage evaluator that combines evidence-grounded verification with ambiguity-aware score correction.
Yuangang Li, Justin Tian Jin Chen, Ethan Yu et al.· arXiv.org· 1 citation
PolicyLong is proposed, shifting data construction towards a dynamic on-policy paradigm, by iteratively re-executing data screening (entropy computation, retrieval, and verification) using the current model, which ensures the training distribution tracks evolving capabilities, yielding an emergent self-curriculum.
A novel Dual Self-Consistency Reinforcement Learning optimization paradigm is introduced, which utilizes Round-Trip Verification to penalize degenerate code and boost overall self-consistency in TikZ code.
Juekai Lin, Yun Zhu, Honglin Lin et al.· arXiv.org· 5 citations
This paper proposes a camera-agnostic, one-shot, post-training pruning method for 3D Gaussian splats that relies solely on attribute-derived neighbourhood descriptors, and introduces a hybrid descriptor framework that captures structural and appearance consistency directly from the splat representation.
Peter O. Fasogbon, Ugurcan Budak, P. R. Alface et al.· arXiv.org· 0 citations
This work formalizes Active Testing in NLP and conducts an extensive benchmarking of existing approaches across 18 datasets and 4 embedding strategies spanning 4 different NLP tasks, revealing variations in method effectiveness across different data characteristics and task types.
Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu et al.· arXiv.org· 1 citation