Skip to content

Category

artificial intelligence

6,499 papers

SDGBiasBench: Benchmarking and Mitigating Vision-Language Models' Biases in Sustainable Development Goals

This work proposes CADE (Contrastive Adaptive Debias Ensemble), a training-free, plug-and-play method that leverages modality-specific answer priors that yields significant gains on the proposed benchmark, which can foster the development of more fair and reliable AI systems for sustainable development.

Zihang Lin, Huaiyuan Qin, Mu Yang et al. · 0 citations

Prompts Don't Protect: Architectural Enforcement via MCP Proxy for LLM Tool Access Control

This work identifies a critical gap: when unauthorized tools are visible in an agent's context, models select them in 48-68% of adversarial scenarios, even when explicitly instructed not to, and proposes a proxy-enforced attribute-based access control layer for MCP that filters tool registries at discovery time.

Rohith Uppala · 1 citation

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces

This work presents SkillSafetyBench, a runnable benchmark for evaluating skill-facing safety failures, and suggests that agent safety depends not only on model-level alignment, but also on how agents interpret skills, trust workflow context, and act through executable environments.

Chang Jin, Anr'an W'ang, Zeming Wei et al. · 11 citations · ⚡1

ABC: Any-Subset Autoregression via Non-Markovian Diffusion Bridges in Continuous Time and Space

This work proposes ABC: Any-Subset Autoregressive Models via Non-Markovian Diffusion Bridges in Continuous Time and Space, and derives SDE dynamics via changes-of-measure on path space, yielding another advantage: path-dependent conditioning on arbitrary subsets of the state history and/or future.

Gabriel Guo, Thanawat Sornwanee, L. Hao et al. · 0 citations

G-Loss: Graph-Guided Fine-Tuning of Language Models

G-Loss is presented, a graph-guided loss function that incorporates semi-supervised label propagation to use structural relationships within the embedding manifold to build a document-similarity graph that captures global semantic relationships.

Aditya Sharma, Vinti Agarwal, Rajesh Kumar · 0 citations

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

LLMs have limitations when it comes to cultural coverage and competence, and in some cases, show specific cultural biases. Although prior studies have examined the cultural capabilities of LLMs, none have specifically investigated their regional preferences in generic culture-related questions. In this work, we propose a new dataset based on a comprehensive taxonomy of Culture-Related Open Questions (CROQ), with questions available in 24 languages. We evaluate LLMs by prompting them to answer questions from CROQ and provide a sample location. The results show that, contrary to previous cultural bias work, LLMs show a clear tendency towards countries such as Japan in their answers. Moreover, our results show that when prompting in languages such as English or other high-resource ones, LLMs tend to provide more diverse outputs. Low-resource languages, on the other hand, show more inclinations towards answering questions highlighting countries for which the input language is an official language. Finally, we also investigate at which point of LLM training this cultural bias emerges, with our results suggesting that the first clear signs appear after supervised fine-tuning, and not during pre-training. Dataset available at https://huggingface.co/datasets/HiTZ/CROQ

Joseba Fernandez de Landa, Carla Pérez-Almendros, J. Camacho-Collados · 1 citation

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks

This work introduces CodeRQ-Bench, the first benchmark for evaluating LLM reasoning quality across three coding task categories: generation, summarization, and classification, and proposes VERA, a two-stage evaluator that combines evidence-grounded verification with ambiguity-aware score correction.

Yuangang Li, Justin Tian Jin Chen, Ethan Yu et al. · 1 citation

PolicyLong: Towards On-Policy Context Extension

PolicyLong is proposed, shifting data construction towards a dynamic on-policy paradigm, by iteratively re-executing data screening (entropy computation, retrieval, and verification) using the current model, which ensures the training distribution tracks evolving capabilities, yielding an emergent self-curriculum.

Junlong Jia, Jiangnan Zhou, Ziyang Chen et al. · 0 citations

Camera-Agnostic Pruning of 3D Gaussian Splats via Descriptor-Based Beta Evidence

This paper proposes a camera-agnostic, one-shot, post-training pruning method for 3D Gaussian splats that relies solely on attribute-derived neighbourhood descriptors, and introduces a hybrid descriptor framework that captures structural and appearance consistency directly from the splat representation.

Peter O. Fasogbon, Ugurcan Budak, P. R. Alface et al. · 0 citations

Select, Label, Evaluate: Active Testing in NLP

This work formalizes Active Testing in NLP and conducts an extensive benchmarking of existing approaches across 18 datasets and 4 embedding strategies spanning 4 different NLP tasks, revealing variations in method effectiveness across different data characteristics and task types.

Antonio Purificato, Maria Sofia Bucarelli, Andrea Bacciu et al. · 1 citation

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.