This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models'training cutoffs.
This work proposes SciJEPA, a citation-free framework that learns through asymmetric within-document prediction: title and abstract representations are used to predict method representations, and method representations are used to predict conclusion representations.
You Zuo, Éric de la Clergerie, Benoît Sagot· 0 citations
The results show that sycophancy can corrupt the reasoning chain independently of the final answer, so answer-level evaluation alone is insufficient, and a failure taxonomy separating reasoning-chain from answer-level sycophancy is introduced, and a complementary sentence-level taxonomy locating where in the chain drift first emerges.
Mahir Numayeer Islam, G. Okuyama, Nikolaus Siauw et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
Layer-by-layer analysis of GenAI VP dialogue logs can reveal process patterns associated with high rated history taking and support process-focused feedback in medical education.
Xinyu Li, Zijian Li, Mengyu Xia et al.· 0 citations
Sequence-to-edit approaches make grammatical error correction (GEC) efficient and locally interpretable by predicting edit labels over the input rather than generating a full corrected sentence. Their interpretability, however, is primarily operational: a label specifies how the string should change, but a single edit vocabulary does not always reveal the type of correction being made. We propose STAGEET, a stage-wise typed edit-tagging framework that reorganizes Seq2Edit supervision into typed executable stages and extends edit operations to correction categories. STAGEET decomposes correction into an ordered sequence of medium-grained typed stages; each stage predicts from its own label space, rewrites the current hypothesis once, and passes the resulting intermediate sentence to the next stage. We instantiate the framework as both an end-to-end shared-encoder multi-head model with stage-specific adapters and a fully specialized variant with one independent tagger per stage. Experiments on QALB-2014 and ZAEBUC show that category-aware staged correction retains competitive edit-based GEC performance while exposing a more inspectable correction trajectory, and attains state-of-the-art results on QALB-2014.
This work curates a syllabus-aligned QA dataset based on NCERT textbooks for classes 9-12, capturing the content, context, and teaching style of Indian curricula, and introduces GurukulAI, an open-access platform that enables Indian students to chat with the model, get doubts cleared, practice exam-style questions, receive contextual answers, and interact in both English and Hindi.
I. Narang, Sneha S. Gosai, Mayank Singh· 0 citations
This work ground perceptual memory in the model, decomposing recall into two subproblems: a vision-language model grounds the referent in context (what and where), and a dedicated encoder extracts an identity key (who), stored as one inline token read by attention at generation with no external round-trip.
The research here utilizes Natural Language Processing methods like Named Entity Recognition (NER), BERTopic modeling, and Knowledge Graph development in Neo4j to extract, categorize, and visualize important concepts based on translated versions to make ancient Indian medical wisdom more accessible and understandable.
M. Rajeevan, B. Devi, V. Anoop et al.· 2 citations
Multi-Objective In-context Knowledge Editing (MO-IKE), a multi-objective RL algorithm that formulates prompt construction for in-context knowledge editing as a Constrained Markov Decision Process, enabling more balanced and globally coherent prompt construction.
Xu-Zhong Wang, Maiqi Jiang, Tejal Nair et al.· 1 citation
LUCAID is an agentic AI system for precision lung cancer pathology that combines diagnostic reasoning with nine modules that cover the full routine workflow, from quality control, tumor detection and segmentation, histological subtyping, tumor microenvironment profiling, tumor cellularity quantification, and predictive biomarker scoring to automated structured report generation.
M. Eich, K. Standvoss, Timo Milbich et al.· 0 citations
The core design principle of DELE-w0.5 is to model how the physical world changes under robot actions, rather than how its visual appearance evolves frame by frame, which enables cheaper training and low-latency inference.
Fenghao Lei, Zhixiong Huang, Long Yang et al.· 0 citations