SchemaGUI, a template-based benchmark for controllable GUI generation evaluation, synthesizing paired natural language instructions and deterministic function-call references from parameterized interface schemas can generate thousands of deterministically annotated tasks in seconds without human labeling.
Jiarui Dong, Yin Cai, Zhouhong Gu et al.· 0 citations
This work introduces Instruction-Followed Function Calling (IFFC), a novel framework that decouples function-calling logic from the primary LLM and delegates it to a dedicated smaller model operating within the instruction-following paradigm, establishing a new paradigm for reliable, resource-efficient function calling in edge-computing scenarios.
Yalda Taheri, Mohammad Hassan Heydari, Erfan Naaman et al.· 0 citations
Dual-Layer Agentic Memory is proposed, a framework that shifts memory management to the write phase through cost-aware epistemic routing and periodic parametric consolidation, allowing the router to adaptively suppress redundant writes as the model's epistemic boundaries evolve.
Wenzhi Li, Dong Nie, Ruiyi Lan et al.· 0 citations
Two state-of-the-art multimodal models, Gemma-3 and Qwen-VL, are assessed on their ability to interpret mechanical problem images by eliciting a step-by-step chain of thought (CoT) and a final answer, and final answers are compared to verified solutions to measure accuracy.
Henry Fordjour Ansah, Shreya Banerjee, Pranish Ghimire· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
NeuroPrefetcher is presented, a storage-backed LLM inference system that exploits that MLP activity during autoregressive decoding has strong temporal locality, and achieves 7.9-12.0x speedup over llama.cpp across constrained memory budgets.
Nobel Dhar, Md Romyull Islam, Xuechen Zhang et al.· 0 citations
This work adapts Hugging Face's SmolVLA for Universal Robots lightweight robots, and releases the open-source repository ROS2SmolVLA that implements an interface for ROS 2 to SmolVLA, and makes it applicable for industrial-grade hardware.
Nils Mandischer, Noah Böckmann, Ludwig Holl et al.· 0 citations
A training-free, model-agnostic memory graph that repairs continuity entirely at the prompt-text layer by extracting a multi-dimensional state for every shot, retrieving only causally prior context over the resulting graph, filtering it selectively, and injecting the surviving constraints by natural-language prompt rewriting.
Jiaqi Liu, Maolin Ran, Xiaoyan Lu et al.· 0 citations
This work investigates two representative model families, LLaVA-1.5 and Qwen2.5, and provides a token-, layer-, and head-level account of how VLMs transform object grounding into spatial relations, showing that knowing where objects are is not equivalent to knowing how they relate.
Xiwei Liu, Yulong Li, Xinlin Zhuang et al.· 0 citations
It is shown that policies within a bounded $\chi^2$ divergence from the proxy-feasible reference distribution admit an $N$-independent safety-hacking bound, and instantiate this general coverage-control principle with constrained pessimistic sampling.
Experimental results on instruction-following tasks show that the proposed self-distillation framework, SelFusion, substantially outperforms other KD methods with external LLM and DLM teachers, providing a practical path toward improving DLM generation quality.
Hyeongsoo Lim, Jinyoung Kim, Eunjoo Seo et al.· Annual Meeting of the Associ...· 0 citations
General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimensions. \emph{Environment Scaling} expands the diversity and verifiability of executable file, search, and code environments, while \emph{Agentic Coordination Scaling} trains agents to decompose long-horizon tasks, delegate parallel work, integrate asynchronous results, and replan. A shared execution harness and AgentOS maintain task state and provenance across tools and agents, and training turns environment trajectories and coordination traces into reliable behavior. Across complex professional work, finance, scientific research, mathematics, coding, and search, Apodex 1.1 reaches the leading performance band despite using a substantially smaller model than many frontier systems. The 35B-parameter Apodex 1.1 Mini further retains strong working capability in a locally deployable form. These results ground agentic intelligence in useful, verifiable work completed over time and advance our goal of building a \emph{Heavy-Duty Solver} for ambitious, long-running tasks.
Apodex Team B. An, B. Li, B. Wang et al.· 0 citations
It is argued that WG is more plausible as an adversarial threat-requiring careful data engineering-rather than as a significant hazard inherent to routine fine-tuning.
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.