Two language-model-based strategies are proposed for semantic code document segmentation, including a line-by-line approach that classifies each line of code separately before grouping the results into functional units, and a range-based approach that aims to directly determine groups of code lines from the input.
Abdelhalim Hafedh Dahou, A. Scherp, Sebastian Kurten et al.· Proceedings of the 2026 ACM...· 0 citations
This paper proposes the Evaluation Context Protocol (ECP), an early-stage, vendor-neutral framework intended to act as a portable evaluation contract layer for agentic systems and describes an open-source reference implementation that includes adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI.
Small language models can grade open-ended examination answers as reliably as substantially more expensive models when they grade against an explicit rubric. We test this claim as the design principle behind any-to-bench: a frontier model reads source documents once, at ingestion, to extract each question and its rubric; lower-cost models then perform all repeated grading work. We evaluate six cost-efficient model configurations from two model families at three reasoning-effort levels. Each configuration answers 24 open-ended examination questions, and each also grades every answer sheet three times, yielding 3,456 per-question grades. Scores depend overwhelmingly on the answer being graded: answer identity explains 95.6% of score variance, whereas judge identity explains only 0.2%. Raising a writer's reasoning effort moves earned scores by as much as 0.143 of full marks, while raising a judge's reasoning effort moves assigned scores by at most 0.006. Six frontier-tier judges, added as a check, reproduce these scores and are no more reliable as a panel. Two ablations then decompose the rubric on the same questions and answers. Removing its criteria and levels while keeping the official answer changes nothing measurable. Removing the official answer as well collapses reliability (ICC 0.888 to 0.628), inflates scores, and makes judge reasoning effort matter again. The rubric is what decouples grading from judge intelligence, and within the rubric the official answer does nearly all the work. We find no evidence of length preference or same-family preference under rubric-anchored grading.
This work mathematically formalizes how transformers execute abstract reasoning and provides a novel, strictly geometric signature for evaluating factual reliability, proving that deep LLM latent spaces natively organize into Small-World networks.
Md. Faiyaz Abdullah Sayeedi· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
An augmented reality learning medium integrated with a deep-learning-oriented pedagogy for the Solar System topic supports the feasibility and educational promise of combining interactive AR visualization with cognitively engaging pedagogy.
Bagas Brilian Ramadhan, Mei Wulan, S. Yamtinah et al.· Kognisi: Jurnal Ilmu Kegurua...· 0 citations
BACKGROUND
Personalized meal planning by registered dietitian nutritionists (RDNs) is time-intensive. Large language models (LLMs) may automate drafting meal plans, but their nutritional accuracy in clinical practice is uncertain.
METHODS
In this proof-of-concept study, five outpatient RDNs and four LLMs (Gemini, CoPilot, ChatGPT 4.0, and customized ChatGPT 4.0) each generated 3-day meal plans for five validated clinical scenarios. Effectiveness was defined as accuracy in meeting pre-specified energy, protein, carbohydrate, fat, and sodium targets. Time to create plans and RDN comfort (self-rated confidence in nutritional accuracy and clinical appropriateness on 1-5 Likert scale) were recorded. Three independent RDNs, blinded to source, analyzed nutrient content using Nutritionist Pro. Group differences were assessed with t-test and ANOVA.
RESULTS
All LLMs and RDNs produced feasible meal plans. LLMs generated meal plans in under 1 min, whereas RDNs required a mean of 44 min per scenario. RDNs reported comfort levels ranging from 3.8 to 4.8. Across most scenarios, LLM plans delivered a smaller proportion of requested energy than RDN plans, which more consistently approached energy targets. Both groups performed similarly for the Mediterranean diet scenario. Overall, protein accuracy did not differ. However, in chronic kidney disease, LLMs undershot the guideline-based protein target, while RDNs tended to modestly exceed it. Accuracy for low-carbohydrate, fat, and sodium diets was comparable.
CONCLUSION
LLMs can rapidly generate clinically plausible meal plans but are less reliable than RDNs in achieving prescribed energy and selected macronutrient goals. Prompt precision is essential for nutrient-specific targets. A hybrid model in which RDNs refine LLM-generated drafts may leverage efficiency without sacrificing clinical accuracy.
M. Mundi, Osman Mohamed Elfadil, Danielle P. Johnson et al.· Nutrition in clinical practi...· 0 citations
This paper presents FlashAttention-V, a blocked FlashAttention for scalable vector architectures that adapts efficiently from short to very long vectors by exploiting parallelism across attention heads, inter-head packing to enable efficient utilization of vector lengths beyond the head dimension, and improving vector register utilization and memory access locality.
This paper introduces *TranslatePsy-AfriSLM*, a collection of open-source MT resources for 19 Sub-Saharan African languages, including curated parallel data, African-specialized synthetic data, and a family of fine-tuned SLMs.
Milan Gritta, Patrik Lambert, Jihye Back et al.· 0 citations
COSTA leverages the domain gap through proven test-time adaptation, and groups each batch of target-domain points into a small set of semantic clusters based on the similarity distribution in the adapted feature space, and propagates high-confidence pseudo labels obtained from an open-vocabulary vision-language model to all points through cluster-level voting.
Yanghong Lin, Li Fang, Tianyu Li et al.· 0 citations
The results suggest that bitsandbytes 4-bit quantization can impose an additional cost on applications relying on long, updatable, semantically dense contexts, even when aggregate benchmark accuracy appears largely unaffected.
Large Language Models show potential in their diagnostic accuracy and consequent ability to reduce clinician burden, and may provide the greatest benefit when used to optimise referral quality at source, improving both clinician and potentially LLM triage downstream.
K. Surendran, I. Aziz, Glyndwr Jenkins· Current Surgery Reports· 0 citations
The Institutional Newspapers Pipeline is presented, a modular system designed to extract high-quality, structured datasets from historical newspaper scans that was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware.
Matteo Cargnelutti, Catherine Brobston, Eben English et al.· 0 citations
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
MIT News · Artificial Intelligence· news.mit.eduAug 31, 2026
With millions of users across the world, Julia has been used to conduct cutting-edge research and to design new drugs, jet engines, heat pumps, and more.