Skip to content

Category

artificial intelligence

6,499 papers

#artificial intelligence Preprint Aug 2026

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

This work formalizes probabilistic alignment as a distributional criterion for world models and introduces PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics, and introduces PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors.

Yuandong Pu, Le Zhuo, Sayak Paul et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents

LoopHarness is presented, which restores a persistent, non-decaying safety state at the loop level at the loop level, and gives a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.

Chenmin Wu, H. Jia, Yang Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

LiveVVT: High-Fidelity Video Virtual Try-On in Real Time

LiveVVT is introduced, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation, and a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference.

Yushe Cao, Shikun Feng, Ru-Xiang Duan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Comparing Chunking and Embedding Strategies for Turkish RAG Systems

This work compares Turkish document question answering across three chunking strategies, five embedding models, and two LLMs, over three documents with contrasting layouts, finding the faster LLM is not the more accurate one.

Mustafa Sertac Turkel, Fatma Nur Korkmaz, Ahmet Tugrul Bayrak · 0 citations
#artificial intelligence Preprint Jul 2026

Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation

Evaluating 13-21 models across six presentation operationalizations and four task-domain operationalizations suggests that, despite confounds, some models possess practical SGTR capabilities, and that SGTR should be monitored and considered in the design of safety-critical AI applications.

J. St-Amand, Callum Canavan, S. Imran et al. · 0 citations
#artificial intelligence Preprint Jul 2026

AI Models Can Predict and Collaboratively Modulate Human Memory Search

This study explores and evaluates the ability of LLMs to follow and enhance human mental trajectories during semantic memory search and demonstrates that an LLM's abilities to track and predict human memory trajectories in this task exceed those of other humans.

Eric Lacosse, Mariana Duarte, Graham Todd et al. · 0 citations
#artificial intelligence Preprint Aug 2026

TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development

This work introduces TraceML, which pairs human and agent work on the same competitions under one version-level schema, and releases the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.

J. Yan, Weiwei Sun, Si-Jie Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

MathAdv: What Theorem Provers Know, Reason, Formalize, and Generalize

This work introduces MathAdv, a diagnostic benchmark spanning 13 domains across undergraduate- and graduate-level mathematics, and shows how component-wise evaluation can reveal model capabilities and failure modes that aggregate theorem-proving accuracy obscures.

Jiajie Yuan, C. Lockhart, Xiao-Yun Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

On-policy Distillation with Verifiable Reward

This work reformulates the implicit reward of sampled-token OPD based on trajectory correctness, then applies a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, making it readily combinable with any policy gradient algorithm, such as GRPO.

Wenze Lin, Jiale Zhao, Xi-Tai Jiang et al. · 0 citations

From tech blogs

See all →

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.