This work formalizes probabilistic alignment as a distributional criterion for world models and introduces PAWBench, a benchmark for evaluating video generators as stochastic samplers of world dynamics, and introduces PAWEval, an outcome-level protocol that converts repeated video rollouts into empirical distributions over possible physical behaviors.
Yuandong Pu, Le Zhuo, Sayak Paul et al.· 0 citations
LoopHarness is presented, which restores a persistent, non-decaying safety state at the loop level at the loop level, and gives a complete evaluation protocol on native Agent-SafetyBench tasks with paired clean and attacked episodes, an outer-state attack suite whose decisive evidence exists only across iterations, per-module ablations, and an adaptive white-box red team.
LiveVVT is introduced, a rolling streaming diffusion framework that preserves bounded bidirectional modeling within causal recurrent generation, and a progressive distillation framework integrating bidirectional VVT learning, teacher-trajectory regression for causal few-step adaptation, and Collaborative Matching Distillation, which couples teacher-distribution matching with rolling flow matching on real videos to align optimization with recurrent inference.
Yushe Cao, Shikun Feng, Ru-Xiang Duan et al.· 0 citations
This work presents an end-to-end AI system that collapses the software-to-silicon stack into a single optimization loop, where hardware and software are co-designed and verified under one objective.
Architect Labs· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
This work compares Turkish document question answering across three chunking strategies, five embedding models, and two LLMs, over three documents with contrasting layouts, finding the faster LLM is not the more accurate one.
Mustafa Sertac Turkel, Fatma Nur Korkmaz, Ahmet Tugrul Bayrak· 0 citations
Evaluating 13-21 models across six presentation operationalizations and four task-domain operationalizations suggests that, despite confounds, some models possess practical SGTR capabilities, and that SGTR should be monitored and considered in the design of safety-critical AI applications.
J. St-Amand, Callum Canavan, S. Imran et al.· 0 citations
This study explores and evaluates the ability of LLMs to follow and enhance human mental trajectories during semantic memory search and demonstrates that an LLM's abilities to track and predict human memory trajectories in this task exceed those of other humans.
Eric Lacosse, Mariana Duarte, Graham Todd et al.· 0 citations
This work introduces TraceML, which pairs human and agent work on the same competitions under one version-level schema, and releases the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.
A consolidated memory that states a decision constraint and whose source record has since been superseded by a record that withdraws it is studied: provenance is immutable, the current record has changed, and the memory is stale.
This work introduces MathAdv, a diagnostic benchmark spanning 13 domains across undergraduate- and graduate-level mathematics, and shows how component-wise evaluation can reveal model capabilities and failure modes that aggregate theorem-proving accuracy obscures.
Jiajie Yuan, C. Lockhart, Xiao-Yun Liu et al.· 0 citations
SpecMine lets the community study, for the first time, how software is specified in the age of AI agents through two censuses: a broad census of spec.md files and a census-wide index of typed references.
Shyam Agarwal, Anmol Singhal, Travis D. Breaux et al.· 0 citations
This work reformulates the implicit reward of sampled-token OPD based on trajectory correctness, then applies a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, making it readily combinable with any policy gradient algorithm, such as GRPO.
Wenze Lin, Jiale Zhao, Xi-Tai Jiang et al.· 0 citations