This work developed a methodological workflow using Item Response Theory to evaluate VLM and human rater proficiency against expert-established ground truth, suggesting that top-performing VLMs can approximate ground-truth ratings at levels comparable to human raters.
Lana Do, Gio Jung, J. F. Barajas et al.· 0 citations
It is shown that passive EEG, fused online with behavioral evidence, can meaningfully extend the number of targets users detect and engage beyond their unaided action bandwidth, and that OLIVE Pareto-dominates prior test-time adaptation frameworks, achieving the highest convergence rate at comparable convergence speed.
Co-Annotator is presented, which distills expert gaze and dictation into two guidance components: a gaze-aligned Vision Transformer producing fixation-aligned areas of interest (AOIs), and an ontology-bounded vision-language model (VLM) that pre-fills editable biomarker summaries for retinal optical coherence tomography (OCT).
ZihengLeoLi, Benjamin Freeman, Akshay Raman et al.· 0 citations
GuardianAgent, a policy-conditioned anonymization framework that couples structured risk assessment with verified adaptive rewriting, achieves the strongest privacy-utility trade-off among published baselines and is the only method to reach more than 0.90 privacy in all three domains, remaining robust under a backbone switch.
Ruiyi Yang, Gayathri Lihinikaduarachchi, Rahat Masood et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
It is suggested that cognitive processes can provide a promising basis for support selection in proactive writing systems and is hypothesized that Flower and Hayes' cognitive process theory of writing offers an interpretable bridge between observable writing behavior and appropriate support types.
Masahiro Yoshida, Atsuya Kobayashi, Kei Tateno et al.· 0 citations
Simulations show that over-reliance on a weak AI is especially harmful, and that diversifying AI signals across users can better keep the crowd informative, and conclude with implications for understanding human-AI interaction in information spread and designing misinformation interventions.
Zhuoran Lu, Weilong Wang, Yang-Yang Yu et al.· 0 citations
It is argued that AI-SEL studies should systematically specify Who should act, What actions are recommended, Why these actions are needed, When and Where they apply, and How strongly they are framed, thereby strengthening the translation of AI x SEL innovation into educational policy and practice.
C. V. Tran, Yi-Hang Liu, Tuong Van Nguyen· 0 citations
A provenance-aware pipeline that converts task telemetry and human-authored reports into a shared typed task-state representation, aligns and reconciles their facts, detects conflicts, and generates structured handover reports supports provenance-aware state reconciliation as a design pattern for safer AI-assisted handover.
Kayleigh Bishop, Maria P. Stull, Breanne Crockett et al.· 0 citations
This work model the AI-mediated process of writing a student--instructor email at the highest level of involvement, and compares it with an unaided model of writing the same messages, built from participants' accounts and a classic model of the writing process.
The results suggest a novel way of approaching automated evaluation, by offering a faster, more explainable, and less ambiguous alternative to black-box rubric evals, particularly in high-stakes domains such as healthcare and banking where precision and auditability are critical.
Kaustubh D. Dhole, Charles L. A. Clarke, E. Agichtein· 0 citations
Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible. We present PsychoAgent, a cognitive architecture for LLM agents that separates factual and affective memory and integrates both through a conflict-aware executive controller. Affective memories are first filtered by semantic relevance and then re-ranked by salience, preserving topical fit while allowing emotionally important traces to enter the prompt. Across three controlled conflict scenarios, the full architecture retrieved more conflict-critical memories than semantic-affective and single-memory RAG baselines (0.933 vs. 0.500 and 0.667), with a small semantic-similarity cost. Five blinded raters evaluated 27 outputs. After within-rater standardization, the full architecture had the highest overall mean (+0.22 SD), but corrected pairwise differences were not significant. A three-day illustrative trace further shows persistent affect, offline memory recombination, and selective memory reweighting. The findings support affect-sensitive retrieval as an inspectable mechanism for modeling human-like conflict effects in LLM agents.
Mohammad Amanlou, Parham Abed Azad, Farbod Davoodi et al.· 0 citations
A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form. We ask whether that reading is faithful: when an item's intent is held fixed and only its meaning-preserving surface form varies, does the canonical-form score estimate model behavior well, and how much of any variation is decoding/judge noise rather than signal? We instantiate this in safety, a high-stakes setting with no gold label to average toward. To avoid prior confounds, we pre-author the reformulations (refusal-free, mostly non-LLM: machine back-translation and a Matrix-Language-Frame code-switch generator) so an identical surface form reaches every model, score all responses with one human-anchored, vendor-neutral judge (Claude, kappa = 0.86 vs. human on unsafe compliance, stable across languages, cross-checked by GPT-4o), and verify intent preservation. On 370 seeds x 5 surface forms x 5 models, no single transformation is uniformly most dangerous (6 of 20 per-transformation McNemar tests survive correction, most protective). Yet evaluating only the canonical prompt underestimates unsafe compliance: the union of unsafe outcomes across forms exceeds even the worst single form by 3.3-12.9 pp, with bootstrap 95% CIs excluding zero for all five models, and 5-13% of seeds safe on canonical are unsafe under some reformulation -- above a zero stochasticity floor (canonical resampled five times at temperature 0 gives 0/370 new exposures). The size of this gap is model-dependent (largest on Gemini 2.5 Pro). One form recovers only ~53% of a model's observed unsafe surface and about three reach 85% -- a redundancy characterization of this form set, not of a defined population. A benign control (XSTest) suggests the instability is bidirectional, though the benign and harmful pools are not item-matched. We release the dataset, code, and per-response labels.
Yongxin Zhou, Jun-Wei Yao, Yuanzhe Liu et al.· 1 citation