TIDE-Bench is introduced, a benchmark for conversational text-to-SQL under chain ambiguity and intent drift evaluation, targeting two recurring patterns: chain ambiguity, where an underspecified question triggers layered clarification with conditional dependencies, and intent drift, where the user retracts and replaces a previously committed request element.
Yu-Jia Liu, Jia-Yan Lin, Zijin Hong et al.· 0 citations
It is shown that the vector representations of a variety of neural networks can be closely approximated with symbolic structures, providing a potential way to reconcile longstanding symbolic conceptions of intelligence with the vector-based nature of modern AI.
R. Thomas McCoy, Paul Soulos, Tal Linzen et al.· 1 citation
Preliminary evidence is given that argument structure is a useful intermediate representation for aligning specialized normative texts in cross-standard control mapping and a neuro-symbolic pipeline is built that combines neural text representations with Toulmin features.
In the complete markdown sweep, hard negatives lower accuracy more than length-matched random documents at both context sizes for gpt-5-mini, while random documents stay near the no-distractor baseline, and for gpt-5-mini hard negatives significantly underperform length-matched random distractors when pooled across context sizes.
Jason Luo, Saibilila Abudukelimu, Judy Song et al.· 0 citations
Reach audiences
Advertise in front of researchers, engineers, and readers.
The results suggest that CoT monitoring may be less reliable when preference information arrives through tools or must be inferred from raw artifacts, within the single-call, prefilled-tool setting tested here.
Findings suggest that AI can be easily persuaded by what people say, who says it, and how the opinion is presented, enabling its safe and reliable use in high stakes medical decision making.
Jiayuan Zhu, Jiazhen Pan, Feng-Lin Liu et al.· 0 citations
The results support model-specific operating-point selection: set a deployment target and retain only interventions that improve it, and retain only interventions that improve it.
This work introduces StageWell, a process-aligned Chinese corpus for positive psychology dialogue together with HQS, a structured protocol for data construction and evaluation, and highlights the value of modeling supportive dialogue as a structured multi-turn support process rather than as single-turn response generation.
Yuxun Wang, Zihan Lin, Bo Wang et al.· 0 citations
This work splits each flagged answer into individual factual claims, checks each against the retrieved source, and compares leaving the answer untouched with three repair strategies of increasing richness: deleting an unsupported claim, replacing it with source text, and rewriting it.
Sai Krishna Reddy Mulakkayala, Niki van Stein, A. Plaat· 0 citations
SA-Pass (*Semantic Alignment Pass*), which tests formal statements using auxiliary statements called shadows that characterize the intended statement, is proposed, which achieves binary agreement with expert judgments.
Hojae Han, Jongyoon Kim, Sanghyuk Park et al.· 0 citations
A2R, a 30B model optimized on Counterfactual Audio with Speaker-level Hard negatives (CASH), a dataset designed to guide the model to prioritize acoustic vocal cues over linguistic signals, achieves strong performance on HEAR and exhibits zero-shot generalization to diverse multi-speaker downstream tasks, demonstrating that learned speaker attribution unlocks the model's latent capacity for speaker-aware reasoning.
Dongwook Lee, Sangkwon Park, Eunwoo Song et al.· 0 citations
A novel method that dynamically leverages target-related statements for conversational stance detection by employing a stepwise, entropy-guided backtracking mechanism to selectively activate memory from historical conversations and dynamically constructs a target-aware graph to model the stance relations among utterances is proposed.
Yifan Xiang, Bin Liang, Yu-Qi Huang et al.· 0 citations