Skip to content

Author

C. Sripada

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning

Native visual reasoning treats visual generation as the medium of reasoning itself: visual states (i.e. images and videos) are not merely inputs to be understood or outputs to be rendered, but first-class substrates for problem solving beyond language. Yet progress remains bottlenecked by the lack of scalable training tasks, reliable feedback, and controlled comparisons across generative substrates. In this work, we introduce VBVR-Pro, a closed-loop testbed that makes native visual reasoning through generation trainable, verifiable, optimizable, and experimentally controllable. 1) Task scaling. VBVR-Pro turns visual reasoning into a controlled task space of 300 procedurally generated tasks. Models trained on VBVR-Pro show strong transfer beyond the proposed suite across seven external visual reasoning benchmarks such as RISE-Video, MME-CoF-Pro, and BabyVision. 2) Verifiable rewards. VBVR-Pro provides verifiable reward scorers for task-grounded evaluation. Through a systematic study of leading MLLMs as judges, we identify recurring failure modes of the prevalent VLM-as-a-judge paradigm. In contrast, the proposed scorers are grounded in deterministic, task-specific rules, achieve fine-grained alignment with human judgments. Importantly, they serve as reliable reward signals for large-scale multi-task reinforcement learning and demonstrate stronger post-RL performance across visual reasoning tasks. 3) Mechanism study. VBVR-Pro enables controlled modality studies across more than 30 image, video, and interleaved generators. Our analysis shows that video generation remains strongest for tasks requiring persistent spatiotemporal state tracking, while interleaved generation provides a compute-efficient alternative. Critically, ablations and probing suggest the presence of vision-native trajectories that are crucial to visual reasoning. We release all data, models, scorers, and code.

Junhua Xu, Ruisi Wang, Fanyi Pu et al. · 0 citations
Preprint Jul 2026

Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition

LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human cognition are therefore often seen as the result of anthropomorphic projection. We argue that this framing is mistaken. LLMs clearly differ from humans in important respects, including their physical substrate, learning history, and the environments with which they interact. These differences make it all the more striking that contemporary LLM-based systems converge with human cognition on a number of principles of cognitive organization with longstanding support in cognitive science. We identify structural correspondences across five dimensions: inferential organization, computational architecture, representational structure, prediction-driven learning, and reinforcement-learning-like mechanisms supporting goal-directed action. These correspondences support a broader model of intelligent cognition in which core principles long used to explain human intelligence also characterize contemporary LLM-based systems.

C. Sripada, Richard L. Lewis · 0 citations
Review Open access Jul 2026

Scalable, context-sensitive psychiatric assessment with large language models and brief diaries

Abstract Background Accurate psychiatric assessment requires understanding a person’s unique experience within their psychosocial context. Clinical interviews have been the gold standard for assessment as the only methods capable of this complex task, but they are time and resource-intensive. Consequently, psychiatric assessment typically relies on patient report surveys that are decontextualized and narrow in scope. This comprehensiveness-scalability tradeoff is a major bottleneck in studying and treating psychopathology. We propose using large language models (LLMs) to score psychopathology from brief personal narratives as a low-burden, context-sensitive solution. Methods Participants (N = 108) completed brief (~1 minute), freeform audio diaries daily for 2 weeks. We used six LLMs to score wide-ranging psychopathology (Internalizing, Detachment, Disinhibition, Antagonism, Anankastia) from the diary transcripts. Leveraging an array of self-report and clinical interview measures, we tested the convergent, discriminant, concurrent, and clinical validity of LLM ratings for between-person differences and within-person fluctuations in psychopathology. Results Supporting convergent and discriminant validity, LLM ratings correlated most strongly with corresponding self-report domains at the between (average convergent r = .42) and within-person (r = .28) levels. LLM and self-report ratings had similar patterns of associations with external variables, except for Anankastia and Antagonism. Further, every LLM-rated domain related to psychopathology ascertained by clinical interview. Conclusions Across multiple forms of validity, we showed that LLMs can assess most major forms of psychopathology from mere minutes of audio. These results support scoring open-ended narratives with LLMs as a scalable, portable method to translate idiographic diagnostic data into standardized psychiatric assessments.

Whitney R. Ringwald, Aman Taxali, Michael Angstadt et al. · 0 citations