Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating in...
Fan Zhang, Yan-Kai Chen, Zhuo-Han Xie et al.· 0 citations
Automated prediction markets require sponsors to prefund liquidity before observing order flow, creating a financing challenge at launch. We study whether nonnegative charges conditioned on observable payoff direction can improve recovery of this prefunded capital while limiting their effect on informed participation....
Yan-Kai Chen, Bowei He, Zhuo-Han Xie et al.· 1 citation
This work introduces FinCUABuildBench, a benchmark for evaluating financial CUA task construction, and introduces FinCUABuildAgent, a multi-agent system for automatically constructing dynamic financial CUA evaluation tasks.
Jing-Pu Yang, Feng-Xian Ji, Jinri Guo et al.· 0 citations
Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report. The family includes, to our knowledge, the largest...
Mohamed Anwar, A. Freihat, George Ibrahim et al.· 3 citations
When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-...
R. Elbadry, Ahmed Heakl, Saeed Almheiri et al.· 0 citations
Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.
Chang-Jiang Jiang, Qiannian Zhao, Lei Xin et al.· 2 citations
Professional financial examinations require models to combine domain knowledge, calculation, and judgment, yet no benchmark covers the full CFA and FRM structure under one protocol. We introduce FinExam-10K, to our knowledge the largest reported English benchmark for this setting, with 10,198 expert-reannotated questio...
Yan Lin, Jingyu Sun, Zhong-Liang Guo et al.· 0 citations
FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages...
Zhuohan Xie, Yu-Yang Dai, R. Elbadry et al.· arXiv.org· 1 citation
Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the final diagnosis is incorrect. We introduce...
Abin Roy, Afthab Salam Kanniyan, Jawadh Abdul Kabeer et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.