Skip to content

Author

Zhuohan Xie

We have 10 of 56 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

Contract inference requires multiple judgments about a shared document, but aggregate accuracy can conceal changes in the individual decisions. Repeated agreement is also insufficient: a model may consistently return the wrong answer. In this paper, we compare Jev with nine language models on ContractNLI, evaluating in...

Fan Zhang, Yan-Kai Chen, Zhuo-Han Xie et al. · 0 citations
Preprint Sep 2026

Prediction-Market Seed Capital Recovery from Noise-Dominant Flow

Automated prediction markets require sponsors to prefund liquidity before observing order flow, creating a financing challenge at launch. We study whether nonnegative charges conditioned on observable payoff direction can improve recovery of this prefunded capital while limiting their effect on informed participation....

Yankai Chen, Bowei He, Zhuohan Xie et al. · 1 citation
Preprint Jul 2026

Jais 2: A Family of Arabic-Centric Open Large Language Models

Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report. The family includes, to our knowledge, the largest...

Mohamed Anwar, A. Freihat, George Ibrahim et al. · 3 citations
#artificial intelligence Preprint Sep 2026

Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close

When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-...

R. Elbadry, Ahmed Heakl, Saeed Almheiri et al. · 0 citations
Preprint Aug 2026

Think with Structured Grounding: Perceptual Reinforcement Learning for Chart and Visual-Tabular Understanding

Think with Structured Grounding (TwSG), a novel fine-grained image perception framework designed to internalize complex images's tool-use capabilities within the model, is proposed, endowing models with native fine-grained region description and flexible reasoning capabilities.

Chang-Jiang Jiang, Qiannian Zhao, Lei Xin et al. · 2 citations
#natural language process... Preprint Aug 2026

FinExam-10K: When Retrieval Helps Financial Reasoning?

Professional financial examinations require models to combine domain knowledge, calculation, and judgment, yet no benchmark covers the full CFA and FRM structure under one protocol. We introduce FinExam-10K, to our knowledge the largest reported English benchmark for this setting, with 10,198 expert-reannotated questio...

Yan Lin, Jingyu Sun, Zhong-Liang Guo et al. · 0 citations
Review Jul 2026

Overview of FinMMEval 2026 Task 2: Multilingual Financial Short-Answer Question Answering

FinMMEval 2026 Task 2 evaluates short-answer financial question answering over multilingual evidence over multilingual evidence using document retrieval-augmented generation, cross-lingual evidence handling, structured prompting, answer compression, and validation strategies.

Zhuohan Xie, Xueqing Peng, Georgi N. Georgiev et al. · 1 citation
Review Jul 2026

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi. The task tests whether systems can select the correct answer to finance questions involving domain terminology, numerical interpretation, and conceptual financial reasoning across languages...

Zhuohan Xie, Yu-Yang Dai, R. Elbadry et al. · 1 citation
Jun 2026

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the final diagnosis is incorrect. We introduce...

Abin Roy, Afthab Salam Kanniyan, Jawadh Abdul Kabeer et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.