Skip to content

Author

Yufei Cui

We have 5 of 23 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

DimPO: Dimensionality Reduction for Attention using Preference Optimization

A linear projection can reduce the dimension of query and key vectors without updating the pretrained model, but it remains unclear which training objective best preserves model behavior. We ask whether preferences over keys and attention mass on the highest-weighted keys provide a better signal than matching the full...

Vojtěch Lanz, Yu-Fei Cui, Prasanna Parthasarathi · 0 citations
#artificial intelligence Preprint Sep 2026

ARM: Attention with Routed-Memory for Learnable Sparse Control

Despite advances in long-context inference, large language models (LLMs) remain fundamentally limited by the key-value (KV) caching mechanisms that are necessary for stable computation. Techniques such as selective token eviction and pruning have vastly mitigated these issues, but often discard core information to mana...

Qiu-Hao Zeng, Jerry M. Huang, Peng Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

FINSKILLOPS: A Self-Evolving Multi-Agent System for SEC Filing QA

Financial QA systems are typically improved before deployment through better retrieval, prompting, or agent coordination, leaving their reliability behavior fixed thereafter. In practice, new SEC-filing questions repeatedly expose heterogeneous errors in period, entity, evidence use, and calculation. Existing self-impr...

Yanzhang Ma, Zhenghan Tai, Han-Wei Wu et al. · 0 citations
Jul 2026

FinSAgent: Corpus-Aligned Multi-Agent RAG Framework for Evidence-Grounded SEC Filing Question Answering

Financial question answering over U.S. Securities and Exchange Commission (SEC) filings requires retrieving and synthesizing heterogeneous evidence dispersed across long, standardized, and highly redundant disclosures. Existing retrieval-augmented and multi-agent systems typically derive retrieval queries directly from...

Jijun Chi, Zhenghan Tai, Hanwei Wu et al. · 0 citations
Jul 2026

Stable FP4 Training via Transposition-Invariant Block Quantization

This work proposes a low-precision training framework based on 2D block FP4 quantization, which enforces transposition-invariant scaling and preserves consistency between forward and backward computations, and combines this with truncation-free scaling and stochastic rounding to control quantization error and maintain...

Mehdi Rahimifar, Amin Darabi, Mehran Taghian Jazi et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.