Skip to content

Author

H. Shimodaira

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Oct 2026

My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

Agentic reinforcement learning (RL) has emerged as a powerful approach for training large language model agents on multi-step tasks, yet reliance on terminal outcome rewards creates two credit-assignment problems, particularly in long-horizon tasks. First, same-outcome rollout groups provide no learning signal from ter...

Yi-Hua Zhu, Qian-Ying Liu, Wei Qiao et al. · 0 citations

Language Model Maps for Prompt-Response Distributions via Log-Likelihood Vectors

A method that represents language models by log-likelihood vectors over prompt-response pairs and constructs model maps for comparing their conditional distributions is proposed, which supports the analysis and prediction of input-dependent model behavior.

Yusuke Takase, Momose Oyama, H. Shimodaira · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.