Skip to content

Author

Yuta Oshima

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation

Text-to-image (T2I) models can synthesize realistic, prompt-aligned images, yet samples generated for the same prompt often cover only a small subset of visually distinct modes. This limits the diversity of images, and for person-centric prompts, can reflect or amplify demographic skew. We formalize this problem as coverage of a predefined set of semantically specified modes, which we call target-mode coverage. We then propose multi-axis max@K, a group-based reinforcement learning objective for improving such coverage in diffusion-based T2I models. Given a group of samples and one score per target category, multi-axis max@K first takes the maximum score across samples for each category and then sums these category-wise maxima. The resulting credit assignment gives a sample positive weight on a category only when it increases that category's group-wise maximum, allowing different samples to contribute to different categories. We first validate the credit-assignment mechanism on a synthetic mixture and on SD3.5-M using deterministic pixel-based color rewards. We then evaluate the same objective on perceived-appearance fairness. Across three automatic evaluators on held-out prompts, multi-axis max@K improves the Fairness Score by 0.23-0.36 relative to the base model, while maintaining image quality and text alignment.

Ku Onoda, Paavo Parmas, Hiroki Furuta et al. · 0 citations
Open access Mar 2024

SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces

The ablation study shows that when using SSMs for temporal modeling, incorporating bidirectionality and selective scans enhances video generation performance, and SSM-based models incur lower computational cost to achieve the same Fréchet Video Distance as attention-based models.

Yuta Oshima, Shohei Taniguchi, Masahiro Suzuki et al. · 13 citations