Skip to content

Author

Matthias Seeger

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Learning how to Forget: Fine-tuning for Long-Context Sparse Attention

This work provides a new method for fine-tuning models with sparse attention that works for any KV cache policy, runs on a moderate hardware budget, and allows the model to co-adapt with the policy, often outperforming models trained with exact attention (sequence parallelism).

Matthias W. Seeger, Zeyu Zhang, Vihang Patil et al. · 0 citations