Skip to content

Author

S. Avestimehr

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Aligned Data Can Induce Misalignment via Context Confusion

It is argued that it is difficult to predict the alignment state of a model after training by inspecting the training data alone, which highlights the importance of comprehensive post-training alignment evaluations.

Yavuz Faruk Bakman, D. Yaldiz, Baris Askin et al. · 0 citations
Preprint Aug 2026

Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms

Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is bandwidth-bound. Hence, reducing the cache size can raise both decoding speed and serving capacity. The challenge is to reduce cache size while preserving the attention prod...

Samuel Fernández-Menduiña, Amir Ziashahabi, Eduardo Pavez et al. · 0 citations
#machine learning Preprint Sep 2026

ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

Long-context LLM inference is bottlenecked by attention, whose repeated KV-cache reads make decoding memory-bound. Self-speculative decoding alleviates this by drafting tokens with sparse attention and verifying them with full attention, but existing batched methods remain synchronized: all requests in a batch share a...

Amir Ziashahabi, Hossein Entezari Zarch, Lei Gao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.