A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes
Hugging Face Blog
· huggingface.co · August 17, 2022
Read on Hugging Face Blog →
Opens the original article in a new tab.
More from the blog
Microsoft Research Blog
· microsoft.com
Aug 31, 2026
GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models
What if pathology foundation models could do more with less? GigaPath-Flash and GigaTIME-Flash cut computational demands while maintaining strong performance, opening the door to larger studies and broader exploration. The post GigaPath-Flash and GigaTIME-Flash: Toward population-scale discovery with efficient pathology foundation models appeared first on Microsoft Research.
Hugging Face Blog
· huggingface.co
Aug 26, 2026
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
Hugging Face Blog
· huggingface.co
Aug 25, 2026
Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
A Blog post by Multiverse Computing on Hugging Face
Hugging Face Blog
· huggingface.co
Aug 18, 2026
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
We’re on a journey to advance and democratize artificial intelligence through open source and open science.