JustQuant: You Don't Need Smoothing, SVD, or Rotation for 4-Bit Activation Quantization
Recent generative models have become increasingly powerful, but their inference cost continues to grow. Model quantization offers a promising way to compress these models and accelerate inference. However, at 4 bits, activation quantization is substantially more challenging than weight quantization. Recent post-trainin...