Benchmarking Vector Quantized Auto-Encoders for Multimodal Vision-Language Tokenization of Medical Images
This work benchmarks both reconstruction quality, codebook collapse and representational capacity of VQ-VAEs across a variety of settings, surpassing state of the art in the reconstruction task and providing a stepping stone for further development of medical multimodal auto-regressive techniques.