Beyond Truncation: Rethinking LLM Decoding as Ensemble Pruning
This work introduces Mahalanobis-Ensemble Decoding (ME-Decoding), a novel Large Language Model (LLM) decoding framework that frames candidate token selection as ensemble pruning and devise an efficient greedy selection algorithm with near-linear complexity in the candidate size under early stopping, while establishing...