Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original [LAB]
Directly relevant to the reader's interest in quantization and the performance of compressed models for local inference. https://huggingface.co/blog/MultiverseComputingCAI/quantization-aware-healing (huggingface.co)