Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
Covered by 2 sources
Read full postResearchers introduced Quantization-Aware Healing (QAH), a method that improves compressed, 4-bit large language models. Applied to a GPT-OSS 120B model compressed to 60B parameters, QAH produced a smaller, cheaper, and more accurate model than its full-precision original.




