Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
Multiverse Computing published Quantization-Aware Healing (QAH): A Practical Recipe for Recovering Compressed, 4-Bit LLMs, introducing a method that distills directly from the original pre-compression model instead of the recovered bfloat16 checkpoint.
Applying QAH to a GPT-OSS 120B model compressed to 60B parameters and quantized to MXFP4 produced a model that beat its own bfloat16 source on 7 of 9 benchmarks. Gains included +7.4 on AA-LCR long-context reasoning, +5.6 on AIME 2025 math, +2.7 on the Aider agentic coding benchmark, +2.3 on τ²-bench tool use, +1.7 on GPQA Diamond, +1.5 on IFBench, and +1.0 on LiveCodeBench. It trailed on MMLU-Pro by 0.2 and SciCode by 1.4. Compared with the 120B teacher, the QAH model surpassed it on LiveCodeBench (66.5 vs. 66.0) and came within 1.6 points on GPQA Diamond.
In a separate head-to-head against quantization-aware training (QAT) on a GPT-OSS 9B quantized to MXFP4, both methods peaked similarly: 54.9 for QAH versus 54.6 for QAT. QAH reached its peak in about 100 steps against QAT's roughly 700, and stayed within about two points of that peak, while QAT lost nearly 19 points by step 1,200.
In comments, a reader argued the headline comparison lacked a control: the bf16 60B received one distillation pass while the MXFP4 60B received that pass plus a second against the 120B teacher, and said the table could not separate the recipe from training longer. The reader also noted the table labels the teacher "120B teacher (MXFP4)" while the approach section describes it as full-precision. An article author acknowledged the unequal training, said the caveat would be added, and described the work as an applied case study whose supported claim is that the shipped 4-bit checkpoint scores above the bf16 checkpoint it was quantized from. The author said retraction was not warranted and that a revised version would add the suggested comparison arms.
Based on reporting from the original publisher. Visit the source for full context and later updates.