Benjamin Marie writes about a comprehensive benchmark of 15 GGUF quantizations of Qwen3.8 27B, ranging from Q4_K_XL down to IQ1_M, evaluated using over 150 million tokens generated across roughly 8 days on an NVIDIA RTX Pro 6000. Using 950 prompts subsampled from MMLU-Pro, LiveCodeBench, and GPQA Diamond, he measures both accuracy and token efficiency to identify the lowest quantization level that retains at least 95% of the original BF16 model's performance.
| Chart pt | Quantization | Provider | GGUF file | Size (GB) | Accuracy recovery vs BF16 | Tokens generated | ≥ 95% threshold? |
|---|---|---|---|---|---|---|---|
| 1 | IQ3_XXS | bartowski | Qwen3.8-27B-IQ3_XXS.gguf | 12.39 | 97.7% | 11.54M | Yes |
| 2 | IQ4_XS | bartowski | Qwen3.8-27B-IQ4_XS.gguf | 15.33 | 99.1% | 10.18M | Yes |
| 3 | IQ2_S (AD) | AtomicChat | Qwen3.8-27B-AD-IQ2_S.gguf | 10.85 | 95.9% | 12.24M | Yes |
| 4 | IQ3_S (AD) | AtomicChat | Qwen3.8-27B-AD-IQ3_S.gguf | 13.60 | 101.1% | 10.48M | Yes* |
| 5 | Q4_K_M (AD) | AtomicChat | Qwen3.8-27B-AD-Q4_K_M.gguf | 16.84 | 99.9% | 9.99M | Yes |
| 6 | IQ2_S (GSQ-RCO) | ISTA-DASLab | Qwen3.8-27B-GSQ-RCO-IQ2_S.gguf | 9.26 | 92.2% | 11.81M | **No** |
| 7 | IQ3_XXS (GSQ-RCO) | ISTA-DASLab | Qwen3.8-27B-GSQ-RCO-IQ3_XXS.gguf | 10.09 | 96.5% | 10.87M | Yes |
| 8 | Ridge 3.7 bpw | empero-ai | Qwen3.8-27B-Ridge-3.7bpw.gguf | 12.26 | 97.4% | 10.29M | Yes |
| 9 | IQ1_M (UD) | unsloth | Qwen3.8-27B-UD-IQ1_M.gguf | 6.73 | 53.4% | 15.41M | **No** |
| 10 | IQ2_XXS (UD) | unsloth | Qwen3.8-27B-UD-IQ2_XXS.gguf | 7.27 | 74.3% | 11.96M | **No** |
| 11 | IQ3_XXS (UD) | unsloth | Qwen3.8-27B-UD-IQ3_XXS.gguf | 10.58 | 95.5% | 12.27M | Yes |
| 12 | Q2_K_XL (UD) | unsloth | Qwen3.8-27B-UD-Q2_K_XL.gguf | 9.48 | 96.0% | 11.89M | Yes |
| 13 | Q3_K_XL (UD) | unsloth | Qwen3.8-27B-UD-Q3_K_XL.gguf | 12.80 | 100.0% | 9.89M | Yes |
| 14 | Q4_K_XL (UD) | unsloth | Qwen3.8-27B-UD-Q4_K_XL.gguf | 17.21 | 101.0% | 9.47M | Yes* |
| 15 | Q4_K_XL abliterated (Huihui UD) | huihui-ai | Huihui-Qwen3.8-27B-abliterated-UD-Q4_K_XL.gguf | 17.03 | 98.8% | 9.93M | Yes |
`* = above 100% BF16 (sampling variance, not a real gain); below-threshold points are 6, 9, and 10.`