Measuring how much quantization affects model quality
I benchmark different variants of Qwen3.8-27B to measure the effect of quantization
It is really fun to run models locally, testing different variants and experimenting with different configurations.
For a given model you generally have, besides the main BF16 checkpoint, several quantized checkpoints whose aim is to reduce its size and increase performance.
When running on large servers with 8xB300 GPUs and 2.1TB of HBM memory, I usually do not spend much time selecting the quantization. In these cases what I mainly look at is the tuning for throughput or latency, using tensor parallelism. When the BF16 checkpoint does not fit in memory, I just choose the FP8 or NVFP4 variants, because they have hardware support in Blackwell and so they perform really well.
But there are cases where I want to run on a more modest server, or even on a 4090 GPU. This is when having a quantized checkpoint helps a lot, because it can fit in the VRAM that you have.
It is also when things start to get a little messy, because there are a lot of ways to quantize a model and, for the popular ones, we have many readily available checkpoints at our disposal.
With so many options, sometimes it is not very clear how much a given quantization affects the quality of the model. In some cases the model card includes information about top-1% accuracy; in others there are no clues about how much the quantization degraded the checkpoint. We have all seen cases where aggressive quantizations drop the quality of the model dramatically.
So the best way to know for sure is to run some benchmarks.
In this post I will show the results for Qwen3.8-27B, one of my favourite models right now at this size. I use it both for running inference and for creating fine-tuned versions.
As you will see, quantization does a pretty good job, but the model card alone is not enough to know which checkpoint is better.
If you want, you can also go straight to the results.
Configuration#
Qwen3.8-27B served with vLLM on an RTX PRO 6000 (the same options in all cases: greedy decoding and reasoning mode enabled).
These are the checkpoints that were evaluated:
| variant | checkpoint | weights | activations | weights on GPU |
|---|---|---|---|---|
| BF16 | Qwen/Qwen3.8-27B ↗ | 16-bit | 16-bit | 51.1 GiB |
| FP8 | Qwen/Qwen3.8-27B-FP8 ↗ | 8-bit float | 8-bit float | 28.5 |
| NVFP4 | Inferact/Qwen3.8-27B-NVFP4 ↗ | 4-bit float, gs 16 | 4-bit float | 24.2 |
| NVFP4 | unsloth/Qwen3.8-27B-NVFP4 ↗ | 4-bit on MLP, 8-bit elsewhere | 8-bit | 21.3 |
| NVFP4 | RadixArk/Qwen3.8-27B-NVFP4 ↗ | mixed 4/8-bit | mixed | 20.0 |
| INT4 | cyankiwi/Qwen3.8-27B-AWQ-INT4 ↗ | 4-bit int, gs 32 | 16-bit | 19.2 |
| INT4 | RedHatAI/Qwen3.8-27B-INT4 ↗ | 4-bit int, gs 128 | 16-bit | 17.7 |
And this is the vLLM command used to run all the checkpoints:
vllm serve "${MODEL}" \
--served-model-name "${SERVED_NAME}" \
--tensor-parallel-size 1 \
--max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.95 \
--max-num-seqs 256 \
--max-model-len 262144 \
--kv-cache-dtype fp8 \
--no-enable-prefix-caching \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--mm-encoder-tp-mode databashFor the one-shot benchmarks I use lm-eval ↗.
Benchmarks#
-
Agentic performance:
- SWE-bench Verified: the model drives a terminal to fix real GitHub issues in real repositories until the project’s own tests pass.
-
General accuracy in one-shot tasks:
- MMLU-Pro: multiple-choice general knowledge and reasoning across 14 academic and professional domains.
- BBH (BIG-Bench Hard): mostly multi-step symbolic, logical and algorithmic reasoning.
- Minerva-MATH: competition mathematics problems requiring worked derivations, graded on the final boxed answer.
- MBPP+: short Python programming problems.
- IFEval: instruction following under verifiable constraints (“answer in exactly three bullets”, “no commas”).
- GPQA Diamond: 198 graduate-level biology, chemistry and physics questions written to be hard even for a non-expert with web access.
Results#
Agentic performance#
Four of the five quantizations are statistically indistinguishable from BF16 on the agentic coding benchmark; only NVFP4-Inferact degrades considerably.
Understanding why NVFP4-Inferact performs much worse than the other NVFP4 checkpoints took some research. The model cards do not make the difference clear, and the reason only shows up in the model repo: Inferact quantizes activations alongside the weights.
quantization_config in the model repo.
Quantization does a pretty good job. The only thing to avoid is quantizing the activations to 4 bits, because that degrades model performance in agentic coding tasks. Unfortunately, you have to manually open
config_groupsin the checkpoint config and read it to see which checkpoints do that.
General accuracy in one-shot tasks#
Two BF16 runs are plotted on every chart, so the difference between them represents the run-to-run noise in that benchmark.
Each bar in the figures is split into correct answers, incorrect answers and no-answer. I prefer to separate failed answers into wrong and no-answer, because answering wrongly is not the same as consuming all the output tokens before answering.
To have a better view of how the checkpoints compare, apart from the raw benchmark score, I also compare each one against the base BF16 variant using a paired McNemar test. The test ignores the cases where the two models agree and focuses on the disagreements, i.e. the questions that they answer differently.
A bar is marked in red when a paired McNemar test against BF16 returned
p < 0.05 on that benchmark, which means that there is an important difference from the BF16 checkpoint; a ▲ marks the one case where a variant is
significantly better.
INT4-RedHat has the lowest raw score on this family of benchmarks, last on MMLU-Pro and last on GPQA, but on the paired test it is the only 4-bit checkpoint not significantly worse than BF16. This is why having both the raw scores and the McNemar test results is useful.
Looking at the details, INT4-RedHat’s low score comes from its high percentage of no-answer results. This happens when the model reasons past its token budget and so returns no answer at all. On the questions it does answer, it is the least degraded 4-bit checkpoint.
Looking at some of the specific questions that it does not answer, the model is not looping. The traces show reasoning that is making progress, cut off mid-stream by the budget. Increasing the budget improves the results. For some reason this checkpoint needs more reasoning than the others.
So when ranking on the raw benchmark numbers INT4-RedHat is the worst checkpoint, but when ranking on the paired test it is the best 4-bit checkpoint in the set.
Surprisingly, on SWE-bench the same checkpoint had zero non-termination and performed on par with BF16.
Memory usage#
This is the memory usage for each checkpoint, as reported by vLLM on one GPU:
The GPU memory budget is fixed, and I like to use a higher utilization setting (--gpu-memory-utilization 0.95), so whatever the weights do not use is assigned to the KV cache:
| variant | on disk | weights (+overhead) | peak activations | CUDA graphs | KV cache | KV tokens | vs BF16 |
|---|---|---|---|---|---|---|---|
| BF16 | 51.8 | 51.9 | 3.27 | 1.82 | 35.0 | 1,120,627 | 1.00× |
| FP8 | 28.8 | 29.4 | 3.34 | 1.86 | 57.5 | 1,842,673 | 1.64× |
| NVFP4-Inferact | 24.6 | 25.1 | 3.27 | 1.87 | 61.8 | 1,979,110 | 1.77× |
| NVFP4-unsloth | 21.8 | 22.3 | 3.27 | 1.86 | 64.6 | 2,069,557 | 1.85× |
| NVFP4-RadixArk | — | 21.1 | 3.34 | 1.86 | 65.8 | 2,107,883 | 1.88× |
| AWQ-INT4 | 19.6 | 20.1 | 3.34 | 1.86 | 66.8 | 2,138,543 | 1.91× |
| INT4-RedHat | 18.1 | 18.3 | 3.34 | 1.86 | 68.6 | 2,196,797 | 1.96× |
Conclusions#
Quantization performs really well in most cases, close enough to the original model that you cannot tell them apart, but you have to be careful with the checkpoint that you choose.
In the case of Qwen3.8-27B:
- FP8 is indistinguishable from BF16.
- 4-bit variants are also good, but you should choose carefully based on the use case.
- NVFP4-Inferact is the one that performs worst, due to 4-bit activations.
- INT4-RedHat performs very well being the smallest in size, but it has a very high no-answer rate in one-shot benchmarks.
This is the final summary:
| variant | weights | saved | measured cost |
|---|---|---|---|
| BF16 | 51.1 GiB | — | reference |
| FP8 | 28.5 GiB | 44% | none |
| NVFP4-Inferact | 24.2 GiB | 53% | −8.0 pts SWE-bench, significant on 2 of 6 benchmarks |
| NVFP4-unsloth | 21.3 GiB | 58% | significant on BBH only |
| NVFP4-RadixArk | 20.0 GiB | 61% | significant on BBH only |
| AWQ-INT4 | 19.2 GiB | 62% | significant on BBH and MMLU-Pro |
| INT4-RedHat | 17.7 GiB | 65% | small, but highest no-answer rate |