Javier Cacheiro

Back

It is really fun to run models locally, testing different variants and experimenting with different configurations.

For a given model you generally have, besides the main BF16 checkpoint, several quantized checkpoints whose aim is to reduce its size and increase performance.

When running on large servers with 8xB300 GPUs and 2.1TB of HBM memory, I usually do not spend much time selecting the quantization. In these cases what I mainly look at is the tuning for throughput or latency, using tensor parallelism. When the BF16 checkpoint does not fit in memory, I just choose the FP8 or NVFP4 variants, because they have hardware support in Blackwell and so they perform really well.

But there are cases where I want to run on a more modest server, or even on a 4090 GPU. This is when having a quantized checkpoint helps a lot, because it can fit in the VRAM that you have.

It is also when things start to get a little messy, because there are a lot of ways to quantize a model and, for the popular ones, we have many readily available checkpoints at our disposal.

With so many options, sometimes it is not very clear how much a given quantization affects the quality of the model. In some cases the model card includes information about top-1% accuracy; in others there are no clues about how much the quantization degraded the checkpoint. We have all seen cases where aggressive quantizations drop the quality of the model dramatically.

So the best way to know for sure is to run some benchmarks.

In this post I will show the results for Qwen3.8-27B, one of my favourite models right now at this size. I use it both for running inference and for creating fine-tuned versions.

As you will see, quantization does a pretty good job, but the model card alone is not enough to know which checkpoint is better.

If you want, you can also go straight to the results.

Configuration#

Qwen3.8-27B served with vLLM on an RTX PRO 6000 (the same options in all cases: greedy decoding and reasoning mode enabled).

These are the checkpoints that were evaluated:

variantcheckpointweightsactivationsweights on GPU
BF16Qwen/Qwen3.8-27B ↗16-bit16-bit51.1 GiB
FP8Qwen/Qwen3.8-27B-FP8 ↗8-bit float8-bit float28.5
NVFP4Inferact/Qwen3.8-27B-NVFP4 ↗4-bit float, gs 164-bit float24.2
NVFP4unsloth/Qwen3.8-27B-NVFP4 ↗4-bit on MLP, 8-bit elsewhere8-bit21.3
NVFP4RadixArk/Qwen3.8-27B-NVFP4 ↗mixed 4/8-bitmixed20.0
INT4cyankiwi/Qwen3.8-27B-AWQ-INT4 ↗4-bit int, gs 3216-bit19.2
INT4RedHatAI/Qwen3.8-27B-INT4 ↗4-bit int, gs 12816-bit17.7

And this is the vLLM command used to run all the checkpoints:

vllm serve "${MODEL}" \
    --served-model-name "${SERVED_NAME}" \
    --tensor-parallel-size 1 \
    --max-num-batched-tokens 8192 \
    --gpu-memory-utilization 0.95 \
    --max-num-seqs 256 \
    --max-model-len 262144 \
    --kv-cache-dtype fp8 \
    --no-enable-prefix-caching \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --reasoning-parser qwen3 \
    --mm-encoder-tp-mode data
bash

For the one-shot benchmarks I use lm-eval ↗.

Benchmarks#

  • Agentic performance:

    • SWE-bench Verified: the model drives a terminal to fix real GitHub issues in real repositories until the project’s own tests pass.
  • General accuracy in one-shot tasks:

    • MMLU-Pro: multiple-choice general knowledge and reasoning across 14 academic and professional domains.
    • BBH (BIG-Bench Hard): mostly multi-step symbolic, logical and algorithmic reasoning.
    • Minerva-MATH: competition mathematics problems requiring worked derivations, graded on the final boxed answer.
    • MBPP+: short Python programming problems.
    • IFEval: instruction following under verifiable constraints (“answer in exactly three bullets”, “no commas”).
    • GPQA Diamond: 198 graduate-level biology, chemistry and physics questions written to be hard even for a non-expert with web access.

Results#

Agentic performance#

no measurable loss BF16 reference significant loss
60% 64% 68% 72% resolve rate — 500 SWE-bench Verified instances FP8 71.8% BF16 71.6% INT4-RedHat 71.6% NVFP4-unsloth 71.4% AWQ-INT4 71.0% NVFP4-Inferact 63.6% −8.0 points · p = 3.5×10⁻⁵
Paired McNemar test against BF16, counting only instances where both variants reached a verdict.

Four of the five quantizations are statistically indistinguishable from BF16 on the agentic coding benchmark; only NVFP4-Inferact degrades considerably.

Understanding why NVFP4-Inferact performs much worse than the other NVFP4 checkpoints took some research. The model cards do not make the difference clear, and the reason only shows up in the model repo: Inferact quantizes activations alongside the weights.

WEIGHTS ACTIVATIONS FP8 8-bit 8-bit AWQ-INT4 4-bit int 16-bit INT4-RedHat 4-bit int 16-bit NVFP4-unsloth 4-bit 8-bit 8-bit NVFP4-Inferact 4-bit 4-bit weights left block: MLP only; activations dashed: full precision
What each checkpoint actually quantizes, read from quantization_config in the model repo.

Quantization does a pretty good job. The only thing to avoid is quantizing the activations to 4 bits, because that degrades model performance in agentic coding tasks. Unfortunately, you have to manually open config_groups in the checkpoint config and read it to see which checkpoints do that.

General accuracy in one-shot tasks#

Two BF16 runs are plotted on every chart, so the difference between them represents the run-to-run noise in that benchmark.

Each bar in the figures is split into correct answers, incorrect answers and no-answer. I prefer to separate failed answers into wrong and no-answer, because answering wrongly is not the same as consuming all the output tokens before answering.

To have a better view of how the checkpoints compare, apart from the raw benchmark score, I also compare each one against the base BF16 variant using a paired McNemar test. The test ignores the cases where the two models agree and focuses on the disagreements, i.e. the questions that they answer differently.

A bar is marked in red when a paired McNemar test against BF16 returned p < 0.05 on that benchmark, which means that there is an important difference from the BF16 checkpoint; a ▲ marks the one case where a variant is significantly better.

0% 25% 50% 75% 100% MMLU-Pro BF16 (a) 83.1% BF16 (b) 82.8% FP8 82.8% NVFP4-unsloth 82.9% NVFP4-RadixArk 83.1% AWQ-INT4 81.9% p = 0.0111 INT4-RedHat 78.8% NVFP4-Inferact 80.9% p = 0.00349
INT4-RedHat has the lowest score, but this is mainly due to a large percentage of no-answer results. See below.
0% 25% 50% 75% 100% BBH BF16 (a) 90.8% BF16 (b) 90.7% FP8 91.2% NVFP4-unsloth 89.3% p = 3.0e-07 NVFP4-RadixArk 88.8% p = 8.4e-12 AWQ-INT4 89.7% p = 7.5e-05 INT4-RedHat 90.6% NVFP4-Inferact 84.8% p < 1e-50
Four variants are highlighted because they separate from BF16 in the paired McNemar test. The NVFP4-Inferact checkpoint is the most affected, so this benchmark also seems to be highly impacted by 4-bit activation quantization, as was the case in SWE-bench Verified.
0% 25% 50% 75% 100% Minerva-MATH BF16 (a) 97.1% BF16 (b) 97.0% FP8 97.3% NVFP4-unsloth 96.8% NVFP4-RadixArk 96.9% AWQ-INT4 96.9% INT4-RedHat 96.8% NVFP4-Inferact 96.3%
All perform well in maths.
0% 25% 50% 75% 100% MBPP+ BF16 (a) 96.3% BF16 (b) 96.6% FP8 95.8% NVFP4-unsloth 95.2% NVFP4-RadixArk 96.6% AWQ-INT4 94.4% INT4-RedHat 95.2% NVFP4-Inferact 93.1%
All perform well at creating short Python functions.
0% 25% 50% 75% 100% IFEval BF16 (a) 87.2% BF16 (b) 87.4% FP8 88.9% NVFP4-unsloth 89.3% NVFP4-RadixArk 90.0% ▲ p = 0.0127 AWQ-INT4 87.8% INT4-RedHat 86.9% NVFP4-Inferact 85.0%
NVFP4-RadixArk at p = 0.0127 (▲) shows better performance than BF16 in the McNemar test, but this seems to be just due to chance across 42 comparisons.
0% 25% 50% 75% 100% GPQA Diamond BF16 (a) 63.1% BF16 (b) 66.2% FP8 64.6% NVFP4-unsloth 64.1% NVFP4-RadixArk 68.2% AWQ-INT4 63.6% INT4-RedHat 56.1% NVFP4-Inferact 56.6%
A large percentage of no-answers in all variants, roughly 30%, and just a small set of 198 questions, so the final reported numbers are not statistically significant.

INT4-RedHat has the lowest raw score on this family of benchmarks, last on MMLU-Pro and last on GPQA, but on the paired test it is the only 4-bit checkpoint not significantly worse than BF16. This is why having both the raw scores and the McNemar test results is useful.

MMLU-PRO RAW ACCURACY BF16 83.1% AWQ-INT4 81.9% NVFP4-Inferact 80.9% INT4-RedHat 78.8% ← lowest of any variant NON-TERMINATION — MODEL PRODUCED NO ANSWER AT ALL BF16 3.8% AWQ-INT4 4.7% INT4-RedHat 8.6% ← highest of any variant
Benchmark evaluations score an empty response identically to a wrong one, so a model that fails to finish looks like a model that answers badly. INT4-RedHat is penalized by this because it has the highest no-answer rate.

Looking at the details, INT4-RedHat’s low score comes from its high percentage of no-answer results. This happens when the model reasons past its token budget and so returns no answer at all. On the questions it does answer, it is the least degraded 4-bit checkpoint.

Looking at some of the specific questions that it does not answer, the model is not looping. The traces show reasoning that is making progress, cut off mid-stream by the budget. Increasing the budget improves the results. For some reason this checkpoint needs more reasoning than the others.

So when ranking on the raw benchmark numbers INT4-RedHat is the worst checkpoint, but when ranking on the paired test it is the best 4-bit checkpoint in the set.

Surprisingly, on SWE-bench the same checkpoint had zero non-termination and performed on par with BF16.

Memory usage#

This is the memory usage for each checkpoint, as reported by vLLM on one GPU:

10 20 30 40 50 GiB of weights on one GPU BF16 51.1 GiB FP8 28.5 GiB −44% NVFP4-Inferact 24.2 GiB −53% NVFP4-unsloth 21.3 GiB −58% NVFP4-RadixArk 19.9 GiB −61% AWQ-INT4 19.2 GiB −62% INT4-RedHat 17.7 GiB −65%
Measured at load, TP=1 on one RTX PRO 6000. Percentages are against BF16.

The GPU memory budget is fixed, and I like to use a higher utilization setting (--gpu-memory-utilization 0.95), so whatever the weights do not use is assigned to the KV cache:

varianton diskweights
(+overhead)
peak
activations
CUDA
graphs
KV cacheKV tokensvs BF16
BF1651.851.93.271.8235.01,120,6271.00×
FP828.829.43.341.8657.51,842,6731.64×
NVFP4-Inferact24.625.13.271.8761.81,979,1101.77×
NVFP4-unsloth21.822.33.271.8664.62,069,5571.85×
NVFP4-RadixArk—21.13.341.8665.82,107,8831.88×
AWQ-INT419.620.13.341.8666.82,138,5431.91×
INT4-RedHat18.118.33.341.8668.62,196,7971.96×

Conclusions#

Quantization performs really well in most cases, close enough to the original model that you cannot tell them apart, but you have to be careful with the checkpoint that you choose.

In the case of Qwen3.8-27B:

  • FP8 is indistinguishable from BF16.
  • 4-bit variants are also good, but you should choose carefully based on the use case.
  • NVFP4-Inferact is the one that performs worst, due to 4-bit activations.
  • INT4-RedHat performs very well being the smallest in size, but it has a very high no-answer rate in one-shot benchmarks.

This is the final summary:

variantweightssavedmeasured cost
BF1651.1 GiB—reference
FP828.5 GiB44%none
NVFP4-Inferact24.2 GiB53%−8.0 pts SWE-bench, significant on 2 of 6 benchmarks
NVFP4-unsloth21.3 GiB58%significant on BBH only
NVFP4-RadixArk20.0 GiB61%significant on BBH only
AWQ-INT419.2 GiB62%significant on BBH and MMLU-Pro
INT4-RedHat17.7 GiB65%small, but highest no-answer rate
Measuring how much quantization affects model quality
https://javicacheiro.com/blog/what-quantization-actually-costs
Author Javier Cacheiro López
Published at September 6, 2026