Writing

Quantising Qwen3.8-27B: Where Quality Plateaus

· AI · Local AI

Vendors benchmark full-precision models on datacentre hardware. llama.cpp users download quantised GGUFs, which use fewer bits per weight to fit in less memory, and the files people actually run rarely get task-level evaluation. Uploaders publish KL divergence and similar agreement measures; a source audit of publisher-reported Qwen3.8 measurements found agreement with Qwen3.8-27B BF16 rising with precision, with the warning that this does not prove practical tasks improve.

Executed code measures something different. A nearly correct program still fails its tests. So I ran five quantisations of Qwen3.8-27B through the same coding exam on one RTX 5090 and recorded what each level of precision cost in speed and context.

The setup

  • RTX 5090 with 32GB of VRAM and 96GB of DDR5.
  • llama.cpp with CUDA, build 10524, commit 9ee9fc04c. EvalPlus 0.3.1. Greedy decoding, 16,000-token context, two parallel slots, reasoning disabled.
  • HumanEval+ and MBPP+ through EvalPlus, scored as pass@1.
  • Qwen3.8-27B Q4_K_M, Qwen3.8-27B Q4_K_XL, Qwen3.8-27B Q5_K_M, Qwen3.8-27B Q6_K and Qwen3.8-27B Q8_0 files, with downloads checked against published byte sizes.

Reasoning was disabled so the sweep could finish overnight. That makes the absolute scores lower than other settings and not comparable with them. The relative test is fair: same model, software and decoding rules throughout.

The tests and the columns

TermWhat it is
HumanEval+OpenAI’s 164 hand-written Python problems, with EvalPlus adding about 80 times more test cases per problem than the original.
MBPP+378 problems from Google’s Mostly Basic Python Problems, curated by EvalPlus with about 35 times more test cases.
pass@1The percentage of problems whose first and only answer passes every test. Greedy decoding, no retries.
Exam speedGeneration speed in tokens per second, recorded during the exam runs: 16k context, MTP off.
Largest fitted contextThe largest context window each file loads with on the 32GB card under the everyday serving flags: MTP on, 8-bit KV cache, two slots. Measured in separate load tests, not during the exam.

MTP is multi-token prediction, a decoding speed-up. It was off for the exam so every file ran under identical rules, and on for the fit tests because that is how the card is used day to day. Speed and context are therefore two separate facts about each file, not one controlled comparison.

What each quantisation cost

QuantisationMBPP+HumanEval+Exam speedLargest fitted context
Qwen3.8-27B Q4_K_M55.046.361.5 t/s256k
Qwen3.8-27B Q4_K_XL57.941.562.8 t/s256k
Qwen3.8-27B Q5_K_M57.745.753.6 t/s224k
Qwen3.8-27B Q6_K59.348.849.9 t/s160k
Qwen3.8-27B Q8_056.945.755.2 t/sNot measured

Bar chart showing MBPP plus pass at one scores for five Qwen3.8-27B quantisations

MBPP+ pass@1 across 378 executed problems. Olive marks the two 4-bit files and rust marks Q5, Q6 and Q8.

Where quality plateaued

Quality did not rise steadily with bits. On MBPP+, Qwen3.8-27B Q5_K_M, Qwen3.8-27B Q6_K and Qwen3.8-27B Q8_0 cluster between 56.9 and 59.3, and the differences inside that cluster are within exam noise. There is no measured basis for choosing Qwen3.8-27B Q6_K or Qwen3.8-27B Q8_0 over Qwen3.8-27B Q5_K_M.

Qwen3.8-27B Q4_K_M sits at 55.0, a gap of 1.9 to 4.3 points below the cluster. That is suggestive rather than conclusive: the largest comparison, against Qwen3.8-27B Q6_K, is roughly 2.5 standard errors. Qwen3.8-27B Q4_K_XL breaks any simple four-bit rule by scoring 57.9 on MBPP+, inside the cluster, and 41.5 on HumanEval+, the lowest in the table.

The familiar claim that four-bit quantisation costs about one point comes from knowledge and language benchmarks. Here it cost 1.9 to 4.3 points on executed code, roughly double that figure at the near end, with uncertainty too wide for a universal rule.

HumanEval+ cannot rank the files

Qwen3.8-27B Q6_K scored 48.8, Qwen3.8-27B Q8_0 scored 45.7 and Qwen3.8-27B Q4_K_XL scored 41.5. Added precision cannot credibly cause those swings in true ability. With 164 problems, a few points is a handful of answers and sampling noise can reverse the order. MBPP+ at 378 problems resolves more, but still only broad bands: Qwen3.8-27B Q5_K_M, Qwen3.8-27B Q6_K and Qwen3.8-27B Q8_0 in one noise band, Qwen3.8-27B Q4_K_M with a possible modest penalty, and no exact universal four-bit cost.

What the extra precision bought

Speed fell with precision under matched settings: 61.5 t/s for Qwen3.8-27B Q4_K_M, 53.6 for Qwen3.8-27B Q5_K_M and 49.9 for Qwen3.8-27B Q6_K. Context fell too, in the separate fit tests: 256k, 224k and 160k tokens on the same card.

Qwen3.8-27B Q8_0 is the warning case. It is the near-full-precision reference, and it scored 56.9 on MBPP+ and 45.7 on HumanEval+, below Qwen3.8-27B Q6_K on both and inside the noise band. The extra memory bought no measurable task-level return.

Other measurements

Two other public measurements touch the same question. Neither ran task-level tests on quantised files.

SourceDateFilesWhat it measuredWhat it found
Qwen model card14 AugQwen3.8-27B BF16 onlyLiveCodeBench v6, Terminal-Bench 2.1, SWE-bench Pro, reasoning on90.3, 73.0 and 61.7. Nothing on quantised files.
kingy.ai audit of AtomicChat’s tests17 AugQwen3.8-27B Q4_K_XL to Qwen3.8-27B Q8_0Token agreement with Qwen3.8-27B BF16, no tasksAgreement rises with every step: 96.0, 97.3, 97.9 and 98.9 percent for the closest matches to Qwen3.8-27B Q4_K_XL, Qwen3.8-27B Q5_K_M, Qwen3.8-27B Q6_K and Qwen3.8-27B Q8_0.

The model card numbers are for Qwen3.8-27B BF16 with reasoning on and sampling at temperature 1.0, so they say nothing about the files people download and cannot be compared with the reasoning-off scores here. The agreement table is the closer neighbour, and it makes the plateau sharper: agreement with Qwen3.8-27B BF16 keeps improving above Qwen3.8-27B Q5_K_M, from 97.3 to 98.9 percent, while the executed-code scores do not move outside noise. Above Qwen3.8-27B Q5_K_M the extra bits buy closer token distributions, not more passing programs.

Practical rules

  • Run Qwen3.8-27B Q4_K_M when the context it buys will be used. Its quality cost is suggested, not established exactly.
  • Run Q5 when spare memory should buy quality. It was indistinguishable from Q6 and Q8 here and fits more context than Q6.
  • Distrust small rankings. HumanEval+ could not order the files.
  • Measure the artefact you will run. Qwen3.8-27B BF16 vendor results and KL tables do not answer what a particular GGUF does on your task.

On this card, quality plateaus from Q5 upward and Q4 carries a possible modest penalty. Q4 remains the context-first choice and Q5 is the most precision these results justify. The concrete cost of going higher was 61.5 falling to 49.9 t/s and 256k falling to 160k tokens of context, with no measured coding return.

← All posts