AI & Tools

Four 4-bit Qwen3.8-27B builds on one RTX 3090: the models tied, the context window didn’t

Infographic in two panels. Left, valid passes of four 4-bit builds of Qwen3.8-27B: library Q4_K_M with MTP 24/24, unsloth Q4_0 23/23, abliterated and uncensored Q4_K_M 21/22 each, all at the ceiling. Right, context: 32k would cut off 10.2 % of 3,408 real prompts, 128k 5.9 %; 128k with q4_0 KV fits in 21.3 GB, 262k with q4_0 leaves 57 of 66 layers on the GPU.
The tasks were too easy to separate the four builds; the context size is what changed the plan.

I wanted to split a dual-RTX-3090 box into “one model per card”: two independent 27B models instead of one model stretched across both GPUs at a huge context window. Before doing that, I wanted two answers. Which 4-bit build of Qwen3.8-27B should go on a card? And how much context does a single 24 GB card actually hold? The first question had a boring answer. The second one changed the plan.

The setup

  • One RTX 3090 (24 GB), Ollama 0.35.1, a dedicated server instance pinned to that one card. Nothing else was allowed on the GPU while measuring. A gate checked for foreign processes before and after every request.
  • KV cache q4_0, flash attention on, context fixed at 32,768 tokens at the server level. Never per request; more on why below.
  • I verified every runner’s real flags (-c 32768, KV type, flash attention, speculative decoding) on the process itself, not in a config file.
  • Identical sampling for every model (temperature 1.0, top_k 20, top_p 0.95), two seeds, a unique prompt per request, and every model warmed up before measuring.
  • Twelve tasks: four reasoning tasks with exact answers, four coding tasks whose output is executed against a reference implementation, and four instruction-following tasks with programmatic checks.
  • Every checker had to accept a known-good answer and reject a known-bad one before it was allowed to grade anything.
  • A run that hit the token limit (6,144) or returned an empty answer counts as invalid, not as a failure. The limit is a property of my test, not a verdict on the model.

Finding 1: the four 4-bit builds tie, and the speed gap is packaging

buildvalid passesinvalidgeneration tok/sprompt tok/sVRAM @32k
Ollama library qwen3.8:27b (Q4_K_M, with MTP head + vision)24/24043.51,05618.6 GB
unsloth Q4_023/23135.81,30816.8 GB
community “abliterated” Q4_K_M21/22232.61,22216.4 GB
community “uncensored” Q4_K_M21/22230.61,16416.4 GB
Bar chart of generation speed for four 4-bit builds of Qwen3.8-27B on one RTX 3090: Ollama library Q4_K_M with MTP head 43.5 tok/s, unsloth Q4_0 35.8, community abliterated Q4_K_M 32.6, community uncensored Q4_K_M 30.6. A side panel notes 21–42 % more generation speed from the MTP head and 7–24 % faster prompt reading for Q4_0.
On this test, the MTP head is the only difference that is clearly larger than the noise.

With 24 graded cells per model, one or two cells of difference mean nothing. These tasks are simply too easy to separate 27B models. All four sit at the ceiling. I can neither confirm nor rule out that the “uncensored” builds lose capability.

The robust differences are elsewhere:

  • Speed comes from packaging, not quantization. The Ollama library build ships a multi-token-prediction (MTP) head and runs speculative decoding with it. That is worth 21–42 % more generation speed than the plain GGUFs. Q4_0 reads prompts 7–24 % faster than the Q4_K_M builds.
  • The failures are interesting individually.
    • Three of the five invalid runs were a calendar question (“what date is 1,000 days after 14 March 2027?”). The models thought for up to 13,000 characters and never answered.
    • One code failure was a real bug: a Roman-numeral parser that rejected the letter “I”.
    • One instruction failure is debatable. The answer had two sentences, but the second one had no full stop. My checker counts terminators. I did not re-grade it after the fact, because the checker was fixed before the run.

Finding 2: 32k context is not enough. I measured it on live traffic.

My plan assumed 32k per model. While the benchmark ran, the production server was temporarily limited to one card at 32k. I looked at what its clients had actually sent over the previous seven days (3,408 prompts):

prompt longer thanshare
16k tokens21.0 %
32k tokens10.2 %
64k tokens8.0 %
128k tokens5.9 %
Bar chart of 3,408 real prompts from seven days: 21.0 % were longer than 16k tokens, 10.2 % longer than 32k, 8.0 % longer than 64k and 5.9 % longer than 128k. A side panel lists the median of 245 tokens, the 99th percentile of 238,201, and a logged prompt of 44,602 tokens cut at a limit of 16,386, so the answer used 37 % of the input.
A 32k window would silently cut off 10.2 % of the prompts; the long ones come from coding agents.

The median prompt is 245 tokens, but the 99th percentile is 238,201. Those long prompts are coding agents carrying tool definitions, file contents and history. Eight minutes after the switch, the log showed one of them: truncating input prompt limit=16386 prompt=44602. The client got an answer based on 37 % of what it had sent. It saw no error and no warning.

Finding 3: what one 24 GB card actually holds

The VRAM per context size was measured, one fresh server instance per configuration, for the 18.6 GB library build:

contextKV q4_0KV q8_0
32k✅ 18.6 GB✅ 19.1 GB
64k✅ 19.5 GB✅ 20.5 GB
128k✅ 21.3 GB❌ 63 of 66 layers on GPU
262k❌ 57 of 66 layers on GPU❌ 46 of 66 layers on GPU
Bar chart of measured VRAM on one 24 GB card for the 18.6 GB library build. With a q4_0 KV cache: 18.6 GB at 32k, 19.5 GB at 64k, 21.3 GB at 128k. With q8_0: 19.1 GB at 32k, 20.5 GB at 64k. Marked with a cross because they do not fit: 128k with q8_0 (63 of 66 layers on the GPU) and 262k with q4_0 (57 of 66) or q8_0 (46 of 66).
128k with a q4_0 cache fits on one card; beyond that, Ollama moves layers to the CPU instead of failing.
  • The context cost is cheap and linear. About 27 KiB per token with q4_0, about 42 KiB with q8_0. This model family uses full attention in only a fraction of its layers, so the KV cache is far smaller than for a classic transformer of the same size. The linear model predicted the 262k miss before it was measured.
  • 128k with q4_0 fits. By arithmetic it also fits on the second card, next to the embedding model that lives there (21.3 + 1.4 GB). That part is calculated, not measured on that card. It would cut off 5.9 % of real traffic instead of 10.2 %.
  • Is a q4_0 KV cache still faithful at long context? One needle test: a code at the very start of a 101,482-token prompt, asked for at the end. It came back exactly. Prompt processing dropped to 779 tok/s at that length, but generation stayed at 43.2 tok/s, the same as with a short prompt. One needle is a positive signal, not a proof.
  • When it doesn’t fit, Ollama does not fail. It silently moves layers to the CPU (the “57 of 66” above) and the model keeps answering, just slower. No error, no warning in the API. Only the offload line in the server log tells you.

Where Ollama itself bit me

Ollama 0.35 runs llama.cpp’s llama-server underneath, so the kernels and the speed are the same. The surprises all came from the orchestration layer around it:

  • The default context follows the available VRAM, not my configuration. With 48 GB across two cards, Ollama loaded the model at 256k all by itself. Neither the model file nor my pin job asks for that. So unloading it would not have freed a card: the next client request would have loaded it straight back at 256k across both GPUs. That is why I moved the production instance to one card instead of just unloading the model.
  • Prompts that are too long are silently truncated (Finding 2).
  • A request whose num_ctx differs from the loaded runner disappears. No error, no log line; the client just times out. That is why context is set at the server, never per request.
  • “Doesn’t fit” means “silently slower” (Finding 3).

For measuring, I will run llama-server directly from now on: explicit flags, loud failures, per-request timings, one model per process. For “what do my clients actually get”, Ollama stays in the loop.

My own measurement errors

  • My idle check could never succeed. Before restarting the production server, I waited for the GPU to be idle, but I read utilization for the whole device. A virtual machine renders its desktop on that card, so the device was never idle. Worse, my script restarted anyway, because I had not made the restart depend on the check. The log shows no aborted request, but I cannot prove there was none. The fix: gate on the server process’s own utilization, and do not restart at all if it never goes quiet.
  • My harness crashed on its own report. After the first model, the summary step tripped over the run’s metadata row. The measurements were already on disk, and I resumed from there. Reporting code is now no longer able to end a measurement loop.

What I’m doing with it

  • One model per card stays the plan, but at 128k with q4_0 KV, not 32k.
  • Pick the build with the MTP head. On this test, it is the only difference that is clearly larger than the noise.
  • Write the next test set to be harder. A benchmark everything passes measures nothing.