An Nvidia GB10 has 124.5 GB of shared memory, that CPU and GPU split between them. I ran a single model on it: a 27B-class runner with a large context window, instead of two models on two cards. Two questions were on the table. What does the context window actually cost at 124.5 GB? And will a model that thinks for up to 32.768 tokens about a task still ship a working game of Tetris at the end?
The setup
- One GB10: 124.5 GiB of unified memory, a 20-core CPU, one Blackwell GPU, 3,092 GiB of NVMe. The model lives on disk and is streamed through the page cache into the accelerator (mmap, compressed). Nothing else was allowed on the GPU during measurement; a lock checked for stray processes before and after every cell.
- Model:
Qwen3.8-Flash-Next-abliterated, IQ3_S quantization, two shards (66.97 GiB on disk, roughly 61.8 GB of weights), f16 KV cache, Flash Attention on, context pinned at the server. llama-serverdirectly, not Ollama. The reasoning is in my post on the four 4-bit builds: explicit flags, loud failures, per-request timings, one model per process.- Identical sampling parameters, a separate seed per rep, a cell only counts once it has three valid reps, medians with standard deviation. Rejected probes never enter any median.
- The Tetris task is the same one from October 9th: one single generated HTML file containing a fully playable Tetris, 16 checks, the code executed, not read. Every scorer had to accept a known-good answer and reject a known-bad one before it may grade. Three reps, a 32.768-token budget, thinking on the model default.
Finding 1: The 27 tokens/s still holds at 32.768 context
Six cells across prompt lengths from 512 to 32.768 tokens, three valid reps each:
| Prompt | Prefill PP tok/s | Generation TG tok/s | first token |
|---|---|---|---|
| 512 | – (overhead) | 27.36 | 0.86 s |
| 2,048 | 738 | 27.08 | 2.86 s |
| 8,192 | 716 | 26.95 | 11.58 s |
| 32,768 | 675 | 25.83 | 49.45 s |

- Generation barely sees the context window. From 27.36 tok/s on short prompts to 25.83 at 32.768: minus 5.6%. Even at a context depth of 65.536 tokens it is 21.12 tok/s against 26.77 at depth 0. 64k of context costs about a fifth, not an order of magnitude.
- The cost of context sits in the prefill. The first token moves from 0.86 s to 49.45 s. This is the same 32k prompt that a 24-GB card silently truncated on October 9th — here it just gets read slowly.
- The same prefix twice costs almost nothing. Cache hit: 50.21 s without, 0.101 s with. A factor of 500. For agents that carry the same context through repeated tool calls, this is the difference between lukewarm and usable.
Finding 2: It ties on generation, it loses on prompt loading
Base run, 2.048 prompt / 512 output, idle, freshly loaded — against the measurement on my two-card machine (2× RTX 3090) from October 9th:
| GB10 (124.5 GB) | Two cards (2× 24 GB) | Difference | |
|---|---|---|---|
| Prefill, tok/s | 738 | 970 | −23.9% |
| Generation, tok/s | 27.08 | 26.6 | +1.8% |

So the 124.5 GB do not make the machine faster at decoding — it is at eye level. What they do: model, its 61.8 GB of weights, the KV cache for a big window, and the operating system all fit under 124.5 GB. Nothing gets offloaded. On a 24-GB card, „does not fit” means „silently slower„: Ollama spills layers to the CPU, the model keeps answering, nobody notices. On the GB10, „fits” is simply true.
And to be honest: anyone who needs long prompts in real time does it faster on the old two-card box. 49.5 seconds for a 32k prompt is still a lot.
Finding 3: Tetris — 32.768 tokens reach the thinking, not the building
The task from October 9th, now on the GB10: one single HTML file containing a fully playable Tetris. Three reps, separate seed per run, thinking on the model default, 32.768-token budget:
| Rep | thinking tokens | content tokens | first content | end |
|---|---|---|---|---|
| 1 | 29,912 | 2,856 | after 1,337 s | length cap (TRUNCATED) |
| 2 | 32,768 | 0 | – | length cap (TRUNCATED) |
| 3 | 29,635 | 3,133 | after 1,323 s | length cap (TRUNCATED) |

All three runs finished after roughly 24 minutes with done: length at 22.4 generation tokens/s. In two runs the content only starts after the model has written roughly 30.000 tokens into its thinking chain; in the second run no content token ever arrives. The model did not crash and produced no error — it simply did not finish before the budget ran out. Thinking was on the model default; I neither enabled nor disabled it.
On October 9th, the same task on another machine: two models of this size class finished it with 30/35 and 32/35 checks (86% and 91%). One 30B model failed it outright (12/35), and one coder model hung on its own verifier. The task is solvable — on the GB10 the budget was just too small to show it.
By the rule from the 4-bit post, a run that hits the token limit counts as invalid, not as a failure. That is exactly what happened here: not a verdict on the model, but a measurement of a budget that was too small for a thinking model.
My own measurement errors
- 36 of 409 measurements had rising throttling counters (software thermal or power cap), GPU temperature peaked at 85°C. My sanity checks flagged them and kept them out of the medians. Without the checks they would have slipped in.
- 136 of 307 cells in the larger model battery reached the minimum of three reps; the rest are reported as „incomplete”. I did not interpolate anything.
- One by assumption, not by measurement: I assumed the thinking chain would stop on its own. It can run 32.768 tokens long — longer than a whole coding task. The cap is a property of the test rig, and it was too small.
What I do with this
- Tetris goes into the next nightly window twice: once with thinking explicitly off, once at a 96k budget. A 128k budget would fit on the numbers (124.5 GB); unmeasured.
- Candidates are queued:
Qwen3.8-27Bin FP8 and a 274B MoE at Q2 quantization (97.6 GB — fits with roughly 20 GiB left over for the KV cache and the OS). What 124.5 GB buys with a genuinely bigger model will be the next post. - The production instance stays where it is: Flash-Next on the GB10, port 18180. The benchmark stopped it overnight and restored it by 5 a.m.; the closing probe passed.
- Backlink: how the same 27B family gives out on 24 GB (though only silently) is in “Four 4-bit Qwen3.8-27B builds on one RTX 3090”: the models tied, the context window didn’t.