I gave eleven local models the same three pieces of real work — build a web page, design a campaign as strict JSON, write a marketing one-pager with a source marker behind every claim — and ran the whole thing twice, on two different machines. Same models, same prompts, same seeds, temperature 0.
The question was not which model wins. It was whether I can trust a local model to give me the same answer twice on two different computers. 90 of 99 outputs came back byte-identical. The nine that did not were all from a single model — the only one that did not quite fit in the card.
The setup
Two hosts, each with an RTX 3090 and Ollama 0.32.5 — and beyond those two things, almost nothing in common. One is a Windows 11 workstation with a monitor attached, 128 GB of RAM and an 8-core i7-7820X. The other is a headless Ubuntu container on a Proxmox node with 24 GB and six visible threads. They are not even on the same NVIDIA driver: 610.88 against 580.126.09. Eleven models, three tasks, three seeds each: 99 requests per host, 198 in total, zero errors.
Corrected on 2026-08-02: “three seeds” here describes only the number of runs. Measured afterwards, the seed has no effect on several models — ornith:35b and qwen3.6:27b return byte-identical output across all three, while others vary even from an identical seed. The figures stand; the details are on How I Measure.
Addendum of 2026-08-03: “from an identical seed” is put too loosely. Re-measured, exactly one run per group differs — always the first, in 111 cases out of 111. Every later run is byte-identical to the others, across both machines as well (108 of 108 cell pairs). Five different seeds produce the same text as five repetitions of one seed: the seed has no effect, the position does. The measurements are unaffected.
The hardware in full, and the rules every measurement on this site follows, are on the How I Measure page.
Scoring is structural only. A script checks whether the HTML really starts with a doctype, whether the anchors resolve to IDs that exist, whether the JSON parses and has exactly eight posts, whether every claim in the one-pager carries a source marker. No model judges another model. There is no score for “reads nicely”, because I cannot measure that and neither can anyone else who says they can. Every one of the 198 outputs is kept in full so I can be contradicted.
num_ctx and num_predict were fixed on both hosts, and the prompt token count came back identical in all 99 cells. That last part matters: it is the proof that the two runs are actually comparable and not just similar-looking.
The result: two machines, one answer
96 of 99 cells scored identically on both hosts. That is the weak version of the claim. The strong version is that I compared the generated text itself, byte for byte: 90 of 99 outputs are exactly the same file. Not similar. The same.
That is worth restating, because it is the whole point. These are not two copies of one machine. Different CPU, five times the RAM on one side, Windows against Linux, and different NVIDIA driver branches. The only thing the two boxes agree on is the GPU model. Even there the boards differ: an EVGA card on one side, a Zotac on the other. Ninety times out of ninety-nine, that was enough to produce a file identical down to the byte.
I nearly published the opposite. My first comparison said 2 of 99 matched, which would have been a much more exciting story. It was wrong: the Windows host writes CRLF line endings and the Linux host writes LF, so almost every file differed in a way that had nothing to do with the models. Normalise the line endings and the real number appears. The lesson is the boring one that keeps being true — when a measurement is dramatic, suspect the measurement first.
The exception, and why it is the interesting part
All nine divergent outputs belong to qwq:32b. It is also the only model in the lineup that does not fit entirely into 24 GB: it needs 24.8 GB, so a slice of it runs on the CPU. On the workstation that slice is 6.6 %, on the headless container 3.5 % — because the monitor costs roughly 1.8 GB of VRAM before a single model is loaded.
Two machines, the same model, the same seed, different amounts of CPU offload — and the text comes out different. Not wildly: on the marketing task qwq scored 12 of 13 on one host and 13 of 13 on the other, and on the web page task the workstation was the better one. It goes both ways, which is exactly what you would expect from an arithmetic difference rather than a quality difference.
The practical rule: as long as the model fits in VRAM, seed and temperature 0 give you a reproducible result across machines. The moment it spills to CPU, that guarantee is gone. If reproducibility matters to you — regression tests, audit trails, anything you have to defend later — the model has to fit. Not “mostly fit”.
The ranking
Score is the mean over all nine cells per model. Throughput is the median across the run on the workstation.
| Model | Score | tok/s |
|---|---|---|
| gemma4:26b | 100.0 % | 99 |
| ornith:35b | 97.0 % | 119 |
| qwen3:8b | 96.7 % | 115 |
| qwen3.6:27b | 93.9 % | 39 |
| qwq:32b | 93.2 % | 19 |
| deepseek-r1:14b | 90.1 % | 68 |
| devstral:24b | 88.3 % | 48 |
| glm-4.7-flash | 85.9 % | 120 |
| qwen2.5vl:7b | 82.4 % | 120 |
| qwen3-vl-abliterated:30b (Thinking) | 74.8 % | 133 |
| laguna-xs-2.1 | 53.8 % | 128 |
gemma4:26b is the only model that did not put a foot wrong in 27 attempts, and it did it at 99 tok/s. Note also that qwen3:8b — the smallest model here — lands third. Size did not decide this.
The bottom of the table is a different failure. laguna-xs-2.1 hit the 8000-token ceiling in all nine of its runs; it simply never stopped writing. Its 53.8 % is not a model that answers badly, it is a model that does not finish. The Thinking variant of the VL model ran into the same wall twice.
What one task looks like eleven times
All eleven models were given the same brief: a complete, self-contained HTML page about peptides. Rendered at 1280 by 900 pixels, one seed per model, top of the page. The order follows the ranking above.











The scores measure the stated requirements, not taste. Looking through them, it is still striking that the models at the top do not merely meet the requirements — they throw in a navigation bar, section cards and readable typography that nobody asked for.
What the tasks actually caught
The check that failed most often was “the HTML starts with a doctype” — 21 of 33 attempts. That number is misleading on its own, and I want to be precise about it: not one model forgot the doctype. Fifteen wrapped the page in a ```html code fence, and six put a sentence in front of it. The task said to output the HTML source and nothing else. So the failures are real, but they are obedience failures, not competence failures — and reporting them as “no doctype” would have been simply untrue.

The JSON task was the harshest overall at 79.7 % across all models: six of 33 attempts produced JSON that would not parse at all. Strict output formats are still where local models lose most cheaply.
The three tasks, verbatim
So the numbers can be checked, here are the briefs every model received word for word. They were written in German, because that is the language I work in; each one ends with a fixed format block that I have left out here, since it only governs the output shape.
webseite
Erstelle eine vollstaendige, in sich geschlossene HTML-Seite ueber Peptide und ihre Wirkungen.
Pflichtanforderungen:
- Beginne mit <!doctype html>, setze lang="de" am <html>-Tag und einen Viewport-Meta-Tag.
- Mindestens 4 <section>-Bereiche mit je einer Ueberschrift.
- Eine Navigation, deren Links auf Anker (#id) zeigen, die es im Dokument auch gibt.
- Eine <table> mit mindestens 5 Datenzeilen, die Peptidgruppen und ihre beschriebenen Wirkungen gegenueberstellt.
- Das gesamte CSS in einem <style>-Block im Dokument.
- Keine externen Dateien, keine Bilder von fremden Servern, keine Schriften von CDNs.
- Kein Platzhaltertext (kein Lorem Ipsum). Alle Texte inhaltlich ausformuliert, mindestens 400 Woerter Flaechentext.
- Ein Abschnitt muss ausdruecklich darauf hinweisen, dass die Seite keine medizinische Beratung ersetzt.
Gib ausschliesslich den HTML-Quelltext aus, ohne Erklaerung davor oder danach.
html
kampagne
Entwirf eine Social-Media-Kampagne fuer ein Unternehmen, das Peptid-Praeparate fuer Kosmetik vertreibt.
Gib AUSSCHLIESSLICH gueltiges JSON aus, ohne Text davor oder danach, exakt nach diesem Schema:
{
"campaign_name": "...",
"audience": "...",
"posts": [
{"platform": "linkedin|instagram|x|facebook", "text": "...", "hashtags": ["...", "..."], "cta": "..."}
],
"kpis": [
{"metric": "...", "target": "..."}
]
}
Regeln:
- Genau 8 Eintraege in "posts".
- Jede der vier Plattformen kommt mindestens einmal vor.
- Bei platform "x" darf "text" hoechstens 280 Zeichen lang sein.
- Jeder Post hat 2 bis 5 Hashtags.
- Keine zwei Posts duerfen denselben Text haben.
- Mindestens 3 Eintraege in "kpis", jeder mit "metric" und "target".
- Keine Heilversprechen; formuliere kosmetisch, nicht medizinisch.
json
marketing
Schreibe einen Marketing-Einseiter (Markdown) fuer eine Peptid-Kosmetikserie.
Verwende exakt diese Ueberschriften, jede genau einmal, als Ebene 2:
## Positionierung
## Zielgruppe
## Headline-Varianten
## Vergleich
## Wirkaussagen und Belege
## Rechtlicher Hinweis
Regeln:
- Unter "Headline-Varianten" genau 3 Aufzaehlungspunkte, jeder hoechstens 60 Zeichen.
- Unter "Vergleich" eine Markdown-Tabelle mit Kopfzeile und mindestens 4 Datenzeilen.
- Unter "Wirkaussagen und Belege" jede Aussage als eigener Aufzaehlungspunkt, und JEDER dieser
Punkte endet mit einem Belegmarker in eckigen Klammern: entweder [Beleg: <Quelle>] wenn du eine
Quelle benennen kannst, oder [Beleg fehlt] wenn nicht. Kein Punkt ohne Marker.
- Unter "Rechtlicher Hinweis" mindestens zwei Saetze zur Abgrenzung von medizinischen Aussagen.
Gib ausschliesslich das Markdown aus.
md
What this does not show
Three tasks are three tasks. Everything here was measured with think disabled, which is a real constraint for the reasoning models in the lineup. Both hosts run the same Ollama version on the same GPU model, on boards from two different makers — this says nothing about determinism across different Ollama versions, different quantisations, or different hardware. And structural scoring rewards a model that follows instructions; it cannot tell you whether the marketing copy is any good.
Corrected on 2026-08-02: the switch does not control whether the model thinks. For qwq:32b the output is byte-identical with and without think; for deepseek-r1:14b the generated token count stays exactly the same (1402) and only the reasoning blocks move out of the answer text. For gemma4:26b (562 → 1261 tokens) and qwen3:8b (730 → 1448) it behaves as expected. The caveat “measured with thinking mode off” therefore does not hold for every model. The measurements are unaffected. Figures corrected on 2026-08-03: the two pairs inadvertently compared different tasks (scheduling against the logic puzzle, the invoice against scheduling). The correct figures are 562 → 1261 and 730 → 1448; the claim itself is unchanged.
What it does show is narrow and, I think, worth having: on identical hardware with a model that fits, a local LLM is a reproducible component. You can put it in a pipeline and expect the same bytes tomorrow. Give it 800 MB too little VRAM and it becomes something you have to re-check every time.
Clarified on 2026-08-02: this originally said only “the same GPU model”. True, but understated. The two cards are not the same board — one is an EVGA, the other a Zotac, verified via the PCI subsystem IDs 0x3842 and 19da:1613. Same silicon, different boards. That makes the finding stronger rather than weaker: byte-identical outputs across two board makers say more than across two identical cards. None of the measurements change.
Addendum, 15 August 2026. When I went looking for these outputs for a new comparison, they were not there — the second machine had been wiped on 3 August for a reinstall. The sentence above, that all 198 outputs are available in full, was untrue for a few days. I found them in a backup taken on the day of the reinstall. They are public now: all 198 outputs, separated by machine, together with the per-request logs and the task definitions as a script. Anyone who wants to recompute no longer has to ask me. Why I did not do it this way from the start is in this piece.