AI & Tools

Same Prompt, Two Machines: 90 of 99 Outputs Were Byte-Identical

I gave eleven local models the same three pieces of real work — build a web page, design a campaign as strict JSON, write a marketing one-pager with a source marker behind every claim — and ran the whole thing twice, on two different machines. Same models, same prompts, same seeds, temperature 0.

The question was not which model wins. It was whether I can trust a local model to give me the same answer twice on two different computers. 90 of 99 outputs came back byte-identical. The nine that did not were all from a single model — the only one that did not quite fit in the card.

The setup

Two hosts, each with an RTX 3090 and Ollama 0.32.5 — and beyond those two things, almost nothing in common. One is a Windows 11 workstation with a monitor attached, 128 GB of RAM and an 8-core i7-7820X. The other is a headless Ubuntu container on a Proxmox node with 24 GB and six visible threads. They are not even on the same NVIDIA driver: 610.88 against 580.126.09. Eleven models, three tasks, three seeds each: 99 requests per host, 198 in total, zero errors.

The hardware in full, and the rules every measurement on this site follows, are on the How I Measure page.

Scoring is structural only. A script checks whether the HTML really starts with a doctype, whether the anchors resolve to IDs that exist, whether the JSON parses and has exactly eight posts, whether every claim in the one-pager carries a source marker. No model judges another model. There is no score for “reads nicely”, because I cannot measure that and neither can anyone else who says they can. Every one of the 198 outputs is kept in full so I can be contradicted.

num_ctx and num_predict were fixed on both hosts, and the prompt token count came back identical in all 99 cells. That last part matters: it is the proof that the two runs are actually comparable and not just similar-looking.

The result: two machines, one answer

96 of 99 cells scored identically on both hosts. That is the weak version of the claim. The strong version is that I compared the generated text itself, byte for byte: 90 of 99 outputs are exactly the same file. Not similar. The same.

That is worth restating, because it is the whole point. These are not two copies of one machine. Different CPU, five times the RAM on one side, Windows against Linux, and different NVIDIA driver branches. The only thing the two boxes agree on is the GPU model — and ninety times out of ninety-nine, that was enough to produce a file identical down to the byte.

I nearly published the opposite. My first comparison said 2 of 99 matched, which would have been a much more exciting story. It was wrong: the Windows host writes CRLF line endings and the Linux host writes LF, so almost every file differed in a way that had nothing to do with the models. Normalise the line endings and the real number appears. The lesson is the boring one that keeps being true — when a measurement is dramatic, suspect the measurement first.

The exception, and why it is the interesting part

All nine divergent outputs belong to qwq:32b. It is also the only model in the lineup that does not fit entirely into 24 GB: it needs 24.8 GB, so a slice of it runs on the CPU. On the workstation that slice is 6.6 %, on the headless container 3.5 % — because the monitor costs roughly 1.8 GB of VRAM before a single model is loaded.

Two machines, the same model, the same seed, different amounts of CPU offload — and the text comes out different. Not wildly: on the marketing task qwq scored 12 of 13 on one host and 13 of 13 on the other, and on the web page task the workstation was the better one. It goes both ways, which is exactly what you would expect from an arithmetic difference rather than a quality difference.

The practical rule: as long as the model fits in VRAM, seed and temperature 0 give you a reproducible result across machines. The moment it spills to CPU, that guarantee is gone. If reproducibility matters to you — regression tests, audit trails, anything you have to defend later — the model has to fit. Not “mostly fit”.

The ranking

Score is the mean over all nine cells per model. Throughput is the median across the run on the workstation.

ModelScoretok/s
gemma4:26b100.0 %99
ornith:35b97.0 %119
qwen3:8b96.7 %115
qwen3.6:27b93.9 %39
qwq:32b93.2 %19
deepseek-r1:14b90.1 %68
devstral:24b88.3 %48
glm-4.7-flash85.9 %120
qwen2.5vl:7b82.4 %120
qwen3-vl-abliterated:30b (Thinking)74.8 %133
laguna-xs-2.153.8 %128

gemma4:26b is the only model that did not put a foot wrong in 27 attempts, and it did it at 99 tok/s. Note also that qwen3:8b — the smallest model here — lands third. Size did not decide this.

The bottom of the table is a different failure. laguna-xs-2.1 hit the 8000-token ceiling in all nine of its runs; it simply never stopped writing. Its 53.8 % is not a model that answers badly, it is a model that does not finish. The Thinking variant of the VL model ran into the same wall twice.

What one task looks like eleven times

All eleven models were given the same brief: a complete, self-contained HTML page about peptides. Rendered at 1280 by 900 pixels, one seed per model, top of the page. The order follows the ranking above.

The scores measure the stated requirements, not taste. Looking through them, it is still striking that the models at the top do not merely meet the requirements — they throw in a navigation bar, section cards and readable typography that nobody asked for.

What the tasks actually caught

The check that failed most often was “the HTML starts with a doctype” — 21 of 33 attempts. That number is misleading on its own, and I want to be precise about it: not one model forgot the doctype. Fifteen wrapped the page in a ```html code fence, and six put a sentence in front of it. The task said to output the HTML source and nothing else. So the failures are real, but they are obedience failures, not competence failures — and reporting them as “no doctype” would have been simply untrue.

The JSON task was the harshest overall at 79.7 % across all models: six of 33 attempts produced JSON that would not parse at all. Strict output formats are still where local models lose most cheaply.

The three tasks, verbatim

So the numbers can be checked, here are the briefs every model received word for word. They were written in German, because that is the language I work in; each one ends with a fixed format block that I have left out here, since it only governs the output shape.

webseite

Erstelle eine vollstaendige, in sich geschlossene HTML-Seite ueber Peptide und ihre Wirkungen.

Pflichtanforderungen:
- Beginne mit <!doctype html>, setze lang="de" am <html>-Tag und einen Viewport-Meta-Tag.
- Mindestens 4 <section>-Bereiche mit je einer Ueberschrift.
- Eine Navigation, deren Links auf Anker (#id) zeigen, die es im Dokument auch gibt.
- Eine <table> mit mindestens 5 Datenzeilen, die Peptidgruppen und ihre beschriebenen Wirkungen gegenueberstellt.
- Das gesamte CSS in einem <style>-Block im Dokument.
- Keine externen Dateien, keine Bilder von fremden Servern, keine Schriften von CDNs.
- Kein Platzhaltertext (kein Lorem Ipsum). Alle Texte inhaltlich ausformuliert, mindestens 400 Woerter Flaechentext.
- Ein Abschnitt muss ausdruecklich darauf hinweisen, dass die Seite keine medizinische Beratung ersetzt.

Gib ausschliesslich den HTML-Quelltext aus, ohne Erklaerung davor oder danach.

html

kampagne

Entwirf eine Social-Media-Kampagne fuer ein Unternehmen, das Peptid-Praeparate fuer Kosmetik vertreibt.

Gib AUSSCHLIESSLICH gueltiges JSON aus, ohne Text davor oder danach, exakt nach diesem Schema:

{
  "campaign_name": "...",
  "audience": "...",
  "posts": [
    {"platform": "linkedin|instagram|x|facebook", "text": "...", "hashtags": ["...", "..."], "cta": "..."}
  ],
  "kpis": [
    {"metric": "...", "target": "..."}
  ]
}

Regeln:
- Genau 8 Eintraege in "posts".
- Jede der vier Plattformen kommt mindestens einmal vor.
- Bei platform "x" darf "text" hoechstens 280 Zeichen lang sein.
- Jeder Post hat 2 bis 5 Hashtags.
- Keine zwei Posts duerfen denselben Text haben.
- Mindestens 3 Eintraege in "kpis", jeder mit "metric" und "target".
- Keine Heilversprechen; formuliere kosmetisch, nicht medizinisch.

json

marketing

Schreibe einen Marketing-Einseiter (Markdown) fuer eine Peptid-Kosmetikserie.

Verwende exakt diese Ueberschriften, jede genau einmal, als Ebene 2:
## Positionierung
## Zielgruppe
## Headline-Varianten
## Vergleich
## Wirkaussagen und Belege
## Rechtlicher Hinweis

Regeln:
- Unter "Headline-Varianten" genau 3 Aufzaehlungspunkte, jeder hoechstens 60 Zeichen.
- Unter "Vergleich" eine Markdown-Tabelle mit Kopfzeile und mindestens 4 Datenzeilen.
- Unter "Wirkaussagen und Belege" jede Aussage als eigener Aufzaehlungspunkt, und JEDER dieser
  Punkte endet mit einem Belegmarker in eckigen Klammern: entweder [Beleg: <Quelle>] wenn du eine
  Quelle benennen kannst, oder [Beleg fehlt] wenn nicht. Kein Punkt ohne Marker.
- Unter "Rechtlicher Hinweis" mindestens zwei Saetze zur Abgrenzung von medizinischen Aussagen.

Gib ausschliesslich das Markdown aus.

md

What this does not show

Three tasks are three tasks. Everything here was measured with think disabled, which is a real constraint for the reasoning models in the lineup. Both hosts run the same Ollama version on the same GPU model — this says nothing about determinism across different Ollama versions, different quantisations, or different hardware. And structural scoring rewards a model that follows instructions; it cannot tell you whether the marketing copy is any good.

What it does show is narrow and, I think, worth having: on identical hardware with a model that fits, a local LLM is a reproducible component. You can put it in a pipeline and expect the same bytes tomorrow. Give it 800 MB too little VRAM and it becomes something you have to re-check every time.