AI & Tools

Think Step by Step? On One Task in Four It Makes Things Worse

Grouped bar chart of fully solved cells out of 33 per task, bare prompt against the step-by-step instruction. rechnung rises from 19 to 24, extrakt from 15 to 25, while plan falls from 32 to 25. A side panel lists the cost of the step-by-step variant: median 10.2 seconds against 6.9 (+49 per cent), 769 against 707 tokens (+9 per cent), 11 truncated cells against 3 for the bare prompt, format compliance 115 of 132.
The same instruction adds ten solved cells on extrakt and costs seven on plan.

Three pieces of advice appear in every prompting guide: tell the model to think step by step. Give it a role. Show it an example. All three sound sensible, all three cost typing, and none of them had ever been tested on my hardware. So I tested them — four tasks, four prompt variants, eleven models, 528 measured cells on a single machine.

The result in one sentence: the tricks help substantially — except where the model already solves the task. There they hurt.

What was measured

Four tasks with an unambiguously correct answer rather than a good-looking one: rechnung (a multi-step price calculation), plan (a small assignment problem), extrakt (six fields as JSON, including a date calculation across a month boundary) and raetsel (five people mapped to seat, project and drink, eleven constraints). I confirmed the puzzle has exactly one solution by enumerating all 1,728,000 assignments; every one of the eleven constraints is load-bearing.

Against that, four variants of the same prompt — identical core, identical format block: nackt (the bare task), schritte (“think step by step”), rolle (“you are an experienced …”) and beispiel (one worked example up front). Eleven models, three runs each. 528 requests, zero errors.

Why the format block demands a closing ERGEBNIS: line instead of forbidding explanation: banning explanatory text would have broken schritte before it started. This way every variant may write as much as it likes beforehand and is scored the same.

The rule was fixed before the data

So that I could not pick a convenient reading afterwards: “the trick helps” means the number of fully solved cells rises by at least 3 out of 33 against nackt, on the same task. Below that it is individual models tipping over, and it gets reported as “no effect”. Scoring is per task, never as an overall mean — the four tasks have different ceiling effects and an average would smear them together. A half-correct invoice total is a wrong invoice total.

Results per task

Fully solved cells, out of 33 each:

Tasknacktschritterollebeispiel
rechnung19242421
plan32252727
extrakt15252123
raetsel11141414

On rechnung, schritte and rolle both clear the threshold at +5 while beispiel stays under it at +2. On extrakt all three clear it comfortably, schritte most of all at +10. On raetsel all three land at exactly +3 — clearing the bar precisely, though the absolute level there is low enough that I would not read much into it.

Where it flips

And then there is plan. Bare, the models solve 32 of 33 cells — essentially everything. With “think step by step” they drop to 25. With a role, 27. With an example, 27.

All three pieces of advice break the task that ran best without them. This is the most interesting finding in the run, and an uncomfortable one, because it does not refute the advice — it bounds it. A trick that helps a model reason gets in the way when the model does not need to reason. plan is small enough to be seen directly. Forcing the model to spell out the path gives it seven extra opportunities to get it wrong.

I read the diverging cells individually rather than trusting the number. The losses are genuine assignment errors, not misplaced answer markers.

The price

A trick that buys nothing and costs nothing is a different message from one that buys nothing and stretches the clock by half. Medians across all cells:

VariantTokensvs. nacktSecondsvs. nackt
nackt707—6.9—
schritte769+9 %10.2+49 %
rolle711±0 %8.7+25 %
beispiel662−6 %7.3+6 %

The striking row is schritte: nine per cent more tokens, but forty-nine per cent more wall-clock time. The gap is not in answer length. It is that longer outputs hit the token cap more often, and a capped run keeps generating to the limit — schritte has eleven truncated cells against three for nackt.

beispiel is the quiet winner of the cost table: six per cent fewer tokens than bare, only six per cent more time — and still +8 on extrakt. Putting one worked example up front is the only one of the three tricks that helps meaningfully in this run without costing meaningfully.

Format compliance

Counted separately, because it is a formatting failure and not a reasoning failure: how often did the result actually land on the last line, as demanded? beispiel 123 of 132, nackt 120, rolle 117, schritte 115. Get a model talking and you get more text after the answer.

A mistake of mine belongs here too. My first checker did not recognise ### ERGEBNIS: … and would have scored correctly answering models at zero per cent. I only caught it because I read the outputs instead of the scores. The same class of failure showed up twice more later — once a missing stop token, once a single mistyped letter in the answer marker that scored six fully correct cells as zero. Any suspiciously round zero is suspect, an exact zero-out-of-N most of all.

What this does not show

  • Four tasks are four tasks. They are short, mechanically checkable, and cover no open-ended writing.
  • Variants ran sequentially within a model; an ordering effect is not excluded, only made unlikely.
  • The three runs per cell are not seed variation. Measured afterwards, ornith:35b and qwen3.6:27b return byte-identical output across all three seeds, while qwq:32b answers differently from an identical seed. Three runs catch variance where variance exists — they do not test the effect of the seed.
  • Measured with think=false. For qwq:32b and deepseek-r1:14b it is now established that this switch does not control whether the model reasons, so the caveat “measured with thinking mode off” does not apply to every model in this table.
  • laguna-xs-2.1 runs here in its repaired form with a stop token added. Without it all 48 of its cells would have run to the token cap, and its row would have been a statement about packaging rather than about the model.

The hardware, and the rules every measurement on this site follows, are on How I Measure.

What I take from it

For tasks a model visibly has to work through — extraction with intermediate steps, multi-stage arithmetic — “think step by step” is the strongest of the three instructions and worth its half minute. For tasks the model sees directly, it is a step backwards. And putting one worked example up front costs almost nothing and helps a little almost everywhere.

What I will not do again: switch on a prompting trick by default because a guide recommends it. On plan that would have cost me seven cells out of thirty-three.