Local LLM

Grouped bars compare points per task before and after the stop token was added, in percent: rechnung falls from 88 to 75, extrakt rises from 0 to 100, raetsel from 30 to 39. A side panel carries the totals across all 48 cells: 117 of 336 points before against 215 of 336 after, cells at the token ceiling 48 of 48 against 3 of 48, format compliance 31 against 83 percent, 192,000 against 45,775 generated tokens, 1594 against 326 seconds of compute time, and the plan task unchanged at 88 percent.

The model that couldn’t stop

One model sits at the bottom of the table with 35 percent. The number was measured honestly — and is still wrong, in both directions at…