Blog

Notes on building

Infrastructure, AI, tools, and the decisions behind them.

Four task cards — a multi-stage price calculation, four talks into four slots, six fields into JSON, and a five-seat constraint puzzle — above a full-width bar filled completely and labelled 48 of 48. Result badges note that all 48 cells were correct, that the winning model was nine times faster and half the size, and a panel explains that a model solving everything tells you only that it did not fail.

48 out of 48 — And Why That Is a Problem

Until today, every measurement in this series ran on a single graphics card. 24 gigabytes, and the boundary was clear: anything larger simply went unmeasured. As…

Split dark panel. Left, under 'What I had written down': 65 of 176 groups identical, 111 groups deviate, the run log only reported STREUT, three seeds were meant to catch it. Right, under 'What the data actually held': 0 groups in which all three outputs differ, in 111 of 111 cases the odd one out is seed 42, seeds 43 and 44 never part, the outlier is always the first pass. A strip below covers the two-machine checks: five different seeds on six models yield the same text every time, 108 cell pairs across two machines show not one difference, token counts 503, 508, 508, 508, 508. Tags: temperature 0, Ollama 0.32.5, RTX 3090 on two machines, 528 cells, 11 models, byte-wise comparison.

The Seed Does Nothing. The First Run Does

Every measurement in this series runs at temperature 0 with three seeds. The reasoning was simple: if the model does drift, three passes will catch it.…