The fastest local model builds the second-worst Tetris
A few hours ago I published a throughput table: nine local models, tokens per second, cleanly measured. Its closing paragraph said speed is one dimension of…
Blog
Infrastructure, AI, tools, and the decisions behind them.
A few hours ago I published a throughput table: nine local models, tokens per second, cleanly measured. Its closing paragraph said speed is one dimension of…
Nine local models, three prompt lengths, reasoning on and off, all at a 262144-token context on two RTX 3090s — and the model that has been…
I gave a 120-billion-parameter mixture-of-experts model a second GPU. It got five times slower. Here is the measurement, and the more useful thing I found by…
I had two RTX 3090s and the compute of roughly one. This is the rebuild that put both cards on native PCIe in a single box…
A benchmark ran every night for six nights, reported success every time, and measured nothing at all. The systemd unit exited cleanly. The result files were…
Notes from building an evaluation harness for entity linking against a hierarchical vocabulary Most published model comparisons answer a question I don't have. They tell me…
Why One-Shot Benchmarks Measure the Wrong Thing Most of the coding benchmarks that show up in model cards and leaderboards ask the same question: here’s a…
Thirty-six runs, four model families, four task designs — and three of my own explanations dying along the way I wanted a number. Can the models…
I work with several machines and several AI sessions at once. One on the server, one on the Windows box, occasionally a third somewhere in between.…
In July I published an article about determinism in local models. It contains this sentence: “All 198 outputs are available in full, so that you can…