What generic benchmarks don’t tell you about local LLMs
Notes from building an evaluation harness for entity linking against a hierarchical vocabulary Most published model comparisons answer a question I don't have. They tell me…
Blog
Infrastructure, AI, tools, and the decisions behind them.
Notes from building an evaluation harness for entity linking against a hierarchical vocabulary Most published model comparisons answer a question I don't have. They tell me…
In July I published an article about determinism in local models. It contains this sentence: “All 198 outputs are available in full, so that you can…
Yesterday this blog reported that the test set had hit its ceiling: one model solved all 48 cells, and about a model that solves everything nothing…
Until today, every measurement in this series ran on a single graphics card. 24 gigabytes, and the boundary was clear: anything larger simply went unmeasured. As…
Models that “can think” have a switch for it in Ollama. This series has had it off from the start and said so in every post:…
Every measurement in this series runs at temperature 0 with three seeds. The reasoning was simple: if the model does drift, three passes will catch it.…
An agent is not a chat. The question is not whether the model sounds clever, but whether it reaches for the right tool — and whether…
Three pieces of advice appear in every prompting guide: tell the model to think step by step. Give it a role. Show it an example. All…
An abliterated model scored exactly zero out of 45 on the logic puzzle. "Abliteration destroys reasoning" would have made a good headline. It just was not…
Eleven local models, three real tasks, two RTX 3090 hosts, 198 requests, zero errors. 90 of 99 outputs came back byte-identical across machines — and the…