The challenger with the better argument lost by half
Twelve models, one structured-extraction task, and a challenger with the better argument on paper The companion piece to this one reports a zero: local models, driven…
Blog
Infrastructure, AI, tools, and the decisions behind them.
Twelve models, one structured-extraction task, and a challenger with the better argument on paper The companion piece to this one reports a zero: local models, driven…
Yesterday this blog reported that the test set had hit its ceiling: one model solved all 48 cells, and about a model that solves everything nothing…
Part 8 closed with a promise: next I would measure what abliteration actually takes away from a model. The advice itself is everywhere — take the…
Until today, every measurement in this series ran on a single graphics card. 24 gigabytes, and the boundary was clear: anything larger simply went unmeasured. As…
For seven installments this series has measured what local models answer when you ask them. This one is about the question that comes before that, and…
Models that “can think” have a switch for it in Ollama. This series has had it off from the start and said so in every post:…
Every measurement in this series runs at temperature 0 with three seeds. The reasoning was simple: if the model does drift, three passes will catch it.…
A 2020 desktop build turns into a Proxmox node with 48 GB of VRAM. The interesting part is the move itself: a guest migration that reported…
An agent is not a chat. The question is not whether the model sounds clever, but whether it reaches for the right tool — and whether…
Three pieces of advice appear in every prompting guide: tell the model to think step by step. Give it a role. Show it an example. All…