The model that couldn’t stop
One model sits at the bottom of the table with 35 percent. The number was measured honestly — and is still wrong, in both directions at…
Blog
Infrastructure, AI, tools, and the decisions behind them.
One model sits at the bottom of the table with 35 percent. The number was measured honestly — and is still wrong, in both directions at…
Two RTX 3090s, eleven models with identical digests, 528 requests and no errors. Reserving a 128k KV cache is nearly free; filling it costs 10 to…
Ollama 0.32.5 was out, my workstation was still on 0.31.2, and I had just finished benchmarking 26 models on it. So I had a rare opportunity:…
I have an RTX 3090 sitting in a Windows workstation that is about to be torn down and rebuilt as a Proxmox AI host. Before that…
Two hosts that share a graphics card and almost nothing else, scoring that only counts what a script can verify, and the rules every number in…
The opening piece of a series in which I test what local language models actually do on hardware you can buy. This first part has no…
Venice.ai runs open-source models with zero-knowledge privacy. No prompt logging, no training on your data. The VVV token staking model gives you permanent AI access without…
How I use Tailscale to connect my homelab, data center servers, and everything in between into one encrypted mesh network. ACLs, MagicDNS, subnet routing, and why…
The hosting panel as a product category is about to be disrupted. Not by a better panel, but by AI agents that observe, reason, and act…
A compact Intel NUC running Proxmox, Ollama, and ComfyUI, connected to every server through Tailscale mesh. No cloud GPU bills, no vendor lock-in, unlimited inference for…