Filter All Philosophy Fringe Tech Rant Essay Things I Love
Laguna-XS runs beautifully on my 5090, and I am keeping the model everyone calls outdated
A newer, well-recommended coding model fits my 32GB GPU at max quality and runs faster than the one I keep. I kept the old one anyway, and the numbers say I was right to.
7 min read
Qwen3.6-27B Scores Higher on BFCL, Ties on My Workload, and Runs 4.9x Slower
Qwen3.6-27B scored higher on BFCL. Both models scored 58/59 on the real workload. One answers in 0.87s, the other in 4.27s.
10 min read
GLM-4.7-Flash on One Consumer GPU: Why It Runs My Homelab Agent
A ~31B Mixture-of-Experts model at 4-bit, on one 32GB card, that beat a dense 32B on the real BFCL function-calling benchmark. What it is, what it is good for, and the numbers.
8 min read
A Private Coding LLM on My Own Network.
I run Qwen3-Coder on a spare RTX 5090 and point my coding agent at it over the LAN. Private, fast, no per-token bill, secured with an API key, and watched on a dashboard. Here is the whole build.
14 min read