Filter All Philosophy Fringe Tech Rant Essay Things I Love
Do LLMs Have Qualia? A Map of the Evidence as of October 2026
Every major result on experience in language models, 2021 to October 2026, with what each one measured and what it did not.
17 min read
Prompt Injection Works Because the Model Can't Tell Who Is Talking
New research traces prompt injection to a single mechanism: LLMs decide who is speaking from writing style, not from role tags. Forged reasoning takes attacks from near-zero to 60% success. Remove the style, and it collapses to 10%.
5 min read
NVIDIA Is Reportedly Buying Hugging Face. What Would It Mean for Local AI?
One anonymous source says $12.9 billion. Both companies are silent. Here is what is actually known, what the communities fear and hope, and the questions worth asking before anyone predicts anything.
6 min read
Laguna-XS runs beautifully on my 5090, and I am keeping the model everyone calls outdated
A newer, well-recommended coding model fits my 32GB GPU at max quality and runs faster than the one I keep. I kept the old one anyway, and the numbers say I was right to.
7 min read
Qwen3.6-27B Scores Higher on BFCL, Ties on My Workload, and Runs 4.9x Slower
Qwen3.6-27B scored higher on BFCL. Both models scored 58/59 on the real workload. One answers in 0.87s, the other in 4.27s.
10 min read
GLM-4.7-Flash on One Consumer GPU: Why It Runs My Homelab Agent
A ~31B Mixture-of-Experts model at 4-bit, on one 32GB card, that beat a dense 32B on the real BFCL function-calling benchmark. What it is, what it is good for, and the numbers.
8 min read
A Private Coding LLM on My Own Network.
I run Qwen3-Coder on a spare RTX 5090 and point my coding agent at it over the LAN. Private, fast, no per-token bill, secured with an API key, and watched on a dashboard. Here is the whole build.
14 min read