Topic

AI

Everything we publish on AI: the models we run on our own machines, what the tools send home, the benchmarks we measure ourselves and the claims we check against the code. Start with the investigations, then the guides.

Ollama’s MLX engine has no 32GB requirement: the rule is three lines of GoNote
Research

Ollama’s MLX engine has no 32GB requirement: the rule is three lines of Go

The Ollama MLX 32GB requirement is not in the source at all. Routing is decided by model format, not by…

7 min
Eight GGUF builds of one model: the filename does not tell you what is insideNote
Research

Eight GGUF builds of one model: the filename does not tell you what is inside

GGUF quant naming, audited across eight builds of one model. A file called UD-Q4_K_XL holds no Q4_K, and none had…

9 min
Run an NVFP4 model in Ollama on a Mac: tags, commands, gotchasNote
Guides

Run an NVFP4 model in Ollama on a Mac: tags, commands, gotchas

How to run an Ollama NVFP4 model on a Mac: the tags that exist, the version you need, and four…

6 min
Bonsai 27B in a coding agent: a 1-bit model that calls tools cleanly and misses the line numberNote
Guides

Bonsai 27B in a coding agent: a 1-bit model that calls tools cleanly and misses the line number

We put Bonsai 27B, a 1-bit model in 3.8 GB, behind the Pi coding agent. Tool calls came back clean…

8 min
Ollama MTP: the speculative decoding that is already on, and the DRAFT command nobody documentedNote
Guides

Ollama MTP: the speculative decoding that is already on, and the DRAFT command nobody documented

Ollama MTP has been on by default since v0.31.1, roughly 90% faster on Gemma 4. What multi-token prediction does, and…

8 min
MoE on a Mac: 4 to 5 times faster, and bigger in memory than the dense model it beatsNote
Guides

MoE on a Mac: 4 to 5 times faster, and bigger in memory than the dense model it beats

A 26B MoE decodes 4 to 5 times faster than a dense model on a Mac and can still take…

10 min
NVFP4 vs MXFP4: the same 4 bits, 36% apartNote
Research

NVFP4 vs MXFP4: the same 4 bits, 36% apart

MXFP4 vs NVFP4 comes down to two choices: block size and scale format. Same 4-bit element, and NVIDIA reports MXFP4…

6 min
NVFP4 vs Q4_K_M: the vendor published no numbers, and the independent ones favour Q4_K_MNote
Research

NVFP4 vs Q4_K_M: the vendor published no numbers, and the independent ones favour Q4_K_M

NVFP4 vs Q4_K_M: the vendor published no numbers to check, and the one independent comparison has Q4_K_M ahead on KL…

9 min
Does llama.cpp support MLX? No, and a maintainer explained why in 2023Note
Research

Does llama.cpp support MLX? No, and a maintainer explained why in 2023

Does llama.cpp support MLX? No, and a maintainer said why in 2023. What Apple Silicon actually runs on, what Ollama…

7 min
NVFP4 on Apple Silicon: an NVIDIA format that only runs on a MacNote
Research

NVFP4 on Apple Silicon: an NVIDIA format that only runs on a Mac

NVFP4 on Apple Silicon is an NVIDIA format Ollama will only serve to macOS. Why the 412 error exists, what…

8 min