AI
Everything we publish on AI: the models we run on our own machines, what the tools send home, the benchmarks we measure ourselves and the claims we check against the code. Start with the investigations, then the guides.
apple-siliconM4 Pro vs M5 Max for local inference on Apple silicon
M4 Pro vs M5 Max for local inference. Memory bandwidth, unified memory ceilings and time to first token, and whether…
cloud-aiLocal vs cloud AI: what to run where in 2026
Local vs cloud AI in 2026. A decision guide for what to run on your own hardware, what to send…
coding-agentsLocal AI coding agent setup with OpenCode and Ollama
Set up a local AI coding agent with OpenCode and Ollama. No API key, no subscription. The install, the model…
edgeGemma 4 E4B on edge hardware: small models catch up
Gemma 4 E4B runs on 8 GB of unified memory. What a small model can now do on edge hardware,…
local-aiWhat Ollama NVFP4 means for local model quality
Ollama NVFP4 and q4_K_M are nearly the same size on disk. What the newer 4-bit format changes for local model…
gemmaGemma 4 26B MoE local: quality per gigabyte on unified memory
Gemma 4 26B MoE runs from 16 GB of weights. We look at quality per gigabyte on unified memory, and…
gemmaGemma 4 31B vs Qwen 3.5 27B on unified memory
Gemma 4 31B vs Qwen 3.5 27B on unified memory. Which fits, which is faster, and which answers better when…
MLX vs llama.cpp vs MetalRT: best inference engine on Apple Silicon in 2026
MLX vs llama.cpp vs MetalRT: one quantized model, three engines, one Mac. Generation speed, memory and setup, measured rather than…
lm-studioLM Studio vs Ollama on Mac: which one to use in 2026
LM Studio vs Ollama on a Mac in 2026. Setup, speed, model management and the API each one exposes, plus…
benchmarksLocal LLM benchmarks on M4 Pro: Gemma, Qwen, Llama speeds
Local LLM benchmarks on an M4 Pro. Gemma, Qwen and Llama measured on tokens per second, memory used, and the…