Topic

AI

Everything we publish on AI: the models we run on our own machines, what the tools send home, the benchmarks we measure ourselves and the claims we check against the code. Start with the investigations, then the guides.

RHAE explained: how ARC-AGI-3 scores an agentarc-agi-3
Research

RHAE explained: how ARC-AGI-3 scores an agent

RHAE is not accuracy. It scores your agent on moves against a first-time human, squared. Here is the formula, the…

7 min
LM Studio empty response: when reasoning eats the answerGemma 4
Guides

LM Studio empty response: when reasoning eats the answer

A local model returned nothing four times in a row. What an LM Studio empty response means, why a bigger…

8 min
ARC-AGI-3 local model test: what one move costsarc-agi-3
Guides

ARC-AGI-3 local model test: what one move costs

We ran the MIT licensed ARC-AGI-3 toolkit on a 48 GB M4 Pro with a local model, measured what a…

10 min
An agent’s system prompt tokens can outgrow the context window it runs in, by nearly five timesNote
Research

An agent’s system prompt tokens can outgrow the context window it runs in, by nearly five times

We counted the agent system prompt tokens in four real context files. Together they are 38,922 tokens, which is 475…

8 min
Run two models at once on 48 GB and they both fit, but the thing that breaks is the one nobody measuresNote
Research

Run two models at once on 48 GB and they both fit, but the thing that breaks is the one nobody measures

We run two models at once on a 48 GB Mac. RAM was never the limit: both fit in 26.78…

7 min
A 64K context window costs 4.4 GB of RAM on top of the weights, and your model manager never shows itNote
Research

A 64K context window costs 4.4 GB of RAM on top of the weights, and your model manager never shows it

Measured on a 48 GB Mac: context window ram usage for a local llm runs 80 KB per token. LM…

8 min
Prompt processing speed on an M4 Pro is 65 tokens a second, and your benchmark is timing the cache insteadNote
Research

Prompt processing speed on an M4 Pro is 65 tokens a second, and your benchmark is timing the cache instead

Measured on an M4 Pro: local llm prompt processing speed is 65 tokens a second. Repeat the same prompt and…

8 min
Two Unsloth UD quant names, one GGUF file, and not a single IQ1 or IQ2 tensor insideNote
Research

Two Unsloth UD quant names, one GGUF file, and not a single IQ1 or IQ2 tensor inside

An Unsloth UD quant name describes the recipe, not the GGUF tensors. Two Nemotron builds are byte-identical under different names,…

7 min
Nemotron 3.5 Lightning’s 4-bit GGUF is 12 percent 16-bit, and the weight is a head most people never switch onNote
Research

Nemotron 3.5 Lightning’s 4-bit GGUF is 12 percent 16-bit, and the weight is a head most people never switch on

The Nemotron 3.5 Lightning GGUF from NVIDIA is 22.46 GB, and 2.67 GB of it is BF16. That block is…

8 min
Your 4-bit vision model sees in 16-bit, and the best anyone offers is 8Note
Research

Your 4-bit vision model sees in 16-bit, and the best anyone offers is 8

An mmproj GGUF holds a whole vision encoder. Four publishers ship it unquantised at BF16 or F32, and the one…

7 min