Topic

AI

Everything we publish on AI: the models we run on our own machines, what the tools send home, the benchmarks we measure ourselves and the claims we check against the code. Start with the investigations, then the guides.

ARC-AGI-3 scores: under 1% is six months out of datearc-agi-3
Research

ARC-AGI-3 scores: under 1% is six months out of date

ARC-AGI-3 scores from 0.51% at launch to 62.7% shared harness and 99.9% vendor harness in September 2026, in date order…

5 min
ARC-AGI-2 vs ARC-AGI-3: what changed, and why the numbers do not transferarc-agi-2
Research

ARC-AGI-2 vs ARC-AGI-3: what changed, and why the numbers do not transfer

ARC-AGI-2 vs ARC-AGI-3: static puzzles scored on accuracy against interactive games scored on action efficiency. Why a 24% and a…

6 min
Reasoning tokens are output tokens, and five APIs count them differentlyapi
Research

Reasoning tokens are output tokens, and five APIs count them differently

Reasoning tokens are output tokens on LM Studio, OpenAI, Anthropic, Gemini and OpenRouter, and each names and caps them differently.…

5 min
How to set max_tokens for a reasoning model, from measurementLM Studio
Guides

How to set max_tokens for a reasoning model, from measurement

max_tokens for a reasoning model is a measurement, not a guess: run the largest input once at a huge cap,…

5 min
LM Studio reasoning effort, temperature and schemas: what changes the thinkingGemma 4
Guides

LM Studio reasoning effort, temperature and schemas: what changes the thinking

LM Studio reasoning effort, temperature, JSON schemas and tool calls measured on one board: low effort thought most, a schema…

6 min
Debug a local model with a four call control ladderLM Studio
Guides

Debug a local model with a four call control ladder

Four calls, from a trivial reply to the real input, to debug a local model in minutes. What each rung…

6 min
Grid to text for a local model: four lines of numpy, measuredLM Studio
Guides

Grid to text for a local model: four lines of numpy, measured

Grid to text for a local model, measured: raw hex costs 442.7s a move, a four line numpy colour summary…

5 min
Score an ARC-AGI-3 run locally: the scorecard without the APIarc-agi-3
Guides

Score an ARC-AGI-3 run locally: the scorecard without the API

The ARC-AGI-3 scorecard is computed locally and offline by the package that plays the game. What get_scorecard returns, how it…

5 min
ARC-AGI-3 recordings: what save_recording writes, and how big it getsarc-agi-3
Guides

ARC-AGI-3 recordings: what save_recording writes, and how big it gets

ARC-AGI-3 recordings are one argument to make(). What each JSONL line holds, why frames make the file forty times bigger,…

5 min
The ARC-AGI-3 human baseline: all 183 levels, read from diskarc-agi-3
Research

The ARC-AGI-3 human baseline: all 183 levels, read from disk

The ARC-AGI-3 human baseline is the upper median first-time player, not the best. All 183 baselines read from disk: 6…

5 min