AI
Everything we publish on AI: the models we run on our own machines, what the tools send home, the benchmarks we measure ourselves and the claims we check against the code. Start with the investigations, then the guides.
arc-agi-3ARC-AGI-3 scores: under 1% is six months out of date
ARC-AGI-3 scores from 0.51% at launch to 62.7% shared harness and 99.9% vendor harness in September 2026, in date order…
arc-agi-2ARC-AGI-2 vs ARC-AGI-3: what changed, and why the numbers do not transfer
ARC-AGI-2 vs ARC-AGI-3: static puzzles scored on accuracy against interactive games scored on action efficiency. Why a 24% and a…
apiReasoning tokens are output tokens, and five APIs count them differently
Reasoning tokens are output tokens on LM Studio, OpenAI, Anthropic, Gemini and OpenRouter, and each names and caps them differently.…
LM StudioHow to set max_tokens for a reasoning model, from measurement
max_tokens for a reasoning model is a measurement, not a guess: run the largest input once at a huge cap,…
Gemma 4LM Studio reasoning effort, temperature and schemas: what changes the thinking
LM Studio reasoning effort, temperature, JSON schemas and tool calls measured on one board: low effort thought most, a schema…
LM StudioDebug a local model with a four call control ladder
Four calls, from a trivial reply to the real input, to debug a local model in minutes. What each rung…
LM StudioGrid to text for a local model: four lines of numpy, measured
Grid to text for a local model, measured: raw hex costs 442.7s a move, a four line numpy colour summary…
arc-agi-3Score an ARC-AGI-3 run locally: the scorecard without the API
The ARC-AGI-3 scorecard is computed locally and offline by the package that plays the game. What get_scorecard returns, how it…
arc-agi-3ARC-AGI-3 recordings: what save_recording writes, and how big it gets
ARC-AGI-3 recordings are one argument to make(). What each JSONL line holds, why frames make the file forty times bigger,…
arc-agi-3The ARC-AGI-3 human baseline: all 183 levels, read from disk
The ARC-AGI-3 human baseline is the upper median first-time player, not the best. All 183 baselines read from disk: 6…