Topic

AI

Everything we publish on AI: the models we run on our own machines, what the tools send home, the benchmarks we measure ourselves and the claims we check against the code. Start with the investigations, then the guides.

Why your local model stops at exactly 8,192 tokensbenchmarks
Guides

Why your local model stops at exactly 8,192 tokens

Twelve of thirty six local model runs stopped at exactly 8,192 tokens with a 16,384 cap. The LM Studio context…

6 min
The two Gemma 4 31B entries: how Reki and forge play ARC-AGI-3agents
Research

The two Gemma 4 31B entries: how Reki and forge play ARC-AGI-3

Second and third in the ARC-AGI-3 milestone both ran Gemma 4 31B on vLLM: four frames at eight times scale,…

6 min
Infinite play by eviction: how a 27B plays past its context windowagents
Research

Infinite play by eviction: how a 27B plays past its context window

Infinite play by eviction: how the winning ARC-AGI-3 harness keeps a 27B model playing past a 64K context by dropping…

5 min
The Duck harness, read line by line: how a 27B won an ARC-AGI-3 milestoneagents
Research

The Duck harness, read line by line: how a 27B won an ARC-AGI-3 milestone

The Duck harness won the first ARC-AGI-3 milestone with a 27B model and a Python REPL, scoring 1.21% official and…

6 min
Two agents on one LM Studio endpoint: what concurrent requests costagents
Guides

Two agents on one LM Studio endpoint: what concurrent requests cost

Two identical LM Studio concurrent requests to a 31B on a 48 GB Mac: each ran at half speed and…

5 min
A REPL agent in its smallest form: one model call writes the playeragents
Guides

A REPL agent in its smallest form: one model call writes the player

A REPL agent in its smallest form: one local model call wrote a 44 line policy that played 60,000 ARC-AGI-3…

5 min
Image vs text for a local vision model: one board, two costsGemma 4
Guides

Image vs text for a local vision model: one board, two costs

Image vs text for a local vision model, measured: the same 64x64 board cost 4,179 prompt tokens as hex and…

5 min
A local model agent on one ARC-AGI-3 level: 40 moves, 230 minutes, nothingagents
Guides

A local model agent on one ARC-AGI-3 level: 40 moves, 230 minutes, nothing

A local model agent played 40 moves of one ARC-AGI-3 level in 230 minutes, 154,673 reasoning tokens, and cleared nothing.…

5 min
ARC-AGI-3 on a Mac: what a 48 GB M4 Pro can and cannot doApple Silicon
Guides

ARC-AGI-3 on a Mac: what a 48 GB M4 Pro can and cannot do

ARC-AGI-3 on a Mac, measured on a 48 GB M4 Pro: the toolkit does 4,000 actions a second, a 31B…

5 min
Enter ARC Prize 2026 with a local model: the Kaggle track, from a Macarc prize
Guides

Enter ARC Prize 2026 with a local model: the Kaggle track, from a Mac

ARC Prize 2026 pays $700,000 for 100 percent; its first milestone went to local open weight models. The prizes, the…

5 min