Stridenalysis·Research · Analysis · Insights

We read it, so you don't have to.

Close readings of the documents, papers, and filings shaping the future, with the receipts.

10Research
14Analysis
12Insights
36Total pieces
Showing 36 pieces
The largest distillation attack yet: the Alibaba numbers№ 036 · Research
Research · AI security

The largest distillation attack yet: the Alibaba numbers

Somewhere between late April and early June, a machine somewhere opened a fresh conversation with Claude roughly every seventh of…

Jul 25 · 5 min · 2026
Bonsai 27B vs Gemma 4 31B on a Mac: which local model to run№ 035 · Analysis
Analysis · 1-bit LLM

Bonsai 27B vs Gemma 4 31B on a Mac: which local model to run

Two models sat loaded in the same copy of LM Studio this week, and one of them was a fifth…

Jul 25 · 6 min · 2026
Bonsai 27B 1-bit on a Mac: a 27B model in 3.8 GB№ 034 · Insights
Insights · 1-bit LLM

Bonsai 27B 1-bit on a Mac: a 27B model in 3.8 GB

Most 27-billion-parameter models arrive as a download you plan an evening around. The Bonsai 27B 1-bit build from PrismML lands…

Jul 24 · 7 min · 2026
How AI companies hide tracking code, from Claude Code to your local tools№ 033 · Research
Research · Note

How AI companies hide tracking code, from Claude Code to your local tools

AI tracking code hides in plain sight. We break down the covert tracker Anthropic shipped in Claude Code, the telemetry…

Jul 11 · 5 min · 2026
Graphify has 81,000 stars. We read the code instead of running it№ 032 · Analysis
Analysis · graphify

Graphify has 81,000 stars. We read the code instead of running it

Graphify has 81,000 GitHub stars in three months. We read the Graphify source instead of running it: zero telemetry, local-first…

Jul 11 · 7 min · 2026
Graphify on a local model: we ran it and watched the network№ 031 · Insights
Insights · graphify

Graphify on a local model: we ran it and watched the network

We ran Graphify on a local model and watched every socket it opened. 358 went to LM Studio, none went…

Jul 10 · 6 min · 2026
OpenCode vs OpenJarvis: which local AI agent phones home?№ 030 · Research
Research · local-ai

OpenCode vs OpenJarvis: which local AI agent phones home?

OpenCode vs OpenJarvis: we captured what each one sends home. One has a working off switch. The other, the agent…

Jul 08 · 6 min · 2026
OpenJarvis on a Mac: we retested Stanford’s local AI agent№ 029 · Analysis
Analysis · local-ai

OpenJarvis on a Mac: we retested Stanford’s local AI agent

We reinstalled OpenJarvis on a Mac, captured its telemetry with a local sink, and found the undocumented off switch. What…

Jul 08 · 9 min · 2026
Brand voice for AI agents: the layer that beats a better model№ 028 · Insights
Insights · coding-agents

Brand voice for AI agents: the layer that beats a better model

We built a five-file brand voice for AI agents. Why the output spec, not the model, is the layer that…

Jul 08 · 6 min · 2026
How exo Clusters Macs to Run a 671B Model Locally№ 027 · Research
Research · apple-silicon

How exo Clusters Macs to Run a 671B Model Locally

exo is open-source software that pools several Apple Silicon Macs into one machine, so a cluster can run models like…

Jul 08 · 5 min · 2026
Set up local LLM tool calling in Hermes and OpenCode on a Mac№ 026 · Analysis
Analysis · Gemma 4

Set up local LLM tool calling in Hermes and OpenCode on a Mac

We set up local LLM tool calling in Hermes and OpenCode on a 48 GB Mac mini. Why small models…

Jun 25 · 5 min · 2026
M4 Pro vs M5 Max for local inference on Apple silicon№ 025 · Insights
Insights · apple-silicon

M4 Pro vs M5 Max for local inference on Apple silicon

M4 Pro vs M5 Max for local inference. Memory bandwidth, unified memory ceilings and time to first token, and whether…

Jun 21 · 7 min · 2026
Local vs cloud AI: what to run where in 2026№ 024 · Research
Research · cloud-ai

Local vs cloud AI: what to run where in 2026

Local vs cloud AI in 2026. A decision guide for what to run on your own hardware, what to send…

Jun 21 · 7 min · 2026
Gemma 4 E4B on edge hardware: small models catch up№ 023 · Analysis
Analysis · edge

Gemma 4 E4B on edge hardware: small models catch up

Gemma 4 E4B runs on 8 GB of unified memory. What a small model can now do on edge hardware,…

Jun 21 · 6 min · 2026
What Ollama NVFP4 means for local model quality№ 022 · Insights
Insights · local-ai

What Ollama NVFP4 means for local model quality

Ollama NVFP4 and q4_K_M are nearly the same size on disk. What the newer 4-bit format changes for local model…

Jun 21 · 6 min · 2026
Gemma 4 26B MoE local: quality per gigabyte on unified memory№ 021 · Research
Research · gemma

Gemma 4 26B MoE local: quality per gigabyte on unified memory

Gemma 4 26B MoE runs from 16 GB of weights. We look at quality per gigabyte on unified memory, and…

Jun 21 · 8 min · 2026
Gemma 4 31B vs Qwen 3.5 27B on unified memory№ 020 · Analysis
Analysis · gemma

Gemma 4 31B vs Qwen 3.5 27B on unified memory

Gemma 4 31B vs Qwen 3.5 27B on unified memory. Which fits, which is faster, and which answers better when…

Jun 21 · 7 min · 2026
MLX vs llama.cpp vs MetalRT: best inference engine on Apple Silicon in 2026№ 019 · Insights
Insights · apple-silicon

MLX vs llama.cpp vs MetalRT: best inference engine on Apple Silicon in 2026

MLX vs llama.cpp vs MetalRT: one quantized model, three engines, one Mac. Generation speed, memory and setup, measured rather than…

Jun 21 · 7 min · 2026
LM Studio vs Ollama on Mac: which one to use in 2026№ 018 · Research
Research · lm-studio

LM Studio vs Ollama on Mac: which one to use in 2026

LM Studio vs Ollama on a Mac in 2026. Setup, speed, model management and the API each one exposes, plus…

Jun 21 · 7 min · 2026
Local LLM benchmarks on M4 Pro: Gemma, Qwen, Llama speeds№ 017 · Analysis
Analysis · benchmarks

Local LLM benchmarks on M4 Pro: Gemma, Qwen, Llama speeds

Local LLM benchmarks on an M4 Pro. Gemma, Qwen and Llama measured on tokens per second, memory used, and the…

Jun 21 · 8 min · 2026
What the Ollama MLX shift means for local AI on Mac№ 016 · Insights
Insights · local-ai

What the Ollama MLX shift means for local AI on Mac

Ollama is moving toward MLX on Apple Silicon. What the shift changes for speed, memory and model choice, and whether…

Jun 21 · 7 min · 2026
Best Local LLMs for Coding on a Mac in 2026: Benchmarked and Ranked№ 015 · Research
Research · benchmarks

Best Local LLMs for Coding on a Mac in 2026: Benchmarked and Ranked

We benchmarked and ranked the best local LLMs for coding on a Mac in 2026, on real refactors, with speed,…

Jun 21 · 6 min · 2026
Privacy-first coding: why you should run AI agents entirely offline№ 014 · Analysis
Analysis · cloud-ai

Privacy-first coding: why you should run AI agents entirely offline

Your config files hold keys and endpoints. Here is why we run AI agents offline, what a cloud agent can…

Jun 21 · 8 min · 2026
How to choose between Ollama, LM Studio, and MLX for local models№ 013 · Insights
Insights · lm-studio

How to choose between Ollama, LM Studio, and MLX for local models

Three ways to serve a local model on a Mac. We compare Ollama, LM Studio and MLX on setup, speed…

Jun 21 · 7 min · 2026
The Hidden Costs of ‘Free’ Cloud AI Tiers№ 012 · Analysis
Analysis · Analysis

The Hidden Costs of ‘Free’ Cloud AI Tiers

A free cloud AI tier does not cost zero. It costs something other than dollars: your data, your reliability, your…

Jun 07 · 9 min · 2026
Agentic AI explained: what it changed for local AI№ 011 · Insights
Insights · Insights

Agentic AI explained: what it changed for local AI

Agentic AI explained without the hype. Stripped down it is one real shift: a model that answers questions became a…

Jun 07 · 10 min · 2026
Local Deep Research vs Perplexity and the Cloud№ 010 · Insights
Insights · Insights

Local Deep Research vs Perplexity and the Cloud

Local deep research vs Perplexity: a research engine on your own machine keeps every query private and charges nothing per…

Jun 07 · 10 min · 2026
Open WebUI vs Jan: The Local ChatGPT, Picked№ 009 · Insights
Insights · Insights

Open WebUI vs Jan: The Local ChatGPT, Picked

Open WebUI vs Jan: two ways to put a private ChatGPT in front of someone. A feature-rich browser workspace, or…

Jun 07 · 9 min · 2026
Do you need an agent framework? smolagents, and when to skip it№ 008 · Analysis
Analysis · Analysis

Do you need an agent framework? smolagents, and when to skip it

Do you need an agent framework at all? smolagents vs the alternatives, and why most of the noise is selling…

Jun 07 · 10 min · 2026
Is Local AI Good Enough to Replace the Paid Tools?№ 007 · Insights
Insights · Insights

Is Local AI Good Enough to Replace the Paid Tools?

Is local AI good enough to replace the paid tools? For most of what you do in a day, yes.…

Jun 07 · 10 min · 2026
Which Local LLM Should You Actually Run in 2026?№ 006 · Research
Research · Research

Which Local LLM Should You Actually Run in 2026?

The best local LLM in 2026 is not the biggest one. It is the model that fits your RAM with…

Jun 07 · 11 min · 2026
Cloud AI vs local: the true cost compared№ 005 · Analysis
Analysis · Analysis

Cloud AI vs local: the true cost compared

Stop arguing about whether local AI is cheaper. Add up the real cost of AI subscriptions vs local: a machine…

Jun 07 · 10 min · 2026
Ollama vs MLX vs Jan: Running Local Models on a Mac№ 004 · Research
Research · Research

Ollama vs MLX vs Jan: Running Local Models on a Mac

Three ways to run a model on Apple Silicon: the foundation, the fast one, and the friendly one. We run…

Jun 07 · 11 min · 2026
Aider vs Cline vs OpenCode vs OpenHands, Picked for You№ 003 · Analysis
Analysis · Analysis

Aider vs Cline vs OpenCode vs OpenHands, Picked for You

Four open-source coding agents, four different answers to one question: where do you want to work? Terminal, editor, app window,…

Jun 07 · 12 min · 2026
OpenCode vs Pi vs Cursor: Which Coding Agent Do You Need?№ 002 · Analysis
Analysis · Analysis

OpenCode vs Pi vs Cursor: Which Coding Agent Do You Need?

Three coding agents, three different bets on where you work and where your code goes. We have run all three.…

Jun 07 · 12 min · 2026
Best local model for OpenCode, compared№ 001 · Analysis
Analysis · Analysis

Best local model for OpenCode, compared

Nemotron vs DeepSeek for local coding. DeepSeek scores higher on pure coding benchmarks, and we still run Nemotron as the…

Jun 07 · 11 min · 2026