Terminal coding agent harness routing requests between planning and execution model nodes

Pi 1.0: The Minimal Coding-Agent Harness That Routes Planning to Opus and Execution to GPT

Pi 1.0 shipped on October 1, 2026, and it is the most interesting coding-agent release of the year precisely because it added features by refusing to add most of them. Earendil — the company that acquired the Pi project in April 2026 — tagged the first stable version of a terminal agent that Mario Zechner and Armin Ronacher built around four tools: read to read files, bash to execute commands, edit to modify files, and write to create them. Everything else in 1.0 — MCP integration, model routing, deferred tool loading, cache warming — is layered on that core, and the release notes describe the additions as features that survived months of testing, not features that were requested.

Pi is MIT licensed, installs on macOS, Linux, and Windows, and Earendil reports hundreds of thousands of weekly users. The repository sits near 113,000 stars as of early October 2026, up from roughly 54,000 in May. That curve is the reason a minimal harness matters commercially: the agents that dominate the market carry system prompts in the 7,000–10,000 token range, and Pi’s is under 1,000.

What Pi 1.0 Actually Ships

The 1.0 release folds in eight changes:

  • Codemode — a harness-side sandbox that helps agents issue tool calls, with native support for MCP, decision models, and image models
  • Native MCP support — Model Context Protocol is now a standard feature, not an extension
  • Virtual models — an extension that composes multiple physical models into a single addressable model
  • Deferred tool loading — tool definitions load when a capability is needed, so the prompt does not grow with every tool you add
  • Cache warming for Anthropic models — pre-warms prompt caching to reduce waiting time
  • Mid-conversation system messages — system prompt and available tools change during a session, transcript-aware
  • New terminal UI theme
  • Full-screen mode by default

Alongside Pi 1.0, Earendil released Pi Durable, an experimental separate package aimed at long-running agents. Regular Pi is a terminal coding session; Pi Durable adds the layer that orchestrates multiple conversations and handles the storage backend, using SQLite and JSONL so a task can resume after the process exits. It installs via npm as @earendil-works/pi-durable with @earendil-works/pi-ai and @earendil-works/chord.

The Four-Tool Design Is the Cost Story

Pi’s original pitch was that a coding agent needs very few primitives. read, bash, edit, write covers the entire loop a developer performs inside a repository, and a harness that exposes only those four tools sends a much smaller prompt to the model on every turn. Independent comparisons put Pi’s system prompt under 1,000 tokens against roughly 7,000–10,000 for Cline, OpenCode, and Claude Code. At agent-loop volumes — thousands of turns per task — that difference is the bill.

The cost profile is not just marketing. NVIDIA researchers published SoL-Pi in September 2026: four harness mechanisms discovered by an AI running auto-research loops across 535 environments, applied to the open-source Pi agent. On EdgeBench the mechanisms cut token traffic by 44.7% to 49.0% and API cost by roughly 33%, while retaining about 94% of Pi’s score on GPT-5.6 Sol and Opus 5. The point is structural: harness overhead, not model quality, is where most agent spend leaks.

Deferred tool loading attacks the same leak from the other side. A harness that presents every tool definition in every request pays for those tokens on every request, including the ones where the tool is never used. Loading tools on demand keeps the fixed cost flat as you add integrations.

Codemode, MCP, and the Protocol Reversal

The headline change in 1.0 is that Pi adopted the Model Context Protocol, which Pi creator Mario Zechner had publicly dismissed as unnecessary. Earendil’s explanation is that MCP improved, and that the team’s own changes to the MCP implementation made integration practical — the same changes also made it easier to use Jev, the decision model that has drawn a wave of imitators, inside Pi. The company’s framing: what Pi needs is quite similar to what MCP needs, a sandbox to play with in the form of an interpreter. Codemode is that sandbox.

Practically, Codemode means Pi can call MCP servers, decision models, and image models through one harness-side mechanism rather than a pile of bespoke adapters. OpenAI, Google, and Microsoft back MCP, so a Pi that speaks MCP plugs into the tool ecosystem those vendors have built — databases, browsers, design tools, internal APIs — without custom glue per integration.

Virtual Models: Splitting Planning From Execution

Pi supports more than 15 providers, including Anthropic, OpenAI, Google, Amazon Bedrock, Azure, Mistral, xAI, Hugging Face, OpenRouter, and Ollama, and you can switch models mid-conversation. The new virtual models extension goes further: it composes several physical models and exposes them as one model, distributing requests to the right one per turn.

Earendil’s demo shows the pattern. Pi builds a mechanism where Claude Opus handles planning, and when the Jev decision model classifies the work as having moved into the implementation phase, the harness hands processing to a GPT model. Planning at a frontier price, execution at a cheaper price, with a decision model doing the routing. For anyone running agents at volume, that is the single most valuable feature in the release: routing is a harness concern, not a model concern.

Running Pi Against Local Models

Pi’s local-model path is the reason this release matters for readers who run their own inference. Pi integrates directly with the llama.cpp router, which discovers GGUF files in a directory and loads models on demand. Start the router without a model flag so it runs in router mode:

llama-server 
  --models-dir ~/models 
  --no-models-autoload 
  --jinja 
  --host 127.0.0.1 
  --port 8080 
  -ngl 999 
  -c 32768

Then in Pi, run /login llama.cpp to store the connection, /llama to load a model from the discovered set (including a Hugging Face download search), and /model to select the loaded model for the session. The --jinja flag enables compatible chat templates and tool calling, which is what Pi needs for MCP and Codemode to work against a local model. -ngl 999 offloads as many layers to the GPU as fit; -c 32768 caps context per loaded model, which matters when several models share one router.

For Ollama, LM Studio, vLLM, SGLang, and llama-swap, Pi uses a compatible endpoint in ~/.pi/agent/models.json:

{
  "providers": {
    "ollama": {
      "baseUrl": "http://localhost:11434/v1",
      "api": "openai-completions",
      "apiKey": "ollama",
      "models": [{ "id": "qwen3-coder" }]
    }
  }
}

Ollama also ships a launcher that installs Pi, configures Ollama as the provider, and drops you into a session: ollama launch pi, or ollama launch pi --model qwen3-coder to run directly against a model. Set defaultProvider and defaultModel in ~/.pi/agent/settings.json to make a local model the default. A Hugging Face extension, pi-llama, auto-discovers models from a running llama.cpp server and registers them as a provider.

Four usage forms matter for automation: interactive terminal, JSON output, RPC, and a TypeScript SDK. That means Pi is callable from custom applications, not only from a terminal.

Hardware for Running Pi Locally

Local Pi is a memory problem before it is a model problem. A 7B–14B coder model at Q4 fits in 32GB of unified memory, a 24GB machine handles 7B comfortably with room for tool definitions, and a discrete GPU with 16GB of VRAM offloads most layers of a mid-size model. For the router workflow above, the machine needs enough memory to keep two or three quantized models resident, not one.

For a compact router box, the Beelink SER9 AI Mini PC carries a Ryzen AI 9 HX 370 with 32GB of LPDDR5X and a 1TB PCIe 4.0 SSD, listed at $609 against a $749 list price — enough memory to run a llama.cpp router with a Qwen3-coder quant resident, and small enough to sit on a desk as a dedicated agent host.

The GEEKOM A9 Max AI Boost is the same class of machine with more headroom: Ryzen AI 9 HX 370 rated at up to 80 TOPS with a dedicated 50 TOPS XDNA 2 NPU, 32GB of DDR5 and 2TB of SSD, and vendor documentation that names Ollama and ComfyUI as supported local workloads. The NPU path is worth noting for Pi specifically — deferred tool loading and virtual-model routing are exactly the kind of continuous background work an NPU is built for.

For a macOS host, the Mac mini with the M4 Pro chip, 24GB of unified memory and 512GB SSD is listed at $1,449.99 and is frequently discounted to around $1,279 with a promo code. Unified memory is the advantage: a single 24GB pool serves the router, the model, and the harness, and Apple’s MLX stack is one of the local backends Pi’s provider-agnostic config supports.

If you want the router to serve a larger model with long context, add GPU memory rather than more CPU cores. The ASUS Prime GeForce RTX 5070 Ti with 16GB of GDDR7 is an SFF-Ready, PCIe 5.0 card with a dual BIOS, and 16GB of VRAM is the tier where -ngl 999 offloads most of a 14B model at Q4 without spilling to system RAM.

What This Means

Pi 1.0 is a bet that the coding-agent market is a harness market. The model layer is commoditizing — 15+ providers, local GGUFs, peer-to-peer inference markets, open-weight MoE models shipping weekly — so the durable advantage sits in how a harness spends tokens. Under 1,000 tokens of system prompt, tools loaded on demand, a decision model routing planning to an expensive model and execution to a cheap one, and the same harness pointed at a llama.cpp router on a machine you own: that is a cost structure no cloud subscription matches.

Two practical takeaways. First, if you run local models, Pi is now the cleanest open-source harness for them, because the llama.cpp router integration handles discovery and on-demand loading natively and MCP works against a local model through Codemode. Second, install Pi Durable only when you need resumable long-running tasks — it is experimental, and the storage backend is a real dependency. For terminal coding, Pi 1.0 stable is the right starting point: curl -fsSL https://pi.dev/install.sh | sh on macOS/Linux, powershell -c "irm https://pi.dev/install.ps1 | iex" on Windows, or npm install -g @earendil-works/pi-coding-agent.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply

Your email address will not be published. Required fields are marked *