KI-Agenten

Pi cuts MCP token bloat by hiding tools behind Codemode sandbox

Pi 1.0 integrates MCP servers but keeps tool definitions out of the prompt, using a JavaScript sandbox to discover and execute tools only when needed.

A glass cube with gears represents efficient tool handling, while stacked papers symbolize reduced context bloat.
Für diesen Artikel generierte Illustration

Dieser Artikel ist nur auf Englisch verfügbar.

Pi, the coding agent created by Mario Zechner, has finally integrated the Model Context Protocol (MCP) into its core architecture after spending much of the past year excluding it. This shift arrives in Pi 1.0, following an acquisition by Earendil earlier this year, and introduces a novel method to handle the heavy context costs associated with standard MCP implementations.

The integration does not follow the traditional approach of exposing all server tools directly to the large language model. Instead, Pi uses an internal mechanism called Codemode to mediate access, ensuring that tool definitions remain hidden until they are explicitly required for a specific task.

What happened

For most of the last year, Zechner resisted adding MCP support to Pi despite the protocol becoming a standard across developer tools. His hesitation stemmed from concrete performance metrics. When he measured popular browser-automation servers, he found that connecting Chrome DevTools via MCP consumed approximately 18,000 tokens. This usage accounted for about 9% of a 200,000-token context window before the agent had performed any useful work.

Other tools presented similar burdens. Playwright MCP required roughly 13,700 tokens to describe its 21 tools, consuming 6.8% of the same context window. Each additional server added more overhead, reducing the space available for actual code and reasoning. Zechner also criticized the lack of composability in standard MCP setups, noting that results returned by a server must pass through the agent’s context to be persisted or combined with other data.

During this period, Pi relied on Bash scripts and command-line interfaces for browser automation. These tools required only a 225-token README because the model already understood how to use them. Outputs could be piped, filtered, or saved to disk without clogging the model’s context window. Community extensions like pi-mcp-adapter allowed users to bridge MCP to Pi, but native support remained absent.

Earendil, which acquired Pi earlier this year, decided to revisit MCP. The company stated that the protocol had matured and that the architectural changes required to support it effectively were broadly useful. Rather than adopting the standard exposure model, Pi 1.0 implements MCP through its existing Codemode infrastructure, allowing it to support the protocol while avoiding the context bloat Zechner originally opposed.

How it works

Codemode acts as an intermediary layer between the language model and its available tools. It runs inside a QuickJS sandbox, which means it operates without Node APIs, file system access, network connectivity, or timers. This isolated environment allows scripts to call Pi’s tools and models, execute operations concurrently, and process results before returning anything to the main model.

By default, Pi excludes MCP server tool definitions from the model’s prompt. The system prompt receives only a one-line description of each connected server. When the agent needs to perform a task, it uses Codemode to discover the appropriate tools, execute them via JavaScript, and return only the relevant output. This approach keeps the initial context lightweight and ensures that token usage scales with actual activity rather than potential capability.

Developers can customize this behavior using a setting called toolExposure. This feature allows granular control over how tools are presented. For example, in a GitHub integration, a team might expose the search_code tool directly to the model for frequent use, keep get_* methods behind Codemode for on-demand discovery, and block delete_* methods entirely to prevent accidental data loss. This flexibility lets engineers balance convenience with safety and efficiency.

Key details

  • Chrome DevTools MCP previously consumed ~18,000 tokens, or 9% of a 200k context window, before any action was taken.
  • Playwright MCP required ~13,700 tokens to describe its 21 tools, representing 6.8% of the same context window.
  • Codemode runs in a QuickJS sandbox without Node APIs, file system, network, or timer access.
  • Pi 1.0 reduces the default Codemode footprint from ~5,300 prompt tokens to ~3,300 for GPT-5.6 requests.
  • A default budget of 3,000 tokens is allocated for tool declarations; excess tools remain discoverable via Codemode.
  • The toolExposure setting allows developers to expose, hide, or block specific tools on a per-server basis.

Why it matters

For engineers building AI agents, context window efficiency is a critical constraint. Every token spent on tool definitions is a token not available for code analysis, reasoning, or history. By keeping MCP tools hidden by default, Pi demonstrates a viable path to supporting rich ecosystems without sacrificing performance. This method allows agents to connect to dozens of servers without immediately exhausting their context budget, enabling more complex and multi-step workflows.

The granular control offered by toolExposure also addresses security and usability concerns. Directly exposing all tools to a model increases the risk of unintended actions, such as deleting resources or making unauthorized changes. By forcing less critical or dangerous tools to go through Codemode, developers can add a layer of scrutiny and processing. This pattern encourages a more deliberate design where tool access is tailored to the specific needs of the agent, rather than accepting a blanket exposure model.

What you can do

  • Audit your current MCP integrations to measure the token cost of tool definitions in your prompts.
  • Implement a mediation layer similar to Codemode to intercept tool calls and process results before they reach the model.
  • Use sandboxed environments like QuickJS to execute tool logic securely without granting full system access.
  • Configure granular exposure settings for your tools, hiding rarely used or high-risk functions behind discovery mechanisms.
  • Optimize tool descriptions and documentation to fit within strict token budgets, removing redundant declarations.
  • Test the impact of hiding tools on agent performance, ensuring that discovery latency does not hinder user experience.

Tools aus dem Bytechap-Shop

Weiterlesen

KI-Agenten

Anthropic integriert dauerhaft aktive Cowork-Agenten in Claude

Anthropic hat seinen Agenten für geplante Aufgaben, Cowork, in die Haupt-App von Claude für Pro- und Max-Nutzer integriert. Damit tritt das Unternehmen im Wettrennen um vertrauenswürdige Automatisierung gegen OpenAI Dots und Meta Muse an.

Alle Artikel