Developer tools

GitHub Copilot's new local routing leaves data privacy questions unanswered

Microsoft plans to route GitHub Copilot tasks between local and cloud models by late October, but has not clarified what data leaves the device or how to restrict it.

Illustration of local versus cloud AI processing uncertainty
Illustration generated for this article

GitHub is preparing to automatically route coding tasks between local and cloud models for Copilot users, with the feature expected to arrive by the end of October 2026. While Microsoft describes this as a performance optimization, the company has not disclosed exactly which context data is sent to the cloud or provided a way for developers to force local-only inference.

What happened

Microsoft outlined its plan in a post co-written by GitHub product manager Patrick Nikoletich and Windows platform partner architect Stuart Schaefer. The announcement coincides with the general availability of new sandboxing controls for Copilot, though the level of protection varies depending on the specific tools being used. The core change involves expanding Project HydraFusion, a system that already selects which model to use, to now decide where that model runs.

The new Auto routing feature will weigh task context and cache state to switch between local and cloud inference, even during multi-turn sessions. This functionality will be available in Copilot CLI, the Copilot app, and VS Code. Developers can choose Auto routing or manually select a local model, such as MAI Code 1.1 Flash via the Windows ML provider or other OpenAI-compatible local endpoints. However, Microsoft explicitly acknowledges that local inference does not make the session fully offline.

This ambiguity has raised concerns among teams with strict data-handling policies. Similar issues arose recently when Anthropic admitted to routing Claude Sonnet 5.5 requests to an older model version for higher-risk activities. For GitHub users, the lack of transparency means they do not know how much repository context or conversation history Auto sends to the cloud, nor can they currently inspect these routing decisions.

How it works

To make local inference feasible, Microsoft employed aggressive quantization and speculative decoding. The MAI Code 1.1 Flash model is a mixture-of-experts architecture with 137 billion total parameters, but only 6.8 billion are active at any given time. Microsoft used mixed-precision quantization at roughly 3.3 bits per weight to reduce the model size from its original bfloat16 cloud version to 53GB, an 80% reduction. Speculative decoding further speeds up the process by having a smaller drafter model propose blocks of tokens for the main model to verify.

Despite these optimizations, the hardware requirements remain steep. The initial rollout targets NVIDIA RTX Spark Windows PCs, such as the Surface Laptop Ultra, which offers up to 128GB of unified memory. Microsoft measured peak memory use at 75.5GB with a 256K-token context. This figure excludes the memory needed for the operating system, applications, inference runtime, and the key-value cache, which grows as the agent reads files. Consequently, most developer laptops with 16GB or 32GB of RAM cannot support this local mode.

Security enforcement relies on Microsoft’s open source Execution Containers (MXC) library. On Windows, it uses the BaseContainer tier of the ProcessContainer backend; on macOS, it uses Seatbelt; and on Linux, it uses bubblewrap. These OS-level restrictions apply to shell commands and local MCP servers regardless of whether the task runs locally or in the cloud. However, built-in file tools rely on checks inside the agent harness rather than OS isolation, and remote MCP servers remain outside the local process sandbox entirely.

Key details

  • Auto routing for GitHub Copilot is expected to launch by the end of October 2026.
  • Microsoft has not specified how much repository context or conversation history is sent to the cloud during Auto routing.
  • The MAI Code 1.1 Flash model requires 53GB for weights and peaks at 75.5GB memory usage, ruling out most standard laptops.
  • Local inference does not guarantee an offline session, as agents may still access external services or network tools.
  • Remote MCP servers are not covered by the local process sandbox, relying instead on connection policy checks.
  • The quantized local model scored 70.8% on SWE-Bench Verified, slightly below the 72.6% score of the full-precision cloud version.

Why it matters

For software engineers and technical leads, the primary concern is data sovereignty. Many organizations prohibit sending proprietary code or internal documentation to public cloud models. Without visibility into what Auto routing sends to the cloud, or the ability to restrict inference to local models, compliance teams cannot approve the use of Copilot for sensitive projects. The statement that local inference is not offline undermines the assumption that choosing a local model ensures data privacy.

Furthermore, the hardware barrier limits the accessibility of this feature. Requiring 128GB of unified memory and high-end NVIDIA hardware means that only a small fraction of developers can benefit from local inference. This creates a disparity where only those with top-tier workstations can potentially keep more data on-device, while others remain dependent on cloud routing without clear safeguards. The complexity of managing sandbox policies across different tools and operating systems also adds operational overhead for IT leads.

What you can do

  • Audit your current Copilot usage to identify any workflows involving sensitive code or internal repositories.
  • Monitor Microsoft’s documentation for updates on whether a local-only restriction option will be added to Auto routing.
  • Evaluate your team’s hardware capabilities against the 75.5GB peak memory requirement before planning for local inference adoption.
  • Review sandbox policies for shell commands, local MCP servers, and remote MCP servers to understand where gaps in isolation exist.
  • Consider manually selecting local models if available, but verify that tool access is locked down to prevent unintended network requests.
  • Engage with your security team to define acceptable risk levels for data leakage given the current lack of transparency from Microsoft.

Tools from the Bytechap store

Keep reading

All stories