GitHub Copilot with local models and Auto mode: what Microsoft hasn't explained yet
Back to the blog
Articlegithub copilotagente de programaçãodev toolssegurançamodelos de IAverboo code

GitHub Copilot with local models and Auto mode: what Microsoft hasn't explained yet

MafraOctober 9, 20265 min read

On 10/07/2026 Microsoft announced that GitHub Copilot will run local models and decide on its own, task by task, what stays on your machine and what goes to the cloud. The next day The New Stack published the question the announcement leaves open: what, exactly, goes to the cloud? Both sides are partly right, and the difference matters to anyone with a data policy.

What did Microsoft announce for GitHub Copilot?

Local models in Copilot CLI, the Copilot app and VS Code, expected by the end of October 2026. The announcement was signed by Patrick Nikoletich (GitHub) and Stuart Schaefer (Windows) on Microsoft's Command Line blog.

  • Two ways to use it: Auto mode, where Copilot picks between local and cloud models, or picking the local model yourself.
  • The model: MAI Code 1.1 Flash, a mixture of experts with 137 billion total parameters and 6.8 billion active. Quantized, it shrinks from 265 GB to 53 GB.
  • Where it runs: through the Windows ML provider, or any OpenAI-compatible local endpoint.
  • Launch hardware: Windows PCs with NVIDIA RTX Spark, such as the Surface Laptop Ultra, with up to 128 GB of unified memory.
  • Same day: Copilot's local sandboxing became generally available, configurable with the /sandbox command.

What is the case for Auto mode?

Taking away an infrastructure decision developers don't want to make on every task. The announcement puts it plainly: "With Auto, developers do not need to decide where each task should be run."

According to Microsoft, the router weighs task context and cache state when switching between local and cloud, even in the middle of a long conversation. And the local model is no toy: the quantized version scored 70.8% on SWE-bench Verified, against 72.6% for the full version, and 66.29% against 62.9% on Terminal-Bench 2.1 (Microsoft's own numbers, 10/07/2026). Losing under 2 points while fitting on a laptop is a serious result.

What is The New Stack's criticism?

Microsoft hasn't said how much of the repository Auto mode sends to the cloud, whether you can see each routing decision, or whether Auto can be locked to local only. Amanda Caswell's report (10/08/2026) sums it up: teams with strict data policies still don't know what leaves the machine.

The announcement itself concedes the central point: "Local inference does not make the session offline." The report adds three details that matter day to day:

  • Choosing the local model keeps inference on the machine, but the agent's tools can still make network requests.
  • Remote MCP servers sit outside the local sandbox.
  • Measured peak memory was 75.5 GB at a 256K-token context. That rules out most developer laptops with 16 or 32 GB.

On the Terminal-Bench gain, the report notes the set has 89 tasks: the gap is about 3 tasks. It shows compression didn't break the model, not that it improved it.

Where do both sides agree, and where not?

PointWhat Microsoft statesStill unanswered
Who decides where it runsAuto mode, per taskWhether you can see each decision
Restricting to localYou can pick the local model yourselfWhether Auto can be locked to local only
What goes to the cloud"Local inference does not make the session offline"How much history and code go along
Tool network accessShell and local MCP run in the OS sandboxRemote MCP sits outside the local sandbox
Hardware53 GB of weights, RTX Spark with up to 128 GB75.5 GB peak at 256K rules out a regular laptop
Quality70.8% on SWE-bench Verified, quantizedTerminal-Bench gain is about 3 tasks

How do you check where your coding agent connects?

Watch the process's network connections while it works. It won't tell you what's inside each request, but it shows whether a "local" session is talking to the internet.

On macOS or Linux, with the agent running in another terminal:

lsof -a -i -P -n -p "$(pgrep -f copilot | head -1)"

On Linux, the alternative with ss:

sudo ss -tpn | grep -i copilot

On Windows, in PowerShell:

Get-NetTCPConnection -State Established |
  Where-Object OwningProcess -in (Get-Process -Name *copilot*).Id |
  Select-Object RemoteAddress, RemotePort, OwningProcess

Replace copilot with your agent's process name. If it only shows up as node, filter by PID. A connection to localhost is the local model; any external address is traffic that left the machine, whether from inference or from a tool.

What is Verboo Code's position?

Automatic routing is a good idea that only works if it's auditable. Until Microsoft answers The New Stack's three questions, the safe stance for anyone with a data policy is to treat Auto mode as cloud until proven otherwise, and to use a hand-picked local model when the code can't leave.

We took the opposite path to hybrid: Verboo Code runs in the cloud, with no mode that switches between local and cloud underneath. Every session's startup screen says so on its second line, in the format ● provider · model · cloud, and /model shows which model is answering. You don't get the model on your machine, but you don't have to guess where your code went either.

We compared both tools side by side in GitHub Copilot vs Verboo Code.

Enjoyed this article?
Share knowledge with your network.
// Read also

Related articles