Skip to content
CryptoDecentral

Note

Inference Residency Is Not Action Residency: Tools, MCP, and Local LLMs

Local open-weight sampling can keep prompts on machines you control — that is inference residency. Tools, Ollama web search/fetch, and MCP servers move side effects off-box. Grammar-constrained decoding (GBNF) stays local without remote tools.

2026-09-28 · local-ai, ollama, mcp, tool-calling, trust-boundary, llama-cpp, gbnf, protocol

Dark CryptoDecentral hex mesh — glowing teal core ringed by a bronze boundary, with thin bronze rays crossing outward beyond the ring; no text.

Inference Residency Is Not Action Residency: Tools, MCP, and Local LLMs

A laptop that runs an open-weight model through Ollama or llama.cpp can keep prompts and completions on disk and in RAM you control. That is inference residency. It is not the same claim as "nothing leaves this machine."

Thesis, stated once: once tools, web search, function calling, or Model Context Protocol (MCP) servers are wired into the loop, a local LLM is no longer a closed trust boundary. Inference residency answers where tokens are sampled. Action residency answers where tool side effects land — filesystem reads, HTTP fetches, API calls, subprocesses. Confusing the two is how "air-gapped chat" demos quietly grow egress.

This note is mechanism literacy from primary docs — not a product review and not a how-to for abuse.

What "tool calling" actually wires

Ollama's tool-calling docs describe the concrete loop on POST /api/chat (tool calling). You send a tools array of JSON Schema–shaped functions (type: "function", name, description, parameters). The model may reply with tool_calls instead of (or alongside) plain content. Your host code then executes those calls, appends messages with role: "tool" (and tool_name / content), and calls /api/chat again. Parallel calls, streaming accumulation of partial tool_calls, and a multi-turn agent loop (while there are still tool calls) are first-class patterns in the same page.

Three residency facts follow from that design, not from marketing:

  1. The model proposes; the host disposes. Sampling can stay on localhost:11434. The side effect happens in whatever process runs the function — your Python/JS agent, a shell helper, a browser automation bridge.
  2. Tool results re-enter the context. A role: "tool" payload is untrusted data from the tool's world. A poisoned search snippet or a file-read result becomes prompt material for the next turn.
  3. The agent loop is open-ended by default. Official examples keep calling until the model stops emitting tool_calls. Without an application policy (allowlists, max steps, human approval), "local chat" is an orchestration runtime with a language model as the scheduler.

Local inference without tools can still be a closed box for tokens. Local inference with tools is a control plane over actions. Those are different threat models wearing the same UI chrome.

Web search and web fetch: egress by design

Ollama's web capabilities make the boundary explicit (web search). POST https://ollama.com/api/web_search and POST https://ollama.com/api/web_fetch run against ollama.com, not against your loopback daemon. Authentication is an OLLAMA_API_KEY (Bearer) from a free Ollama account. Search returns title / url / content snippets; fetch returns page title, content, and links.

Their own "building a search agent" example registers web_search and web_fetch as tools in the same agent loop as local chat(). The model can stay on a small local weight (the docs show qwen3:4b); the queries and fetched pages still leave the machine. The same page documents an MCP stdio server wrapper that injects those tools into Cline, Codex, Goose, and similar clients — again with the API key in the server environment.

So: "I run Ollama locally" and "my agent's research path is local" are independent checkboxes. Ticking the first does not tick the second. Query strings, fetched URLs, and result snippets are action-residency events even when every token of the final answer was sampled on your GPU.

MCP: clients trust servers (by design)

The Model Context Protocol standardises how hosts connect LLMs to external tools, resources, and prompts (specification 2025-11-25). Hosts initiate; clients inside the host talk JSON-RPC to servers that expose tools, resources, and prompts. The security section is unusually blunt for an integration spec:

  • Users must consent to data access and operations; hosts should expose clear UIs for review.
  • Tools are arbitrary code execution and must be treated as such.
  • Tool descriptions and annotations are untrusted unless the server itself is trusted.
  • Hosts must obtain explicit user consent before invoking any tool.

The specification repository's SECURITY.md states the trust model in plain language: MCP clients trust the servers they connect to. Local MCP servers are trusted like any other software you install. On stdio transport, the client spawns a command that runs with the client's privileges; "arbitrary command execution via STDIO configuration" is an intended feature, not a CVE-shaped surprise. The SDK does not sandbox peers across stdio — a malicious local server already has code execution by virtue of being run. Deployments that want isolation must enforce it themselves (containers, OS sandboxes); the transport is not a sandbox (community security summary).

Security best practices for implementors add the operational attack surface around that trust: confused-deputy OAuth proxy flows, forbidden token passthrough, SSRF via malicious metadata URLs, session hijack, and local MCP server compromise (malicious startup commands, exfiltration via curl, DNS rebinding to a lax localhost HTTP server). Mitigations they require of serious clients include showing the exact command before one-click local server config, consent before invoke, and sandboxing where feasible (security best practices).

For a residency reader, the compression is simple: an MCP filesystem server that works is a filesystem ACL for the model. An MCP web or browser server that works is intentional egress. Both can be exactly what you want. Neither is implied by "weights are on disk."

Structured output without remote tools: GBNF

Not every structured need requires a tool. llama.cpp's GBNF (GGML BNF) grammars constrain sampling so outputs match a formal grammar — valid JSON, a fixed schema shape, and similar — entirely inside the local decoder (grammars README). You pass --grammar / --grammar-file to CLI tools, or a grammar / json_schema body field to llama-server. A subset of JSON Schema converts to GBNF; the schema constrains tokens and is not injected into the prompt (tool schemas, by contrast, typically are visible to the model).

That contrast matters. Grammar-constrained decoding is still inference residency: no network hop is required for the structure itself. Tool calling is action residency: structure is a request for the host (or an MCP server) to do something in the world. Use GBNF when you need a shape. Use tools when you need a side effect — and then name the side effect's trust island honestly.

Why this matters on constrained machines

Journalists, clinics, and organisers often choose local models precisely because WAN quality is poor, jurisdictional transfer is fraught, or a cloud chat log is an unacceptable residue. That choice still holds for chat-only stacks. It weakens the moment a "helpful" filesystem MCP can read a source corpus into the context window, or a cloud search tool phones home with query text that fingerprints the investigation. Constrained hardware makes the temptation sharper: small local weights plus a remote research tool feel like a free upgrade. They are an upgrade in capability and a downgrade in boundary clarity unless each tool's egress and privilege is listed, consented, and reviewable.

Civil-liberty stakes are concrete without needing a bank metaphor. A model that never leaves the laptop can still exfiltrate notes through a tool you installed because a tutorial said "enable MCP." Residency is a property of the whole agent graph, not of the GGUF filename.

What this note deliberately does not do

It does not rank Ollama, MCP clients, or cloud search products. It does not give investment, token, custody, or trading advice. It does not claim local models are accurate, uncensored, or immune to prompt injection. It does not teach bypassing consent UIs, sandboxes, or lawful process. It does not treat MCP's intentional stdio command execution as a vulnerability report. It does not prescribe a compliance checklist for any organisation.

Educational takeaway: where the weights live is not where the tools act. Audit the tools array, the MCP server command lines, and every URL that is not loopback — then decide whether your stack still matches the residency story you told yourself.

Related on CryptoDecentral

Further reading (primary)

All notes