Skip to content
CryptoDecentral

Note

A Model Tag Is Not a Runner Contract: Ollama’s MLX Default and What Still Differs

Ollama v0.40.0-rc0 defaults supported Apple Silicon models to MLX under the same pull/run tags — a model tag is not a runner contract. Pin versions, name the engine, keep offline plans honest.

2026-09-28 · oss-ai, local-ai, ollama, mlx, llama-cpp, lm-studio, verify, apple-silicon

Dark CryptoDecentral hex mesh — one glowing teal model-tag capsule splitting into divergent teal and bronze runner paths ending in kernel nodes; no text.

A Model Tag Is Not a Runner Contract: Ollama’s MLX Default and What Still Differs

Open-source AI special · week of 28 Sep 2026

The string you type to start a local model is not a contract for which kernel, weight format, or server path will answer. This week’s thesis: a model tag is not a runner contract. On 25 Sep 2026, Ollama shipped pre-release v0.40.0-rc0 with a default flip that makes the point concrete — on Apple Silicon, architectures the MLX runtime supports now automatically run on MLX. The release notes use the familiar pair:

ollama pull qwen3.8
ollama run qwen3.8

Same pull/run vocabulary. Different default engine under the hood for supported Mac architectures — unless you verify which path actually loaded. Around that flip sit llama.cpp v0.5.0 (23 Sep) as the shared-engine major, Ollama’s RC diff preparing to drop llama-server compatibility patches, and LM Studio’s Splash engine as the contrasting pattern: specialized Apple Silicon kernels that stay opt-in, not a silent default.

CryptoDecentral frame: continuity without kill-switch dependency — verify weights and runner identity, keep residency claims honest, and treat defaults as policy you re-read.

What v0.40.0-rc0 actually changed

Primary notes are short. On Apple Silicon, model architectures supported by the MLX runtime will automatically run on MLX. Maintainers will test and enable additional models during the pre-release. That is a default-runner policy change, not a new CLI verb. Operators who assumed “ollama run <name> means the llama.cpp-shaped path on this Mac” need a new habit: confirm which runtime claimed the load after an upgrade.

Cite the surrounding RC work carefully as what appears in the v0.34.4…v0.40.0-rc0 compare, not as polished product docs. That diff includes a commit titled llama-server: prepare to remove compatibility patch, a new compatmigrate/ package (per-architecture migration helpers, tensor layout/ops, writers, large tests), and heavy manifest/ changes consistent with runner-aware or multi-format storage. Read with the release blurb: packaging is being rearranged so the product can stop papering over llama-server compatibility — while the user-facing tag string stays stable. Stability of the name is why runner identity must be checked separately.

Scope limits from the release text itself: the flip is Apple Silicon and architectures the MLX runtime supports. It does not mean every model on every Mac, or that Linux/Docker suddenly grows MLX. RC status means the supported set may still move — another reason to pin and re-verify.

Why the same tag is the scary part

A tag is a lookup key into layers, digests, and runner selection rules. When those rules change under a constant string, three failures become easy:

  1. Evaluation drift. Latency, memory pressure, and quirk surfaces move with the kernel. Replaying last quarter’s “local Qwen” numbers without naming the runner is cargo-cult science.
  2. Offline continuity. Bandwidth-constrained teams who cannot casually re-pull alternate format blobs feel defaults hardest. “I already have qwen3.8” is not “I can reproduce last month’s stack without WAN” if the RC expects a different on-disk layout or runner artefact than the one you cached for travel.
  3. Trust theatre. Hashing weight files still matters — see embedded GGUF alignment: successful load ≠ proof of correct tensors. This week’s layer is orthogonal: even correct bytes can run on a different engine than your runbook assumed. Verify artefact and runner.

Civil-liberty-adjacent failure mode, without bank metaphors: operators who chose local inference so a vendor cannot yank a remote endpoint still depend on local defaults remaining inspectable. A kill-switch is not only a cloud logout. It is also an opaque runner swap that forces a re-download you cannot afford, or a behaviour change you cannot attribute because the tag never changed.

The stack around the default

llama.cpp v0.5.0 (23 Sep; nightly noted as b11146) remains the shared-engine major under much of the local ecosystem even when a given Ollama path routes elsewhere. From the release notes: backend performance/correctness, broader model coverage, and more robust server/router operation. Mechanisms that matter when you bind your own daemons:

  • ggml 0.25.0, with RPC protocol major v7 among the ggml API changes — remote/backend peers need matching expectations.
  • Server --host accepts comma-separated TCP addresses and UNIX sockets — exposure policy becomes an explicit bind list.
  • Router hygiene: do not pass log file or API key file to router-spawned children.
  • API clarity: LoRA from an open FILE*, and docs that llama_model_load_from_file_ptr reads from the current position and requires aligned mmap — packaging discipline stays first-class (related to, not a rehash of, embedded alignment).

Those are engine-contract details. They do not make ollama run qwen3.8 print “MLX” by itself. They remind you the daemon stack has versioned protocols and bind surfaces worth pinning beside the model tag.

LM Studio Splash is the deliberate contrast. Splash is an open-source Inco AI engine for Apple Silicon with model-specific GPU kernels and memory plans for Qwen3.8-27B and Qwen3.6-35B-A3B, each shipping a dedicated DFlash 2 draft for speculative decoding (Splash post). LM Studio 0.4.25 adds Splash on M3+ / macOS 26.4+. Installation is opt-in: Bionic → Settings → Runtime → Experimental backends → download Splash (Metal), then pull Splash-packaged weights (e.g. incoai/Qwen3.8-27B-Splash). Inco’s tests on a 48 GB M5 Pro report large decode/throughput gains versus the next-fastest engine in their comparison — vendor/Inco attribution, not a universal ranking. Mechanism lesson: Splash advertises specialization and requires an explicit backend download. Ollama’s RC advertises convenience by making MLX the default for supported architectures. Both can be legitimate. Only one pattern trains operators to notice the runner.

Today’s cadence note on tools/MCP trust boundaries is a different thesis (inference residency ≠ action residency). This special stays on which engine samples tokens.

Sovereignty checklist (mechanisms)

  1. Record runner identity with the tag. After upgrades — especially RC lines — note which runtime claimed the model. A tag alone is an incomplete inventory line.
  2. Treat defaults as versioned policy. The MLX-on-Apple-Silicon flip is a policy event. Pin Ollama/llama.cpp/LM Studio versions where you pin digests.
  3. Read the RC packaging signal. compatmigrate/ and manifest/ churn warn that on-disk layout and compatibility shims are in motion. Plan migrations; do not assume last month’s offline blob set is forever enough.
  4. Separate opt-in specialization from silent defaults. Splash-style experimental backends make the contract visible. Default flips need the same visibility in your runbook even when the UI never asks.
  5. Keep residency claims scoped. Apple Silicon MLX defaults do not describe Linux servers, Docker hosts, or remote RPC peers on ggml RPC v7.
  6. Read the model card. Cards document intended formats; they do not automatically track your runner’s default after a client upgrade.

Frame: Open weights without a named runner are reproducibility theatre. A named runner on unverified packaging is the mirror failure. Sovereignty is the conjunction: verify the artefact, name the engine, keep sensitive hops on hosts you control.

What this note deliberately does not do

It does not rank MLX against llama.cpp against Splash. It does not guarantee tokens-per-second for any hardware; where Splash speed figures appear, they are attributed to Inco’s published tests, not re-measured here. It does not give investment, token, custody, swap, or trading advice. It does not claim v0.40.0-rc0 is already stable or that every qwen3.8 pull worldwide now resolves to MLX. It does not rehash embedded GGUF alignment or tool/MCP action residency as the main thesis. It does not teach bypassing lawful process. It is not MetaBot coaching and not DreamPiercing.

Educational takeaway: the tag is a lookup key; the runner is the contract. When defaults move under a stable string — as Ollama’s Apple Silicon MLX RC does — pin versions, inspect which engine loaded, and keep offline continuity plans honest about format and kernel, not only about model name.

Related on CryptoDecentral

Further reading (primary)

All notes