Note
A Decision Endpoint Is Not a Chat Completion: Ollama’s /v1/systemone and What You Can Verify
Ollama’s /v1/systemone returns choices, probabilities, and scores instead of chat text. A decision endpoint is a different contract from /v1/chat/completions — what you can check, and what you still trust.
2026-10-05 · oss-ai, local-ai, ollama, decision-models, systemone, api-surface, open-weights, verify, data-residency, sovereignty, civil-liberties

A Decision Endpoint Is Not a Chat Completion: Ollama’s /v1/systemone and What You Can Verify
Open-source AI special · week of 5 Oct 2026
Most local AI stacks have trained operators to think of the model server as one shape: send messages, get text back. This week’s thesis: a decision endpoint is not a chat completion. Ollama v0.35.0 (28 Sep 2026) added decision models behind a new route, /v1/systemone, “based on TypeSafe’s Jev API.” Per the release notes, decision models “return choices, probabilities, and scores instead of text.” v0.35.1 (29 Sep) followed with Cloudflare’s open-source Clef and Clef Flash models, image input, and a capability change that matters more than it looks: decision models now report only decision as their capability, “so clients no longer offer them for general chat, tools, or thinking.”
Same daemon, same port, a different contract. CryptoDecentral frame: local inference without kill-switch dependency only works if you know which contract you are calling — and which parts of it you can check yourself.
What actually shipped
The request shape is not a conversation. You send a model, a state (a string, or a JSON object or array), and between 1 and 64 named questions. Each question has a type — choice, noul (yes/no), or score — plus instructions and, usually, criteria. Choice and score questions take 2 to 26 options. The response comes back as typed answers keyed by your question names:
"label": {
"type": "choice",
"choice": "bug",
"probabilities": { "billing": 0.0125, "bug": 0.9781, "account": 0.0093 },
"confidence": 0.8906
}
The launch models, per the Ollama blog (29 Sep): nimble, a 9B open model from Bespoke Labs, plus tev1 (4B) and tev1:0.8b, both described as experimental, from Together AI. v0.35.1 added clef (27B) and clef-flash (9B), which accept base64 images alongside state, “shared by all questions and scored jointly with it.” The Nimble and Clef Flash library pages both say the models are fine-tuned from Qwen3.5-9B and released under Apache 2.0.
Why it is a different surface, not a chat flavour
Set the System One API reference beside Ollama’s OpenAI-compatibility page and the differences are structural:
- No generation controls.
/v1/chat/completionstakesmessages,temperature,top_p,max_tokens,tools, streaming, and reasoning effort./v1/systemone“returns one JSON response; streaming, video, tools, and generation controls are not supported.” - Bounded output. A chat model can say anything. A decision model, in Nimble’s own model card, “can only pick from the answers you give it. It won’t write explanations, nested JSON, or quotes pulled from the context.” The output space is the set of criteria you supplied.
- No silent truncation. The reference says “Input is never truncated.” Requests over 64 KiB (32 MiB with images) get a 413. A prompt that doesn’t fit the loaded context gets a 400 instead of being quietly cut down.
- Local-only, today. The
modelfield requires “compatible GGUF weights and a scoring-capable runner; cloud and MLX/Safetensors models are not supported,” and the 400 response states “Cloud models are rejected.” That contrasts directly with the chat side, where the compatibility docs note that a signed-in local server can serve cloud models such asgemma4:cloudthrough the same/v1/prefix. The GGUF-only part is already moving: the v0.40.0-rc3 pre-release (5 Oct) runs decision models on MLX too, while still rejecting cloud models. - Numbers with defined meanings.
confidenceis defined as1 − H(p) / ln(N), a measure of how concentrated the distribution is, and is explicitly “not calibrated correctness.”noulis “a number, not a Boolean.”scoreis a probability-weighted average of zero-based levels, “not rounded to a level or normalized to 0–1.” Evenusage.output_tokenscounts internal scoring tokens, “not the length of the JSON response.”
That last point is the core of this note. A chat completion is text you interpret. A decision response is a set of numbers whose meanings are defined in a schema, so you can check them against it.
What an operator can verify
- Version and route. The route ships in v0.35.0, and Clef needs v0.35.1. On an older daemon, the request should fail outright. Record the daemon version alongside the model tag.
- Capability metadata. After v0.35.1,
ollama showshould list onlydecisionfor these models. v0.35.1 also adds ModelfileCAPABILITYdeclarations that persist through GGUF/safetensors creation, inheritance, and export. That makes capability something you can inspect as part of the artefact, rather than something a client UI guesses. - Residency at the endpoint. Send a
:cloudtag to/v1/systemoneand expect a 400. That is a cheap, repeatable check that this route is not forwarding off-box today. - Schema invariants. Probabilities should sum to 1 within floating-point error.
confidencecan be recomputed from the probabilities.scoreshould fall between 0 and N−1. If a response breaks the published maths, you have a real bug report, not a vibe. - Failure modes. Oversized input should produce a 413 or 400, never a shorter prompt. Test that before you depend on it.
What you still trust
- The prompt you don’t write. The Nimble card says “Ollama builds Nimble’s prompt for you.” The template that turns
stateandcriteriainto tokens sits inside the server. You verify the inputs and outputs, not the framing in between. - Per-model scoring mechanics behind one interface. Nimble “reads the prompt once per question and scores the answer tokens directly,” inside an 8,192-token context. Clef Flash advertises scoring “every option of every question … jointly in a single non-autoregressive pass” with a 64K window. Same endpoint, same JSON, different internals. The reference also notes that “answers are not passed to later questions,” and Nimble’s card warns that if two answers need to agree, you should “check that in your code.”
- Calibration. Nimble’s card is blunt: “A probability of 0.9 doesn’t mean the answer is right 90% of the time on your data.” The accuracy tables published on the library pages are vendor and partner measurements. Treat them as claims to reproduce, not guarantees.
- Host policy and roadmap. “Cloud models are rejected” is documented behaviour in a fast-moving release line. The launch post itself says more decision models are coming, “including models served by Ollama’s cloud.” Local-only is current policy, so re-verify it after every upgrade, the same way you would re-read any default.
- Client drift. The library pages show
ollama.systemone(...)snippets for Python and JavaScript, while the same pages say decision models “aren’t in the Ollama CLI or the Ollama Python and JavaScript libraries yet.” Check which client version you actually have rather than trusting a snippet.
Where routing and gating meet sovereignty
Ollama’s examples include model routing (a choice between a small model and a large one) and tool-call moderation (a noul asking whether run_shell(command="rm -rf ~") “could cause harm”). Both put a small local classifier in front of something more consequential. That can be a real sovereignty gain, because the triage decision stays on hardware you control. But the decision model only picks which path to take. Where that path leads is still your policy. If one routing option points at a remote API, the request still leaves the box. And a 0.97 “harmless” score is a probability, not a permission system. Our earlier note on tools and MCP trust boundaries covered where actions execute. This note is about what kind of answer the API returns before anything acts.
What this note deliberately does not do
It does not rank Nimble, Tev1, Clef, or Jev against each other or against chat models. It does not guarantee any latency or accuracy figure; where numbers appear in vendor material, they belong to their publishers. It does not claim decision models replace chat models, or that /v1/systemone is a security boundary. It does not revisit runner defaults, embedded GGUF alignment, or tool residency as its main thesis. It gives no investment, token, custody, swap, or trading advice, and it does not recommend buying any product or service.
Frame: A chat completion hands you prose to interpret. A decision endpoint hands you numbers with published definitions. That makes it more checkable, but only if you actually check the version, the capability, the residency rejection, and the maths, and keep calibration and routing destinations in your own hands.
Educational takeaway: know which contract you are calling. The same local runner now exposes two very different APIs. Verify what it exposes. Trust only what you have tested.
Related on CryptoDecentral
- Decentralised / local AI pillar
- Local and Peer-to-Peer AI: Data Residency Without the Hype
- A Model Tag Is Not a Runner Contract: Ollama’s MLX Default and What Still Differs
- Successful Load Is Not Proof: Embedded GGUF Alignment and Wrong Tensors
- Inference Residency Is Not Action Residency: Tools, MCP, and Local LLMs (contrast: action residency vs API surface)
Further reading (primary)
- Ollama v0.35.0 — decision models via /v1/systemone
- Ollama v0.35.1 — Clef / Clef Flash, image input,
decisioncapability, ModelfileCAPABILITY - Ollama blog — Jev-style decision models
- Ollama API reference — System One
- Ollama docs — Decision capability guide
- Ollama docs — OpenAI compatibility (
/v1/chat/completions) - Ollama library — nimble
- Ollama library — clef-flash
Sources
- https://github.com/ollama/ollama/releases/tag/v0.35.0
- https://github.com/ollama/ollama/releases/tag/v0.35.1
- https://github.com/ollama/ollama/releases/tag/v0.40.0-rc3
- https://ollama.com/blog/ollama-now-supports-jev-style-decision-models
- https://docs.ollama.com/api/systemone
- https://docs.ollama.com/capabilities/decision
- https://docs.ollama.com/api/openai-compatibility
- https://ollama.com/library/nimble
- https://ollama.com/library/clef-flash