Local models (Ollama)
Overdeck can point an agent at a model served locally by Ollama instead of a cloud provider. Ollama 0.14 and newer serves the Anthropic Messages API natively, which is the API Claude Code speaks — so the harness talks to your GPU exactly the way it talks to Anthropic, the model traffic never leaves the host, and the run records$0.
Model ids are ollama:<tag>, for example ollama:gemma4:12b. Overdeck strips the
ollama: prefix before the tag reaches the server; the prefix is what routes the
id to your local endpoint instead of a cloud provider.
Requirements
- Ollama 0.14.0 or newer is the floor Overdeck enforces: earlier releases do not
serve the Anthropic Messages API, and the preflight refuses to launch against them
rather than failing mid-turn. Individual models need more. A model’s manifest
can require a newer Ollama than 0.14, and the pull then fails with
412: The model you are attempting to pull requires a newer version of Ollama—gemma4:12bdid exactly that on 0.19.0. Run the newest Ollama you can; upgrade if a pull returns 412. - About 24 GB of GPU memory for a 12B model at Q4. Verified on an RTX 3090
(24 GB, CUDA) on Linux with
gemma4:12b, which used 9.2 GB of VRAM at a 64K window. Apple Silicon with 24 GB of unified memory (Metal) is a target shape, not yet verified. - Disk for the model.
gemma4:12bis roughly 8 GB at Q4 quantization. - The claude-code harness. Other harnesses are refused for
ollama:models in this release.
Install Ollama
Overdeck never runs an installer for you — acurl | sh fired from a setup script
is exactly the thing worth reading first. Run it yourself:
pan install detects Ollama. If it is missing, the install prints the command for
your platform and carries on. If it is present and you are on an interactive
terminal, it offers to pull gemma4:12b (defaulting to No, because 8 GB is not a
download to start by accident, and because it is not a working work-agent model). Pass
--skip-ollama to skip the step entirely.
Pull a model
Any tag your server has pulled works.gemma4:12b is the tag Overdeck’s install step
offers and the one this page’s examples use, because it is what the wiring was verified
against — not because it works as a work agent. It does not; see the warning above.
Overdeck never substitutes it for a model you asked for.
Set the context length
This is the setting that decides whether local agents work at all. Ollama silently truncates a prompt longer than the model’s loaded context window: there is no error, the agent simply stops seeing the beginning of its own instructions. An Overdeck work agent’s first prompt alone runs to tens of thousands of tokens, so the default window is far too small. The window is the server’s to set, so you have to set it on the server. Overdeck reads back whatever window the server gave the model and pins Claude Code to that number, so the harness’s own compaction agrees with reality — but it cannot raise the window for a server it did not start. SetOLLAMA_CONTEXT_LENGTH:
OLLAMA_CONTEXT_LENGTH from ollama.context_length, so the config
value is enough and the steps above are unnecessary.
pan doctor warns about any loaded model whose window is under 64K, which is the
check to run after changing this.
Configure Overdeck
Point a role — or a workhorse slot — at the local model in~/.overdeck/config.yaml.
Because no local model has completed a work-agent task yet, prefer --model on a single
pan start over pinning roles.work for real work:
ollama: block tunes the endpoint:
base_url with no port gets Ollama’s default 11434.
A non-localhost base_url is refused at config load. The point of a local model is
that no prompt leaves the machine, so a remote endpoint is a configuration error
rather than a supported deployment.
Run an agent
ollama serve if nothing is listening, checks the version, checks the tag is
pulled, and warm-loads the model to read back its real context window. Every
failure names the fix rather than leaving you with a dead agent:
What pan doctor shows
pan doctor prints Ollama rows only when this host has a reason to care — a
configured ollama: model, or the binary installed. It reports the version and
base URL, warns per configured tag that is not pulled, and warns when a resident
model’s window is under 64K. It never warm-loads a model, so running it does not
pull 8 GB into VRAM as a side effect.
What pan up does
pan up starts ollama serve only when your config names an ollama: model. On
every other host it makes no network call at all. If the server will not start,
pan up prints a warning and continues — cloud-model agents are unaffected.
Limits
- claude-code only.
codex,opencode,kimi-code,acp, andmuseare refused forollama:models. - Localhost only. A non-localhost
base_urlis a config error. - Local workspaces only. The preflight checks this host, so remote (Fly.io) workspaces cannot use local models.
- Cost records
$0. Local runs carry no pricing row, by design. - Model traffic is local; the process is not.
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFICis exported, so Claude Code’s own telemetry is off, but any remote MCP servers you have configured are still contacted at startup. If you need a fully offline run, unconfigure them too. - No local model has completed a work-agent task yet. See the warning at the top.
- The dashboard model picker and the Settings provider card do not list local tags
yet; configure them in
config.yamlor pass--modelon the command line.