Models & cost
Nemocode does not sell you tokens. You bring your own provider, and you decide which model does what.
Connecting a provider
/connect
Pick a provider, paste the key, done. Several providers can be connected at the same time, and each role can use a different one.
| Provider | How you connect |
|---|---|
| Anthropic | API key. |
| OpenAI | API key, or OAuth / Codex: you sign in with your ChatGPT account in the browser and no key is needed. Nemocode listens on local port 1455 for the answer of that sign-in, so keep it free. |
| OpenRouter | API key. One key, many models. |
| Moonshot | API key (Kimi models). |
| Ollama | Either an Ollama Cloud key, or the address of a local
or self-hosted server (default http://localhost:11434). A
local server needs no key at all. |
| opencode Zen, opencode Go | Two separate endpoints, each with its own API key from the opencode account. |
| DeepSeek, Groq, Cerebras | API key. |
| Custom | Any OpenAI-compatible endpoint: base URL, model name, optionally an API key, the context window (default 128k) and extra request headers. |
Keys stay on your machine. Instead of typing a key you can also put it in the provider's environment variable — see Settings.
The main model
/model
Picks the model of the main agent from the live list of the provider — the same as the main-agent line under Settings ▸ Models. In the picker, w sets the context window Nemocode assumes for a model. Nemocode usually knows it; set it when you run a model it cannot recognise, for example a custom endpoint.
A model per role
A mission is not one model doing everything. Under Settings ▸ Models each
role can have its own model and its own provider. A role you leave on
“(default)” uses the model you picked with /model; if you never
picked one, a built-in default of the connected provider (Anthropic and Ollama
have separate defaults for the small roles, other providers use one model for
all).
| Role | What it does |
|---|---|
| main_agent | Reads, plans, writes the code. |
| verifier | Checks the finished work before the mission ends (see Verify subagent). |
| curator | Harvests each mission into proposals and talks with you about the project memory. One model for both. |
| advisor | Consulted at decision points — should be at least as strong as the main agent. |
| explore, delegate | Subagents that read broadly or build an independent part. |
| compactor | Summarises the conversation when it gets long. |
| gpc, gpc_gate | Pre-compression of the carried conversation, and the small model that decides per mission whether to do it (only with GPC on smart). See Context & cost. |
| goal_gate | After each turn of a /goal, judges
whether the condition is met. Only used while a goal is set. |
| vision | Describes pasted images for a main model that cannot see them. Without this role, such a model is told to inspect the image with a command. |
| reviewer, synthesis | The review and the joined report of
/ultracode and /ultrareview. |
| prompt_rewrite | Tidies your request when an ultra keyword sat in the middle of a sentence. |
The cheap way to run: a strong model for the main agent, and a small fast one assigned to the roles that summarise, judge and search. Assign them explicitly, because a role on “(default)” follows the main model.
Per session instead of everywhere
Tick Only this session at the top of the Models tab and your picks stay inside that session — models, providers and effort per role. Nothing else changes. In the list, a role that has a session choice is marked (this session). Untick it and the session's choices are deleted; it uses the global ones again. The header shows the model that is really running in this session, so you can always see which one you are paying for.
Effort
/effort sets how hard the model thinks — the levels are a subset
of none, minimal, low, medium, high, xhigh, max. Only the levels your
provider and model actually support are offered; some models have no
medium, some no max. The default is medium.
You can set the effort per role: press e on its
line under Settings ▸ Models. Each role shows its own level in brackets, and a
role without one inherits the global level. /ultrathink runs just the next mission
at maximum.
Seeing the cost
| Where | What it tells you |
|---|---|
| The token line under a mission | Tokens in, out, cached — and the price, if you switched on Show $ cost. |
/usage | The whole session, and what each role and subagent cost of it. |
| Settings ▸ Stats | Per day, across every project. |
Prices and context windows
The cost display multiplies tokens by the price of the model. Where the price comes from, in this order:
- Your own price, set under Settings ▸ Advanced ▸ Pricing. Nemocode never replaces it by itself; x on that line removes it.
- Nemocode's built-in price table.
- A live price list, which you can pick from.
The context window of a model comes from your own value (/model
▸ w) first, then from what Nemocode measured for the
model, then from a guess by model name, and 256k when nothing is known. A custom
endpoint uses the window you entered when you connected it as the ceiling for
guessed values.
Context and compaction
Every model has a limit. Nemocode tracks how full the context is and summarises the conversation before it runs out, keeping working. Two settings shape that: the context budget (how much it may use before compacting) and compaction ask (whether it checks with you first). Everything about it is on Context & cost.
/context shows where the tokens are going — the conversation,
the files, the memory, the tool results.
What Nemocode does to keep it cheap
- The unchanging part of a mission's prompt is arranged so your provider can cache it between calls.
- Re-reading an unchanged file sends a short note, not the file again.
- Old tool output is dropped when carrying it costs more than re-fetching it would.
- Long shell output is compressed before the model sees it.
The last three are switchable under Settings, and are on by default.