Models & cost

Nemocode does not sell you tokens. You bring your own provider, and you decide which model does what.

Connecting a provider

/connect

Pick a provider, paste the key, done. Several providers can be connected at the same time, and each role can use a different one.

ProviderHow you connect
AnthropicAPI key.
OpenAIAPI key, or OAuth / Codex: you sign in with your ChatGPT account in the browser and no key is needed. Nemocode listens on local port 1455 for the answer of that sign-in, so keep it free.
OpenRouterAPI key. One key, many models.
MoonshotAPI key (Kimi models).
OllamaEither an Ollama Cloud key, or the address of a local or self-hosted server (default http://localhost:11434). A local server needs no key at all.
opencode Zen, opencode GoTwo separate endpoints, each with its own API key from the opencode account.
DeepSeek, Groq, CerebrasAPI key.
CustomAny OpenAI-compatible endpoint: base URL, model name, optionally an API key, the context window (default 128k) and extra request headers.

Keys stay on your machine. Instead of typing a key you can also put it in the provider's environment variable — see Settings.

The main model

/model

Picks the model of the main agent from the live list of the provider — the same as the main-agent line under Settings ▸ Models. In the picker, w sets the context window Nemocode assumes for a model. Nemocode usually knows it; set it when you run a model it cannot recognise, for example a custom endpoint.

A model per role

A mission is not one model doing everything. Under Settings ▸ Models each role can have its own model and its own provider. A role you leave on “(default)” uses the model you picked with /model; if you never picked one, a built-in default of the connected provider (Anthropic and Ollama have separate defaults for the small roles, other providers use one model for all).

RoleWhat it does
main_agentReads, plans, writes the code.
verifierChecks the finished work before the mission ends (see Verify subagent).
curatorHarvests each mission into proposals and talks with you about the project memory. One model for both.
advisorConsulted at decision points — should be at least as strong as the main agent.
explore, delegateSubagents that read broadly or build an independent part.
compactorSummarises the conversation when it gets long.
gpc, gpc_gatePre-compression of the carried conversation, and the small model that decides per mission whether to do it (only with GPC on smart). See Context & cost.
goal_gateAfter each turn of a /goal, judges whether the condition is met. Only used while a goal is set.
visionDescribes pasted images for a main model that cannot see them. Without this role, such a model is told to inspect the image with a command.
reviewer, synthesisThe review and the joined report of /ultracode and /ultrareview.
prompt_rewriteTidies your request when an ultra keyword sat in the middle of a sentence.

The cheap way to run: a strong model for the main agent, and a small fast one assigned to the roles that summarise, judge and search. Assign them explicitly, because a role on “(default)” follows the main model.

Per session instead of everywhere

Tick Only this session at the top of the Models tab and your picks stay inside that session — models, providers and effort per role. Nothing else changes. In the list, a role that has a session choice is marked (this session). Untick it and the session's choices are deleted; it uses the global ones again. The header shows the model that is really running in this session, so you can always see which one you are paying for.

Effort

/effort sets how hard the model thinks — the levels are a subset of none, minimal, low, medium, high, xhigh, max. Only the levels your provider and model actually support are offered; some models have no medium, some no max. The default is medium.

You can set the effort per role: press e on its line under Settings ▸ Models. Each role shows its own level in brackets, and a role without one inherits the global level. /ultrathink runs just the next mission at maximum.

Seeing the cost

WhereWhat it tells you
The token line under a missionTokens in, out, cached — and the price, if you switched on Show $ cost.
/usageThe whole session, and what each role and subagent cost of it.
Settings ▸ StatsPer day, across every project.

Prices and context windows

The cost display multiplies tokens by the price of the model. Where the price comes from, in this order:

  1. Your own price, set under Settings ▸ Advanced ▸ Pricing. Nemocode never replaces it by itself; x on that line removes it.
  2. Nemocode's built-in price table.
  3. A live price list, which you can pick from.

The context window of a model comes from your own value (/model ▸ w) first, then from what Nemocode measured for the model, then from a guess by model name, and 256k when nothing is known. A custom endpoint uses the window you entered when you connected it as the ceiling for guessed values.

Context and compaction

Every model has a limit. Nemocode tracks how full the context is and summarises the conversation before it runs out, keeping working. Two settings shape that: the context budget (how much it may use before compacting) and compaction ask (whether it checks with you first). Everything about it is on Context & cost.

/context shows where the tokens are going — the conversation, the files, the memory, the tool results.

What Nemocode does to keep it cheap

The last three are switchable under Settings, and are on by default.