Context & cost

A long mission reads far more than fits in one context window. What Nemocode keeps, what it drops, and when it does the work of shortening — that is the difference between a mission that finishes and one that runs out of room halfway through. These are the parts that manage it, and the choice each one leaves you.

Why this exists at all

Every call sends the whole conversation again. Providers charge far less for the part they have already seen — the cached prefix — than for fresh text. Two consequences run through everything below:

So "the context is getting long, drop something" is the wrong instinct. The question is whether dropping pays for itself before the mission ends.

CAST — context trimming

CAST replaces old tool output with a short placeholder — the file is named, the content is gone. It only does so when carrying that output is more expensive than the shortening costs, and it weighs both sides before it acts:

Once the provider's cache has run out — you paused for longer than the cache TTL below — the cached prefix is gone anyway, so trimming then costs nothing extra and does not wait for the arithmetic.

What it never does

The four settings

Settings ▸ Advanced ▸ CAST context trimming cycles through them. The default is balanced.

SettingWhat it means for your mission
offNo arithmetic at all. Content that has become wrong is still removed, and the hard window limit still applies — a full context window is not a pricing question. Choose this if you would rather decide for yourself, or do not trust the price list your provider publishes.
carefulShortens only when it clearly pays: it assumes the mission is nearly over, wants twice the benefit before it acts and counts the risk of a re-read one and a half times. Suited to long, concentrated work on a handful of files that you keep coming back to.
balancedThe default. Assumes a mission of typical length and takes any advantage that is real.
eagerKeeps the conversation small: assumes many calls are still to come and counts the risk of a re-read only half, so it shortens earlier and deeper. Suited to wide exploration — many files read once and not needed again.

The three active modes run the same reasoning. They differ in how long a mission they expect, how much benefit they demand before acting and how heavily they weigh a re-read — not in what they are allowed to remove. off skips the reasoning altogether. Enter on the line steps to the next setting.

GPC — pre-compression

When you continue a session, everything said before it comes along. GPC shortens that carried conversation before the new mission starts, rather than in the middle of it. The difference matters: shortening mid-mission throws away the cached prefix exactly when the agent is busy, while doing it at the start costs nothing that was not going to be paid anyway.

SettingBehaviour
onEvery mission starts from a shortened history.
offThe default. The conversation is carried as it is.
smartDecided per mission: a small model looks at your latest messages and judges whether the new task depends on the detail of the old ones.

Set it under Settings ▸ Advanced ▸ GPC pre-compression. Shortening rewrites the history, and what was condensed is not read again, which is why it is off unless you choose it. The models that do the work are the roles gpc and gpc_gate under Settings ▸ Models; point them at something cheap — otherwise you pay for the history once more just to make it shorter.

What survives is the substance: decisions, constraints, what was tried and what it did. What goes is the volume — long outputs, repetition, the road not taken.

Compaction

When a session reaches 75 % of its context budget, Nemocode writes a summary of the conversation so far and continues from it; the newest part of the conversation stays word for word. The summary keeps the thread of the work — the goal, the decisions, the state of the files — and you can read it. It is written by the compactor role, which you can give its own model under Settings ▸ Models.

/compact does it now, without waiting for the limit. While a mission is running it says so and leaves the mission alone: it compacts by itself if the conversation grows too big.

SettingWhereWhat it does
Compaction askGeneral; per project under Projectoff compacts silently, always asks first, smart (the default) asks only when the compactor itself is unsure. A project can override the global choice.
Compaction max cutAdvancedHow much of the conversation one compaction may remove in a single step: 10 to 100 %, default 50 %.
Compaction cooldownAdvancedRounds that must pass after a compaction before the next one may start. Default 3.

/context shows where the tokens are going before you decide. It splits them in two groups: what is paid once because it sits in the cached prefix (system rules, paths and scopes, mode and rules, the skill index, tool schemas, the conversation), and what is paid again with every message (file changes, session memory, preloaded files, the memory bundle, your message).

Context budget

How much of the model's window Nemocode is willing to fill before it starts managing it. Settings ▸ Advanced ▸ Context budget steps through presets from 8k to 1M tokens, or Enter lets you type any value (below 4k counts as auto) — models differ, and so do jobs. The default, auto, is derived from the model's own window: the window minus a reserve for the reply, at most 20k and at most a quarter of the window. Compaction starts at 75 % of the budget.

The line shows a bar: green up to the point where compaction starts, amber up to the budget, grey for the rest of the model's window. Set the budget lower to keep responses cheap and quick, higher to let long missions run without interruption. A project can have its own budget under Settings ▸ Project. Whether the context percentage in the header counts against your budget or against the model's window is a Display setting (Context % of budget).

Predictive preload

Experimental, and off by default (Settings ▸ General ▸ Predictive preload). Nemocode watches what the agent reads over several missions of a project. What it keeps returning to — a project memory fact it looks up again and again, a file it opens in mission after mission — is handed over at the start of the next mission, so the agent saves the lookup and the round trip it would have cost. Files come as a short map with one-line hints, not as content that could go stale; the few files that are read and almost never edited may come with their content.

It pays off when your habits are stable and adds noise when they are not. The technical lines it writes to the transcript are hidden unless you switch on Preload / restore lines under Display.

Memory lookups

With PUM loading on map (the default, Settings ▸ General), a mission starts with the standing rules of the project and a list of topics. Everything else the agent looks up itself, with the memory tool, whenever the task calls for it — so what it knows is something it asked for. On full the whole memory is loaded upfront instead, which costs more on every call. The search runs on your machine, on an embedding model you choose in the settings — nothing about your project is sent anywhere for it.

Re-read dedup

If the agent reads the same unchanged file twice, the second copy is not added again — a short note says the file is unchanged and already above. A long mission touches the same handful of files repeatedly; without this, the same content accumulates in the window several times over. Settings ▸ General ▸ Re-read dedup, on by default.

RTK token saver

Long shell output — test runs, build logs, directory listings — is mostly noise to the model. With the RTK helper on your PATH, Nemocode compresses it before the model sees it. Settings ▸ General ▸ RTK token saver, on by default; the installer brings the helper, and without it nothing is compressed. It only handles the common developer commands it knows, and never long-running ones such as servers or watchers.

Lazy tools

Every tool offered to the model costs tokens on every call, used or not. With Lazy tools on (Settings ▸ Advanced), the rarely needed tools appear only as a one-line index, and the agent loads a tool's full description when it decides to use it. Off by default: all tools are offered upfront.

Prompt cache TTL

How long the provider keeps the conversation warm between calls. A longer window survives you going for coffee — the next message is still cheap. It makes the first write more expensive (twice the normal input price instead of 1.25 times), so it pays off when you work in bursts with pauses, and costs when you work in one continuous run. Settings ▸ Advanced ▸ Prompt cache TTL, 5m by default. Only Anthropic models use it; other providers manage their own caching.

Watching it work

The session header shows how full the context is. When Nemocode compacts, it says so in the conversation. When CAST hides old tool output, the Context-trim notice (Settings ▸ Display, off by default) says so and why — switch it on if you want to see every decision; if one looks wrong to you, the CAST setting is where you change it.