Routing and spend
Choose an agent, inspect why a model was selected, and distinguish measured API usage from activity DeckSpace cannot price.
Agent panes use the CLI accounts you connect. Built-in API features use the provider settings in DeckSpace. Those paths have different usage visibility; the app’s total is not your provider’s invoice.
Choose the right control
| Control | What it changes |
|---|---|
| Explicit agent and model | Your chosen provider and supported model argument for that launch. |
| Auto routing | A model and effort recommendation, using task shape, routing mode and available provider information. |
| Team plan gate | Per-task provider, model and effort before you approve the workers. |
| Fast Mode | The registered CLI’s fast-model arguments, where its definition supplies them. |
| Agent run cap | The number of launches allowed through the normal frontend launch path in this app session. |
| Resources worker limit | How many workers may run concurrently. |
| Monthly API spend cap | Admission to the covered metered API paths, using the local usage ledger. |
Auto routing is connected
The composer’s Auto choice reaches the shared routed launcher. An explicit usable agent choice remains authoritative. The router can use a recent provider scan to avoid a known unavailable or signed-out CLI; missing or stale information is not proof that an account works.
Choose Economy, Balanced or Quality in Routing settings. The mode influences model and reasoning effort for quick edits, builds, reviews and architecture work. The first-run wizard asks for the preference; Quality is the fallback when no choice is saved. Economy prefers cheaper configured models and does not guarantee free usage or automatically select a local model.
The route includes a reason. If the optional routing call fails or takes longer than 20 seconds, the launcher uses its local task-shape rules. Inspect the agent/model actually shown on the pane, especially after a provider substitution.
Model and Fast support varies by CLI
A model pin is applied only when that agent definition has a model argument. Fast Mode likewise requires a fast argument. The built-in definitions provide model pins for Claude, Codex, Antigravity, legacy Gemini, OpenCode and Grok. Fast arguments are present for Claude, Antigravity and legacy Gemini. Other built-in entries do not gain a model or Fast flag merely because a launch preference exists.
Custom agents define their own supported arguments. Use Setup & Support to inspect the installed version and run an explicit diagnostic request. A saved argument definition cannot establish that a newly installed CLI version accepts it.
Launch effort and session commands
The launch-time Std / High / Ultra preference exports MAX_THINKING_TOKENS for Claude: Std leaves its default, High supplies 16000, and Ultra supplies 31999. Other CLIs do not receive that Claude-specific variable. Auto routing can choose effort for a routed task; explicit launch overrides take priority.
Running Claude panes also expose model and effort commands for the existing conversation. Changing a live session is separate from choosing defaults for the next launch. Unsupported controls should not be treated as evidence that the provider changed its settings.
Estimates, recorded usage and unknown cost
The plan gate’s approximate dollar figure uses local model-rate and workload assumptions. It is planning information, not a charge, reservation or provider quote. External CLI usage cannot be read reliably across the process boundary, and the estimate is not reconciled to that CLI’s final bill.
Routing settings separately show the monthly API ledger. Its unknown-rate, unmetered and incomplete-call counts matter: unknown cost is not zero, and a failed or interrupted request may still have been billed. Your provider account remains the source for billing and account limits.
The cap refuses new covered metered calls in chat, inline completion and the provider-registry path when the measured ceiling is reached. It does not stop an in-flight call. Atlas calls and provider embeddings are measured but not refused by this cap. Agent CLI panes, hosted Memory indexing, voice/cloud speech and diagnostic CLI calls have additional measurement or enforcement gaps listed in the Routing screen. A cap here cannot limit your entire provider account.
Launch count and concurrent workers
The status bar’s run cap counts launches, not tokens or dollars. Its range is 0–500, with 0 meaning unlimited. The counter resets when the app quits. Normal frontend launches, including Team worker launches, use it; the headless Iterate driver is a separate native path. A relaunch counts again.
Use Resources for concurrency. The automatic worker limit is 2 on machines with up to 8 GiB of RAM (or unknown capacity), 4 above 8 through 16 GiB, and 6 above 16 GiB. You can set an explicit limit from 1–12. Team also keeps its requested run limit and the pane capacity; the applicable lower limit governs admission. Increasing concurrency does not reduce a provider’s token cost.
Iterate token budgets use ledger deltas
The native mission judge reads a delta of recorded input and output tokens from the local ledger. It no longer treats the iteration number as a token count. A configured positive budget can stop later iterations once the measured delta reaches it; zero means unlimited.
This is an imperfect measurement boundary: concurrent recorded API work contributes to the same ledger delta, while unmetered CLI usage does not provide a token count. The baseline is held in the running app and is re-established after restart, with that limitation surfaced on the mission. Do not use this field as an account-wide spending guarantee. Iteration count and deadline remain separate limits.
A practical setup
- Confirm the intended CLI/account in Setup & Support and make an explicit first request.
- Choose a routing mode and inspect the actual route reason and model.
- Review Team’s task assignments and checks before approving the plan.
- Set concurrency in Resources and a launch cap for the paths you use.
- Set the monthly API cap if useful, then read its exclusions and compare with your provider’s dashboard.