Skip to content

Configuration

Itervox is configured entirely through a single WORKFLOW.md file in your project directory. The file has two parts: a YAML front matter block (between --- delimiters) that controls all runtime behaviour, and a Liquid template body that becomes the prompt sent to the agent for each issue.

Run itervox init in your project directory to generate a starter file for Linear or GitHub.


---
itervox_schema_version: 2
tracker:
kind: github
api_key: $GITHUB_TOKEN
project_slug: owner/repo
active_states: ["todo"]
completion_state: "in-review"
polling:
interval_ms: 60000
agent:
max_turns: 60
max_concurrent_agents: 3
turn_timeout_ms: 3600000
workspace:
root: ~/.itervox/workspaces/my-project
hooks:
after_create: |
git clone git@github.com:owner/repo.git .
before_run: |
git fetch origin main && git checkout main && git reset --hard origin/main
server:
port: 8090
---
You are an expert engineer working on the codebase.
## Your issue
**{{ issue.identifier }}: {{ issue.title }}**
{% if issue.description %}
{{ issue.description }}
{% endif %}
Issue URL: {{ issue.url }}
{% if issue.comments %}
## Comments
{% for comment in issue.comments %}
**{{ comment.author_name }}**: {{ comment.body }}
{% endfor %}
{% endif %}
## Steps
1. Explore the codebase relevant to this issue.
2. Create a branch: `git checkout -b {{ issue.branch_name | default: issue.identifier | downcase }}`
3. Implement the change.
4. Run tests and lint.
5. Commit and open a PR.

Everything after the closing --- is the Liquid prompt template. See Liquid template variables for all available variables.


Controls which issue tracker Itervox polls and how issue states are mapped.

FieldTypeDefaultDescription
kindstringrequiredTracker backend. "linear" or "github".
endpointstringhttps://api.linear.app/graphqlGraphQL endpoint (Linear only). Omit for GitHub — the client uses the GitHub API automatically.
api_keystringrequiredAPI token. Use $VAR_NAME to read from an environment variable (recommended).
project_slugstringrequiredFor Linear: the project slug from your Linear URL (e.g. "ENG"). For GitHub: "owner/repo".
active_statesstring[]["Todo", "In Progress"]Issue states Itervox actively polls and dispatches agents for.
terminal_statesstring[]["Closed", "Cancelled", "Canceled", "Duplicate", "Done"]Issue states that are complete. Issues in these states are not re-dispatched.
working_statestring"In Progress"State Itervox transitions an issue to when an agent is dispatched. Empty string = no transition.
completion_statestring""State Itervox transitions an issue to after the agent finishes successfully (e.g. "In Review"). Empty string = no transition. When set, the issue leaves active_states so it is not re-dispatched.
backlog_statesstring[]["Backlog"] (Linear), [] (GitHub)States shown as the leftmost column(s) on the Kanban board. Issues in backlog states are displayed but not dispatched.
failed_statestring""State to move an issue to when agent.max_retries is exhausted. When empty, issues are paused in the dashboard instead of transitioning.
outboxbooltrueWrite-ahead outbox for tracker state transitions and comments. Writes are persisted durably to .itervox/outbox.json and flushed by an independent worker, instead of being made synchronously from the orchestrator’s completion/failed-state paths — so a tracker outage or rate-limit stall cannot lose a transition. Set false as a kill switch to restore synchronous behaviour. With outbox: false, input-required questions and replies are posted with a single direct attempt instead of being queued, so they are not durable and not ordered, and replies are matched to the most recent question by position. Operator comments posted from the dashboard (POST /api/v1/issues/{id}/comment) also go through the outbox when it is enabled; with outbox: false they are posted directly, and entries left from an earlier outbox-on run stay visible with Retry/Discard but are delivered only once the outbox is enabled again. Load-time only (no runtime setter).

Outbox visibility: pending and degraded entries appear in the dashboard’s Outbox panel and LiveOps tile, each with Retry and Discard controls. An entry enqueued with no observed from-state baseline (currently only the issue-discard path) is exempt from supersede-reconciliation, so Discard is the operator remedy for a stuck entry.

Comments are never posted twice. Each queued comment carries a unique key. If an attempt fails, the next retry first asks the tracker whether that comment already landed, and marks the entry delivered instead of re-posting it. If the tracker cannot answer, the entry waits for the next retry rather than posting blind. On Linear the key is the comment’s own id; on GitHub it is a hidden marker in the comment body.

Rate-limited entries are not failures. When a write is rejected by a rate limit, the entry waits until the tracker’s published reset (plus a few seconds of jitter) and shows an amber rate limited until HH:MM chip. These deferrals do not count toward the degraded badge, which is reserved for writes that are genuinely failing.

GitHub note: GitHub Issues does not have built-in workflow states. Itervox simulates states using labels. Create labels in your repository settings that match the values you configure (e.g. create a label named todo, another named in-review).

# Linear example
tracker:
kind: linear
api_key: $LINEAR_API_KEY
project_slug: ENG
active_states: ["Todo", "In Progress"]
terminal_states: ["Done", "Cancelled", "Duplicate"]
completion_state: "In Review"
backlog_states: ["Backlog"]
failed_state: "Failed"
# GitHub example
tracker:
kind: github
api_key: $GITHUB_TOKEN
project_slug: owner/repo
active_states: ["todo"]
terminal_states: ["done", "cancelled"]
completion_state: "in-review"

Controls how frequently Itervox checks the tracker for new or updated issues.

FieldTypeDefaultDescription
interval_msint30000Polling interval in milliseconds.
rate_limit_reserve_percentint10Share of the tracker’s request budget held back for writes. Set 0 to disable.
polling:
interval_ms: 60000 # check every minute
rate_limit_reserve_percent: 10

Reads and writes draw on one tracker budget, and the polling loops scale with the number of stuck issues — so on a busy project polling could exhaust the hourly budget early, after which every state transition, comment and dependency audit failed. Unsticking an issue requires a write, so the reads starved the very operations that would have drained the queue.

When the remaining budget falls into rate_limit_reserve_percent, Itervox stops spending requests on polling reads for that tick — the candidate poll and the input-required reply check — while continuing to admit writes and the input-resume fetch that lets a queued resume proceed. Dispatch pauses until the budget resets rather than failing every write in the meantime, and the daemon logs a warning naming the remaining budget and reset time so a stalled board is never mistaken for a hung daemon.

Adapters that do not report rate-limit counters are unaffected: the check fails open and polls as before.

When a tracker actually rate-limits a request, Itervox records the window once and every caller — agents, the poller, and the outbox — respects it:

  • Linear reports a rate limit as HTTP 400 with a RATELIMITED error code, and publishes the reset time in X-RateLimit-Requests-Reset / X-RateLimit-Complexity-Reset. GitHub uses 429 or 403: Retry-After for secondary limits, and X-RateLimit-Reset once the hourly budget is spent.
  • Writes are admitted ahead of reads when the window lifts. On Linear, where every request is a POST, requests are classified from the GraphQL operation: mutations are writes, queries are reads.
  • If the reset is more than 60 seconds away, calls return a rate-limit error immediately instead of waiting and retrying into a closed window. The outbox holds its writes until the reset and skips flushing until then.
  • A recorded window is capped at 2 hours, so a malformed reset header cannot stall tracker writes indefinitely.

Controls where Itervox creates per-issue working directories and how they are managed.

FieldTypeDefaultDescription
rootstring~/.itervox/workspaces/<project>Directory where per-issue workspaces are created, namespaced per project by default. Supports ~ expansion and $VAR_NAME env var references.
auto_clearboolfalseWhen true, the workspace directory is deleted only when the issue reaches a terminal tracker state — completion_state after success, or failed_state after retries are exhausted. The workspace persists across retries, input-required pauses, stalls, and pipeline mid-states so chained profiles can share .itervox/handoff/ files on the same branch. Logs live in a separate dir and are unaffected. Compatible with agent.auto_review — the clear is deferred until after the reviewer also completes (v0.2.0 breaking change from the legacy “clear after every successful run” semantics).
worktreeboolfalseWhen true, Itervox uses git worktree to create per-issue branches inside a shared clone at root instead of making separate empty directories. Requires a git repository to already exist at root (or clone_url to be set).
clone_urlstring""Git remote URL used to initialise the bare clone when worktree: true and the root directory does not yet contain a git repository.
base_branchstring"main"Branch that new worktrees are created from when worktree: true.
workspace:
root: ~/.itervox/workspaces/my-project
auto_clear: true
worktree: true
clone_url: git@github.com:owner/repo.git
base_branch: main

When multiple profiles run on the same issue — for example a research profile followed by an implementer followed by a reviewer — they can share structured deliverables via the per-issue workspace’s .itervox/handoff/ directory.

How it works:

  1. The orchestrator stamps each worker run with an ISO8601 timestamp and computes a canonical handoff path: .itervox/handoff/<timestamp>_<profile-name>.md. These two values are appended to the worker’s prompt as a ## Run Context block (fields run.timestamp and run.handoff_path).
  2. Before dispatching any worker, the orchestrator reads every existing .md file in that directory (including .partial.md files), sorts them by filename (chronological because of the ISO8601 prefix), and inlines them into the prompt as a ## Prior Agent Handoffs block. A 30 KB token budget caps the section; if exceeded, the oldest files are dropped with a [earlier handoffs truncated] marker.
  3. The agent’s INSTRUCTIONS.md “Handoff Protocol” section tells it to read the prerendered prior-handoffs block and write its own deliverable to run.handoff_path before exiting. The orchestrator does not call out to the agent — it relies on the agent following INSTRUCTIONS.md.
  4. If the worker exits with TerminalFailed or TerminalStalled, the orchestrator renames the most recent matching <timestamp>_<profile>.md to <timestamp>_<profile>.partial.md. Subsequent agents see partials in their handoff context and can distinguish them from clean deliverables. TerminalInputRequired does not mark partial — the agent intentionally paused and may resume.

Git policy: .itervox/handoff/** is committable. itervox init and itervox init --update patch the root .gitignore to whitelist it alongside .itervox/agents/**. Commit the pipeline trail into PRs so reviewers can read the chain.

Filenames are deterministic from run.timestamp and the profile name:

.itervox/handoff/2026-05-26T14-30-45Z_researcher.md
.itervox/handoff/2026-05-26T14-42-12Z_implementer.md
.itervox/handoff/2026-05-26T14-58-30Z_reviewer.md

Profile names with spaces are slugified ("story writer" → story-writer). An empty profile name falls back to agent.

See the Agent Handoff guide for a worked example chaining three profiles end-to-end.


Controls the agent runner: which CLI to invoke, concurrency limits, timeouts, retry behaviour, and advanced features like SSH dispatch and named profiles.

FieldTypeDefaultDescription
commandstring"claude"CLI command used to launch the agent. Can include flags (e.g. "claude --model claude-opus-4-6").
backendstring""Explicitly sets the runner backend ("claude" or "codex"). Only needed when command is a wrapper script and Itervox cannot infer the backend from the command name. A backend that disagrees with a command whose binary is claude or codex (this field, a profile backend, a per-issue backend pin, or a rate_limited switch_to_backend) is refused: the command keeps its own backend and the daemon logs a warning.
max_concurrent_agentsint10Maximum number of agent workers running simultaneously across all issues.
max_concurrent_agents_by_statemap{}Per-state concurrency limits that override max_concurrent_agents. Keys are lowercase state names. See example below.
max_automation_queue_lengthint100Maximum durable automation dispatch entries waiting for capacity or dependency resolution. 0/negative values fall back to the default; the queue is never unlimited.
max_retriesint5Maximum retry attempts before an issue is moved to tracker.failed_state (or paused if failed_state is empty). 0 means unlimited retries. When the failed turn carried a vendor limit signal, the retry waits for the vendor’s reset time or retry delay (at most 6 hours) if that is longer than the normal back-off.
max_retry_backoff_msint300000Cap on exponential retry back-off (5 minutes). Back-off progresses as 10s × 2^(attempt-1), capped at this value. 0/negative values fall back to the default; use max_retries to control retry count.
max_turnsint20Maximum number of agent turns per session.

These retry and rate-limit settings are also editable from the Settings dashboard:

Retries settings showing max retries per issue, on-exhausted-retries behavior, and rate-limit switch cap
FieldTypeDefaultDescription
turn_timeout_msint3600000Hard wall-clock limit for an entire agent session (all turns combined). When exceeded, the subprocess is killed and the issue is retried. 0 disables the timeout.
read_timeout_msint30000Per-read timeout on the subprocess stdout pipe (30 seconds). If no bytes arrive within this window, the subprocess is killed. Catches OS-level pipe hangs before the stall detector fires.
stall_timeout_msint300000Orchestrator-level inactivity timeout (5 minutes). If no SSE events are produced within this window, the worker context is cancelled and the issue is retried. Operates on the parsed event stream and detects semantic stalls (e.g. agent looping without progress). Set to 0 or less to disable.
deps_analyzer_timeout_msint600000Wall-clock limit for one dependency-analyzer job end to end, across all chunks. Matches the dashboard’s 10-minute poll deadline. ≤ 0 falls back to the default.
deps_analyzer_chunk_sizeint75Maximum issues sent to the agent in one analyzer turn. Larger backlogs are split into sequential chunks; relations spanning two chunks are not examined (an accepted blind spot, logged at analysis time). Raise it for full-graph fidelity on a larger backlog if you can tolerate a longer, costlier turn. ≤ 0 falls back to the default.
dependency_audit_refresh_interval_msint600000How often the off-loop dependency audit refreshes blocker state from the tracker. Startup-only — not runtime-editable. See the rate-limit note below before lowering.
dependency_audit_refresh_timeout_msint30000Bounds a single off-loop refresh batch. Startup-only.
dependency_audit_refresh_batch_sizeint100Caps how many audit rows one batch may fetch. Startup-only.
FieldTypeDefaultDescription
inline_inputboolfalseWhen an agent needs human input, its question is always posted as a comment on the tracker issue, and a comment on the issue resumes the agent — normally in the same session; after a daemon restart that had to rebuild the entry from tracker comments, a fresh session starts with the question and your reply as context. false (default): the dashboard also offers a reply box. true: the tracker is the only place to reply — the dashboard reply box is hidden and POST /api/v1/issues/{id}/provide-input returns 409 inline_input_enabled. Automation replies (itervox action provide-input) are unaffected. A reply written before the agent’s question has actually reached the tracker still counts: with the outbox enabled, any comment created after the question was queued resumes the agent. Runtime-editable from Settings → General.
base_branchstring""Remote branch used as the base for git diffs when enriching PR context (e.g. "origin/develop"). When empty, Itervox auto-detects via git symbolic-ref refs/remotes/origin/HEAD, falling back to "origin/main".
reviewer_promptstringbuilt-in(Deprecated) Liquid template for the legacy reviewer. Prefer reviewer_profile.
reviewer_profilestring""Name of the agent profile used for code review. When set, the reviewer runs as a regular worker using this profile’s command, backend, and profile files. Enables the AI Review button in the dashboard.
reviewer_profilesstring[][]Ordered list of reviewer profiles for multi-reviewer fan-out. With more than one entry, each reviewer runs sequentially and independently over the same issue and records a verdict, combined per review_quorum. Empty falls back to reviewer_profile.
review_quorumstring"any_block"How reviewer verdicts combine: any_block (one block blocks), majority (strictly more than half), or unanimous (all must block). A reviewer that records no parseable verdict counts as a block.
auto_reviewboolfalseWhen true, automatically dispatches a reviewer worker after each successful agent run. Requires reviewer_profile to be set. As of v0.2.0, this safely coexists with workspace.auto_clear — the clear is deferred until the reviewer also completes.
max_switches_per_issue_per_windowint2Maximum times a rate_limited automation can switch an issue to a different profile within the rolling window. 0 for unlimited. Also counts backend_fallback switches of implementer runs. This cap limits switching; it is not the trigger condition.
switch_window_hoursint6Rolling window size (in hours) for the switch cap. The switch history, cooldowns and cap-comment dedupe survive a daemon restart.
switch_revert_hoursint0TTL (in hours) after which an auto-applied profile/backend switch is reverted on the next poll cycle, returning the issue to its original profile and backend. 0 (default) disables the revert. Operator-set overrides survive — only auto-switches with a recorded AutoSwitchedAt timestamp are eligible.
backend_fallbackmapabsentDeclarative reroute when a backend hits its usage limit, with reset-time switch-back and a minimum dwell. See “Backend health and fallback” below. Read at load time; not runtime-editable.
rate_limit_error_patternsstring[][]Custom case-insensitive substrings for detecting rate-limit errors, matched verbatim against the whole failure text (agent-reported failure and CLI stderr). Empty falls back to the built-in defaults: anchored 429 forms (http 429, status: 429, (429), 429 too many requests), anchored quota forms (insufficient_quota, quota exceeded for, usage quota), rate_limit_exceeded, rate limit, too many requests and the Claude/Codex usage-limit phrases. A bare 429 counts only in the agent-reported failure, never in CLI stderr; a bare quota never counts.
allow_unchecked_mergeboolfalseWhen false (default), the merge_pr action refuses to merge on a repo with zero required checks configured (reason unarmed_gate:...) instead of merging with no CI coverage. Set true to merge anyway; the daemon still logs a loud warning.
agent:
reviewer_profile: code-reviewer # use the "code-reviewer" profile for AI reviews
auto_review: true # automatically review after each successful run
profiles:
code-reviewer:
command: claude --model claude-opus-4-6
soul_file: .itervox/agents/code-reviewer/SOUL.md
instructions_file: .itervox/agents/code-reviewer/INSTRUCTIONS.md

Backend health and fallback (agent.backend_fallback)

Section titled “Backend health and fallback (agent.backend_fallback)”

Backend circuit breaker. Itervox keeps one breaker per agent backend and worker host (claude, codex, claude@build-1, …): a limit seen on SSH host A says nothing about local runs or host B, which may use other credentials. A breaker opens when:

  • a run stops on a usage limit (Claude rate_limit_event rejected, a result with API status 429/402, a Codex “You’ve hit your usage limit”): limited until the vendor’s published reset, or for default_cooldown_minutes (15) when no reset is published; or
  • Claude reports api_retry for rate_limit/overloaded three times within 5 minutes on that backend and host: limited for 5 minutes (or the vendor’s retry delay, when longer). Fewer retries only mark it warning.

A breaker never holds longer than its source can justify: a rate_limit_event reset is capped at 5h15m for the five_hour window, 7 days

  • 1h for seven_day* windows and 24h for any other window; a reset read from text (Codex, or a Claude result message) at 6h; an api_retry throttle at 1h; an unknown reset uses default_cooldown_minutes (at most 1440). A further reset is capped and logged, and a persisted value is capped again on load. A zone-less text reset from an SSH worker (Codex prints the host’s local time with no zone) is treated as unknown, because the daemon cannot tell the host’s zone; the cooldown and one probe re-learn it. To close a breaker early, use Clear on the dashboard chip or POST /api/v1/backend-health/clear (see the API reference).

While a breaker is open, no issue is started on that backend and host. Workers, retries, reviewers, automation runs and input-required resumes are all checked. An issue that cannot run anywhere is held with the backend_limited ineligible reason (shown as the issue’s “why idle” reason), never paused, and consumes no retry. Queued automations stay queued with backend_limited. When the reset passes, the breaker half-opens: exactly one issue is dispatched as a probe, and the others wait. A probe that succeeds (or stops to ask a question) closes the breaker. A probe that hits the limit again reopens it. Open breakers survive a restart (backend_health.json in the daemon log directory, next to auto_switched.json; runtime state, never commit it).

Declarative fallback. With agent.backend_fallback, a limited target is rerouted instead of held:

agent:
backend_fallback:
chain: [claude, codex] # preference order
profile_map: # counterpart profile per backend
coder: { codex: coder-codex }
reviewer: { codex: reviewer-codex }
default: { codex: coder-codex } # issues running agent.command
on_unmapped: hold # hold | backend_hint
default_cooldown_minutes: 15 # when the vendor publishes no reset
min_dwell_minutes: 30
switch_back: at_reset # at_reset | on_success | manual
FieldTypeDefaultDescription
enabledbooltrue when the block is presentfalse keeps the block but turns rerouting off. Without the block there is no rerouting, but the breaker still holds issues on a limited backend. The whole block may also be a boolean: backend_fallback: false is off, true is on with the defaults. Any other shape ("off", a number, a list, enabled: "false") fails config loading
chain[]string[claude, codex]Backends tried in order when the target is limited. Known, distinct backends only
profile_mapmap{}source_profile: {backend: target_profile}. default is the key for issues that run agent.command. Every target must exist, be enabled and actually run that backend (its command’s binary, else its backend: for a wrapper). The reverse direction is implied: a target profile maps back through its source row
on_unmappedstringholdFor a profile with no mapping: hold (wait with backend_limited) or backend_hint (request the chain backend for the issue’s own command; works only for wrapper commands — a claude ... command is never paired with codex)
default_cooldown_minutesint15How long a breaker stays limited when the vendor published no reset time (1–1440). Also used by the breaker when the block is absent
min_dwell_minutesint30Minimum time a switched issue stays on the fallback before it may switch back. It never blocks a further hop when the fallback itself becomes limited
switch_backstringat_resetat_reset: the override is cleared once the dwell has passed AND the original backend is no longer limited (its reset, or the cooldown, has passed); the next dispatch goes back through the breaker. on_success: cleared by the first successful run after the dwell. manual: never cleared automatically (pin a backend, or use switch_revert_hours)

The block is read at load time: edit WORKFLOW.md to change it (the normal reload applies it). There is no settings API for it.

A switched implementer issue keeps its new profile and backend for later runs (a sticky override, shown on the card as codex (auto) and in the snapshot’s autoSwitches). Before, a success cleared the override, so the next dispatch went back to the still-limited backend. Each fallback switch counts against max_switches_per_issue_per_window. With the cap spent, the issue is held rather than rerouted. Reviewer and automation runs are rerouted per run and leave the issue’s profile alone.

Precedence (fallback vs pins vs rate_limited rules):

  1. An operator’s per-issue backend pin (dashboard issue detail, or POST /api/v1/issues/{identifier}/backend) wins. A pinned issue is never rerouted: when its pinned backend is limited it is held with backend_limited. A pin is refused up front (409 backend_pin_refused) when the issue’s command runs the other backend.
  2. agent.backend_fallback, for profiles it maps (or every profile with on_unmapped: backend_hint). When such a run hits a usage limit, the fallback alone handles it: the issue is rerouted at once (no retry consumed) or held, and rate_limited automations are not evaluated for that exit.
  3. rate_limited automations, for everything else: unmapped profiles, pinned issues, and failures classified from text at retry exhaustion. They keep their own cap, cooldown and switch target. A switch target on a limited backend is queued with backend_limited rather than started.

When both are configured, use profile_map for the common Claude ⇄ Codex pairs and keep rate_limited rules for custom mappings or helper runs.

Itervox can distribute agent work across multiple remote hosts via SSH. Each host runs the agent CLI in a separate SSH session.

FieldTypeDefaultDescription
ssh_hostsstring[][]List of SSH hosts in "host" or "host:port" format. When empty, agents run locally.
ssh_host_descriptionsobject{}Optional display labels for ssh_hosts. Keys are host strings, values are user-facing descriptions shown in the dashboard and TUI.
ssh_strict_host_checkingstring"accept-new"Default StrictHostKeyChecking mode applied to every SSH worker connection. Valid values: accept-new (TOFU — pin on first contact, reject on mismatch), yes (strict — reject any unknown or changed key), no / off (permissive — accept any key, insecure), ask (prompt — incompatible with BatchMode=yes). Defaults to accept-new so a brand-new host’s key is recorded in ~/.ssh/known_hosts on first contact and any subsequent mismatch is rejected.
ssh_strict_host_by_hostobject{}Per-host override for StrictHostKeyChecking, taking precedence over ssh_strict_host_checking. Keys are host addresses (matching entries in ssh_hosts), values use the same set as the default. Use to harden production hosts to yes or temporarily relax a sandbox VM to no.
dispatch_strategystring"round-robin"How issues are routed to SSH hosts. "round-robin" cycles through hosts in order. "least-loaded" sends each new issue to the host with the fewest active workers. Ignored when ssh_hosts is empty.
agent:
ssh_hosts:
- build-host-1.example.com
- build-host-2.example.com:2222
dispatch_strategy: least-loaded
# Default to TOFU for all hosts; harden production specifically.
ssh_strict_host_checking: accept-new
ssh_strict_host_by_host:
"build-host-1.example.com": yes

Worker host requirements. Each turn runs as ssh -T <host> bash -lc '<fixed wrapper>', with the script (prompt included) sent on ssh’s stdin, so the worker needs bash 3.2 or later (with process substitution: /dev/fd or a writable TMPDIR) plus the POSIX userland tools sh, cat, printf, wc and sleep — no setsid, ps, head or base64. The prompt reaches the agent CLI on its stdin, never as a command-line argument (claude ... -p, codex exec ... -), so no prompt size is refused on Linux workers (MAX_ARG_STRLEN); Claude Code caps piped input at 10 MB. The remote user’s login shell can be any of sh, bash, zsh, dash, ksh, csh or tcsh; it only passes the wrapper to bash. Any SSH server works (OpenSSH or Dropbear): the wrapper creates its own process group for the agent instead of relying on the server’s. The worker’s login profile must not read stdin: a profile that reads a line or a fixed number of bytes makes the turn fail with exit 97 (“arrived truncated”), and one that reads stdin to end-of-file blocks the turn until it is cancelled.

Cancelled turns stop the remote agent. itervox keeps the ssh connection’s stdin open for the whole turn. When a turn is cancelled, times out, or the daemon shuts down, the local ssh client is killed, sshd closes the session, and the wrapper on the worker sees end-of-file on stdin. It then sends SIGTERM to the process group it created for the agent (the agent and everything it started), and SIGKILL 2 seconds later. Nothing outside that group is signalled — not the SSH server, other sessions on the same connection (Dropbear, ControlMaster), or background jobs your login profile started — and a normally finished turn signals nothing at all. A readonly TMOUT in the login profile does not affect it. A process that moved itself into another process group or session (for example with setsid) is not reached. If the network drops without the client dying, sshd only notices once its ClientAliveInterval or TCP keepalive expires, and the stop happens then.

The available_models field stores the list of models available for each backend. This is auto-populated by itervox init (which queries claude --list-models / codex --list-models) and used by the web dashboard’s profile editor to suggest models in the dropdown.

FieldTypeDefaultDescription
available_modelsmap{}Map of backend name ("claude", "codex") to a list of model options. Each entry has id (model ID string) and label (human-readable name).
agent:
available_models:
claude:
- { id: "claude-haiku-4-5-20251001", label: "Haiku 4.5 - Fast" }
- { id: "claude-sonnet-4-6", label: "Sonnet 4.6 - Balanced" }
- { id: "claude-opus-4-6", label: "Opus 4.6 - Powerful" }
codex:
- { id: "gpt-5.3-codex", label: "GPT-5.3-Codex - Frontier coding" }
- { id: "gpt-5.2-codex", label: "GPT-5.2-Codex - Long-horizon agentic coding" }

If available_models is empty or missing, the dashboard falls back to a built-in default list.


Named profiles let you configure alternative agent commands selectable per-issue from the web dashboard. Each profile must specify a command. In schema 2, profile text lives in files under .itervox/agents/<profile>/; WORKFLOW.md stores references to those files.

FieldTypeDescription
commandstringCLI command for this profile (e.g. "claude --model claude-haiku-4-5-20251001").
soul_filestringPath to SOUL.md, relative to WORKFLOW.md when not absolute. Holds identity, purpose, boundaries, and collaboration style.
instructions_filestringPath to INSTRUCTIONS.md, relative to WORKFLOW.md when not absolute. Holds operational rules, checklists, and done criteria.
backendstringExplicit backend override for this profile (same as the top-level backend field).
enabledbooleanOptional. Disabled profiles stay in config but are hidden from normal selection and dispatch.
allowed_actionsstring[]Optional daemon-backed actions the profile may invoke: comment, comment_pr, create_issue, move_state, provide_input.
create_issue_statestringRequired when allowed_actions includes create_issue; the tracker state/column used for follow-up issues.
permission_modestring"bypass"

SOUL.md is appended before INSTRUCTIONS.md, and both files support the same Liquid bindings as the main WORKFLOW.md prompt. Automation instructions are appended after the selected profile files. agent.profiles.*.prompt is legacy input for itervox init --update; schema 2 rejects it at daemon startup.

The dashboard profile editor edits SOUL.md and INSTRUCTIONS.md separately and writes those files before refreshing the snapshot.

allowed_actions do not grant shell or tracker access by themselves. They only allow the daemon to mint short-lived per-run bearer grants for the corresponding /api/v1/agent-actions/* routes.

agent:
command: claude
max_concurrent_agents: 5
max_turns: 60
turn_timeout_ms: 3600000
read_timeout_ms: 120000
stall_timeout_ms: 300000
max_concurrent_agents_by_state:
"in progress": 3
"in review": 2
profiles:
fast:
command: claude --model claude-haiku-4-5-20251001
soul_file: .itervox/agents/fast/SOUL.md
instructions_file: .itervox/agents/fast/INSTRUCTIONS.md
thorough:
command: claude --model claude-opus-4-6
soul_file: .itervox/agents/thorough/SOUL.md
instructions_file: .itervox/agents/thorough/INSTRUCTIONS.md
allowed_actions: [comment, move_state]
input-responder:
command: claude --model claude-sonnet-4-6
soul_file: .itervox/agents/input-responder/SOUL.md
instructions_file: .itervox/agents/input-responder/INSTRUCTIONS.md
allowed_actions: [comment, provide_input]
qa:
command: claude --model claude-sonnet-4-6
soul_file: .itervox/agents/qa/SOUL.md
instructions_file: .itervox/agents/qa/INSTRUCTIONS.md
allowed_actions: [comment, create_issue, move_state]
create_issue_state: Todo

Automations dispatch a selected profile when a trigger fires, then layer a small instruction block on top of that profile.

Supported triggers:

  • cron
  • input_required
  • tracker_comment_added
  • issue_entered_state
  • issue_moved_to_backlog
  • run_failed
  • pr_opened — fires when a worker’s PR is detected
  • rate_limited — fires when a worker run hits a vendor usage limit. A structured limit reported by the agent (Claude’s rate_limit_event with status rejected, or a result with API status 429/402; a Codex “You’ve hit your usage limit” error) fires it on the first failure without consuming a retry, also with max_retries: 0. Otherwise it fires when the run exhausts its retries and the failure text is classified as rate-limit-driven. The switch cap limits automated switching; it is not the trigger condition. With no eligible fallback, the issue is retried after the vendor’s reset time (at most 6 hours).
  • blockers_resolved — fires when dependency audit observes a previously blocked issue becoming unblocked.

Tracker event triggers are poll-derived, not webhook-derived. The automation loop runs every 15 seconds. tracker_comment_added compares only the latest observed comment, so multiple comments between polls collapse to the latest comment for trigger purposes.

When a trigger cannot start immediately for a retryable runtime reason such as no_slots, per_state_limit, already_running, input_required, pending_input_resume, or blocked_by, Itervox records a durable automation queue entry instead of dropping the attempt. The queue is capped by agent.max_automation_queue_length. Saturation pauses recurring/cron/polled producer intake; one-shot and internal dispatch attempts are rejected and counted for audit rather than paused. Existing queue entries continue draining until the queue falls below the low-water mark.

FieldTypeDescription
idstringStable automation identifier.
enabledboolWhether the automation is active.
profilestringName of the agent profile to dispatch.
instructionsstringSmall Markdown/Liquid instruction overlay appended after the selected profile files.
trigger.typestringTrigger type.
trigger.cronstringFive-field cron expression for cron triggers.
trigger.timezonestringOptional timezone for cron triggers.
trigger.statestringRequired for issue_entered_state; the tracker state that must be entered.
filter.match_modestringHow populated filters combine: all or any.
filter.statesstring[]Issue-state filter. For cron automations, leave empty to use backlog and active states.
filter.states_anystring[]Alias for filter.states; recommended for blockers_resolved examples to make the source-state policy explicit.
filter.labels_anystring[]Match issues with at least one of the listed labels.
filter.identifier_regexstringRegex matched against issue identifiers like ENG-42.
filter.limitintMaximum number of issues to queue from one cron tick or event poll batch.
filter.input_context_regexstringOnly meaningful for input_required; matched against the blocked-agent question text.
filter.max_age_minutesintOnly meaningful for input_required; skips blocked entries older than this many minutes.
policy.auto_resumeboolFor input_required, allows the helper to resume the blocked run via provide_input. For rate_limited, accepted but prefer policy.auto_switch.
policy.auto_switchboolAlias for policy.auto_resume on rate_limited; allows immediate profile/backend switching without a human approval step.
policy.switch_to_profilestringRequired for rate_limited; profile to use for the switched run.
policy.switch_to_backendstringOptional claude/codex backend override for rate_limited switched runs.
policy.cooldown_minutesintOptional cooldown for rate_limited rules on the same issue/profile tuple. Default is 30 when unset.
policy.move_to_statestringOptional for blockers_resolved; allows the helper profile to move matching unblocked issues to this state when the profile includes move_state.

When switch_to_backend is set, the target profile command must be compatible with that backend. Prefer a dedicated Codex profile such as command: codex / backend: codex, or a backend-aware wrapper command.

Itervox exposes tracker blockers to the prompt and dashboard, and normal issue dispatch skips Todo issues whose blockers are still non-terminal. That is the deterministic blocker behavior shipped in v0.2.0.

Automation rules can opt into a deterministic blockers_resolved trigger. Core dependency audit detects when a previously blocked issue has no unresolved blockers left; tracker mutation still happens only through an enabled automation whose selected profile is allowed to use move_state.

automations:
- id: qa-ready
enabled: true
trigger:
type: issue_entered_state
state: "Ready for QA"
profile: qa
instructions: |
Run the QA routine for this issue.
Comment the results.
If any required check fails, move the issue to Todo.
- id: pm-backlog-review
enabled: true
trigger:
type: cron
cron: "0 9 * * 1-5"
timezone: "Asia/Jerusalem"
profile: pm
instructions: |
Review backlog issues for missing clarity and acceptance criteria.
Leave one concise comment summarising what is unclear.
filter:
states: ["Backlog"]
limit: 20
- id: unblock-backlog-to-todo
enabled: true
trigger:
type: blockers_resolved
profile: pm
instructions: |
All tracked blockers for this backlog issue are terminal.
Move only backlog/Backlog issues to Todo.
Do not move review, in-review, PR-open, or merged issues.
filter:
states_any: ["backlog", "Backlog"]
policy:
move_to_state: "Todo"

For the full mental model, trigger semantics, and examples, see the Automations guide. For runtime behavior, see the Automation Queue guide and Dependency Management guide.

The legacy schedules: block is still parsed and silently upgraded to equivalent cron automations at startup — but this fallback is deprecated and will be removed in a future release. Itervox logs a warning on startup when a schedules: block is found, with the count of upgraded entries.

To migrate: rewrite each schedules: entry as an automations: entry with trigger.type: cron plus the same cron expression, timezone, profile, and state filter. The legacy format has no instructions: block, so migrated entries start with an empty prompt overlay and can optionally add instructions at migration time.


Controls the dependency graph: which inferred (LLM-detected) edges are trusted enough to hold dispatch, how eligible issues are ordered, and when a long-blocked issue escalates for attention.

Tracker-declared blockers (issue.BlockedBy) always hard-block dispatch and are not configurable here. This block governs the inferred layer plus ordering.

FieldTypeDefaultDescription
inferred_gatingbooltrueKill switch for the soft gate. When true, inferred edges can hold dispatch just like tracker blockers. Set false to make inferred edges display-only.
confidence_thresholdfloat0.7Minimum analyzer confidence (0.0–1.0) an inferred edge needs to gate. Lower-confidence edges still appear on the dashboard but never block. Out-of-range values fall back to the default.
staleness_hoursint168How long an inferred edge is trusted before it stops gating. Non-positive values fall back to the default.
orderingstring"critical_path"Dispatch ordering strategy. One of critical_path, critical_path_strict, or simple — see Ordering modes. An unrecognized value falls back to the default with a warning.
escalate_blocked_after_hoursint48How long an issue may sit blocked before it surfaces as needing attention. An explicit 0 disables escalation and is preserved as a meaningful value; only a negative value falls back to the default.
analysis_modestring"auto"How the LLM dependency analyzer is triggered. auto: on the scheduler’s debounce/min-interval rules. manual: only via the Deps tab’s Analyze button or POST /api/v1/deps/analyze. The blocker audit is unaffected. Runtime-editable from Settings → Dependencies. auto_analyze (bool) is a deprecated alias: false = manual. Changing the mode from the dashboard leaves the old auto_analyze line in WORKFLOW.md; delete it by hand to stop the “both set” warning on every reload.
stacked_prsboolfalseBranch an issue’s worktree from its single live blocker’s branch. Does not set the PR base — see below.
auto_analyze_min_interval_minutesint60Minimum gap between scheduled analysis passes. Non-positive values fall back to the default — the analyzer must not run every tick.
auto_analyze_debounce_minutesint5Delay after a dispatch-affecting change before analysis starts, so it waits for state to settle rather than racing an in-flight dispatch.
dependencies:
inferred_gating: true
confidence_threshold: 0.7
staleness_hours: 168
ordering: critical_path
escalate_blocked_after_hours: 48
analysis_mode: auto
auto_analyze_min_interval_minutes: 60
auto_analyze_debounce_minutes: 5

Set dependencies.stacked_prs: true to have an issue’s worktree branch from its blocker instead of workspace.base_branch, so the work starts from the blocker’s commits rather than duplicating them.

dependencies:
stacked_prs: true

It applies only when an issue has exactly one live blocker, and that blocker must carry an identifier — without one it cannot name a branch, so itervox declines to stack rather than guess. An unidentified blocker still counts as a live blocker: it is a real dependency, so an issue with one identified and one unidentified live blocker does not stack either. With several live blockers, any choice of base would be arbitrary: the branch would sit on one blocker while still depending on the others, so the PR would not be reviewable in isolation. A blocker that has already reached a terminal state is skipped — its work is in base_branch already.

Stacking is best-effort by design. If the blocker’s branch does not exist in this checkout — not dispatched yet, worked on another machine, worktree cleared — itervox falls back to base_branch and logs the decision. Review ergonomics must never be the reason a dispatch fails.

Restacking. When a blocker reaches a terminal state, itervox replays the dependent’s worktree onto base_branch automatically, so the stack does not drift behind subsequent merges. Three cases are refused rather than forced:

SituationBehaviour
The dependent is runningSkipped. Rebasing a worktree an agent has checked out moves HEAD under a live process; it is picked up on a later cycle instead.
The worktree has uncommitted changesSkipped. That work is unpushed and unbacked-up.
The rebase conflictsAborted — the branch is left byte-identical — and the issue moves to input_required with the reason. A conflict needs an owner; resolving it automatically would be a guess.

An issue that still has another live blocker is never restacked, matching the rule that stacking only happens with exactly one live blocker.

All three modes share the same final tiebreakers (created_at oldest first, then identifier). They differ only in what they weigh before reaching them:

ModeComparison orderUse when
critical_path (default)priority band → fan-out → chain lengthYou want operator-set priority to stay authoritative.
critical_path_strictfan-out → chain length → priority bandYou want throughput across the dependency graph to outrank the priority field.
simplepriority band only (no graph awareness)You want the legacy pre-graph behaviour.

“Fan-out” is how many issues are transitively unblocked by finishing this one; “chain length” is the longest downstream path. Both are computed per tick over the dependency graph, with cycles collapsed so they cannot skew the counts.

The distinction that matters: critical_path applies graph leverage only as a tiebreaker within a single priority band. If your issues carry consistently distinct priorities, the graph metrics are never consulted and the mode behaves like simple. Choose critical_path_strict when you want a blocker that gates a dozen issues to dispatch ahead of an unrelated urgent leaf.

The tradeoff runs both ways. critical_path_strict deliberately overrides an explicit operator signal — an issue marked urgent for a reason outside the graph (a customer escalation, a deadline) will wait behind a high-fan-out lower-priority blocker. Prefer the default unless you are specifically optimising fleet throughput.

Both graph-aware modes degrade to exactly simple’s ordering when the issue set has no dependency edges, so enabling either on a project without blockers changes nothing.

For the full mental model, see the Dependency Management guide.


Shell commands run at lifecycle points in each issue’s workspace. All hooks run in the issue’s workspace directory.

FieldTypeDefaultDescription
after_createstring""Shell command run after the workspace directory is created, before any agent runs. Typically used to clone the repository.
before_runstring""Shell command run once per worker invocation, before the first agent turn of that attempt. Typically used to sync the branch (git fetch, git reset).
after_runstring""Shell command run after each completed agent turn.
after_run_requiredboolfalseWhen true, a worker whose final after_run hook exits non-zero fails the unit instead of completing it — the hook becomes a per-unit completion gate.
before_removestring""Shell command run before the workspace directory is deleted (when workspace.auto_clear: true or on manual removal).
timeout_msint60000Maximum time allowed for any single hook to complete (60 seconds).

All hooks run via bash -lc in the issue workspace. Itervox does not inject any per-issue environment variables into hook processes; hooks inherit the daemon’s environment as-is. before_run is intentionally per-attempt, not per-turn, so setup hooks do not wipe agent progress between turns. On an in-place input_required resume, Itervox skips before_run; if the workspace had to be recreated first, the normal setup hooks run again.

By default after_run failures are logged and ignored. Set hooks.after_run_required: true to turn after_run into a per-unit completion gate: a unit whose final after_run hook fails does not complete, regardless of the agent’s own clean exit. Use it to run make test (or any operator-owned verification) as part of the definition of done, complementing the CI-governed merge gate.

Multi-line shell scripts are supported using YAML block scalars:

hooks:
after_create: |
git clone git@github.com:owner/repo.git .
pnpm install --frozen-lockfile
before_run: |
git fetch origin main
git checkout main
git reset --hard origin/main
after_run: |
git status --short
before_remove: |
tar -czf ../workspace-backup.tgz .
timeout_ms: 120000

Controls the built-in HTTP server that serves the web dashboard and REST API.

FieldTypeDefaultDescription
hoststring"127.0.0.1"Interface the server listens on. Change to "0.0.0.0" to expose on all interfaces. ITERVOX_SERVER_HOST overrides it (see “Bind overrides from the environment”).
portint8090TCP port. 0 = the OS picks a free port (for several daemons on one machine). If the port is in use, startup fails naming the holder. ITERVOX_SERVER_PORT, then PORT, override it.
allow_unauthenticatedboolfalseBy default Itervox requires bearer-token auth on every request, on every bind — including loopback — and auto-generates an ephemeral ITERVOX_API_TOKEN if none is set. Set to true only for trusted, fully local deployments where the daemon is physically unreachable from anyone else — this has an effect on every bind, including loopback, since a loopback daemon behind a tunnel or reverse proxy is exactly as reachable as one bound to 0.0.0.0. Renamed from allow_unauthenticated_lan; the old key still parses and works identically but logs a deprecation warning at startup. State-changing requests from another origin are still refused with 403 (see Authentication below).
allowed_hostslist of strings[]Extra Host header names the dashboard answers when allow_unauthenticated: true (DNS-rebinding protection). Without a token every route refuses a request whose Host is a DNS name other than localhost, the configured host, or one listed here, with 403 host_not_allowed; IP addresses are always accepted. List the reverse-proxy, tunnel, Tailscale MagicDNS or container service hostname you browse to, without scheme (a :port is ignored). Ignored in token mode.
metrics.enabledboolfalseServe Prometheus metrics at GET /metrics (text format) on the dashboard listener. Requires the same bearer token as /api/v1/*; off, /metrics answers 404. Families, labels and semantics: API reference, “Metrics”. Read at startup.
server:
host: 127.0.0.1
port: 8090

Containers need 0.0.0.0 and PaaS platforms (Cloud Run, Heroku, Render, …) inject the port as $PORT, so the bind can come from the environment:

ValuePrecedence, highest first
hostITERVOX_SERVER_HOST → server.host → 127.0.0.1
portITERVOX_SERVER_PORT → PORT → server.port → 8090
  • There is no daemon command-line flag for the host or port.
  • 0 means “the OS picks a free port” in an env var as in server.port; an explicit server.port: 0 still holds when no env var is set.
  • A variable that is set but invalid — empty, not an integer, outside 0–65535, a host with a scheme (http://…), a path, whitespace or a host:port pair — fails startup with an error naming the variable. It is never silently ignored. An IPv6 literal (::, [::1]) or a host name is accepted.
  • Beware a shell that exports PORT for another project: it moves the daemon’s bind. The HTTP server listening log line names the winning source (host_source, port_source).
  • The environment is fixed when the process starts: a WORKFLOW.md reload re-reads the same values, so changing the bind through the environment needs a restart.
  • The host override feeds the same server.host everything else uses, so the unauthenticated-mode Host guard accepts the env host by name exactly as it would the YAML one, and the dashboard URL written to .itervox/dashboard_url uses the bound address. The cross-origin guard compares Origin with the request’s Host, so it is unaffected.

Every bind — including loopback — requires Authorization: Bearer <token> on every HTTP request and SSE stream, unless server.allow_unauthenticated: true is set. Itervox reads the token from the ITERVOX_API_TOKEN environment variable; when unset, the daemon generates a random ephemeral token at startup, installs it, and logs the tokened dashboard URL to stderr once — only when stderr is a terminal or ITERVOX_PRINT_TOKEN=1 (--no-print-token always suppresses it). Headless (systemd, containers, redirected stderr), an auto-generated token is written to <logs-dir>/api-token (mode 0600, rewritten on every start) instead, and the log line carries only its path and a sha256: fingerprint. The dashboard prompts for the token on first load and persists it (session-only by default, or via the “Remember” checkbox in localStorage). The GET /api/v1/health (static liveness) and GET /api/v1/ready (event-loop readiness) probes are auth-exempt so load balancers, uptime monitors and container health checks can use them without a token — they are the only auth-exempt routes. GET /metrics (opt-in via server.metrics.enabled) always needs the token.

Cross-origin guard in unauthenticated mode. With server.allow_unauthenticated: true there is no bearer check, so Itervox instead refuses cross-origin state-changing requests, which stops a malicious web page in the operator’s browser from driving the daemon. A POST, PUT, PATCH or DELETE to /api/v1/* gets 403 with code cross_origin_forbidden when the browser sends Sec-Fetch-Site: cross-site or same-site, or, for a browser that sends no Sec-Fetch-Site, an Origin whose host and port differ from the request’s Host header (Origin: null never matches). Only host and port are compared, not the scheme, so a TLS-terminating proxy in front of a plain-HTTP daemon keeps working. Requests with neither header (curl, scripts, the TUI) pass, as do GET requests and the SSE streams. The Vite dev server (page on :5173, proxying to the daemon) passes on any browser that sends Sec-Fetch-Site, which all have since 2023. To call the API from a web page on another origin, run with a token instead: token mode has no origin check, because a browser never attaches the bearer token cross-site. The guard is not access control: anyone who can reach the port with curl still has full access.

Host guard (DNS rebinding) in unauthenticated mode. The cross-origin guard cannot stop DNS rebinding: a page on a hostname the attacker controls re-points that name at your daemon’s address and then talks to it same-origin, so it could otherwise store a profile with an arbitrary command, rewrite settings, or read state and logs. With server.allow_unauthenticated: true, every route — API, SSE, the dashboard and static files — therefore requires the request’s Host (port and IPv6 brackets stripped, case-insensitive) to be localhost, an IP address, the configured server.host, or a name in server.allowed_hosts; anything else gets 403 with code host_not_allowed. IP addresses always pass because a rebinding attack needs a DNS name — so browsing to http://192.168.1.20:8090 with host: 0.0.0.0, container probes by pod IP, and SSH tunnels to localhost all keep working, as does the Vite dev proxy (it forwards Host: localhost:<port>). GET /api/v1/health is exempt because it returns a constant, and GET /api/v1/ready because it returns only booleans and the tracker’s rate-limit reset and is probed by whatever name the platform uses. If you reach an unauthenticated daemon through a reverse proxy, tunnel, MagicDNS or container service hostname, add that name to server.allowed_hosts. Token mode has no Host check: a rebound page never holds the bearer token.

See the Remote access guide for setting up persistent tokens and reverse-proxy TLS.


Daemon lifecycle: shutdown, reload and logs

Section titled “Daemon lifecycle: shutdown, reload and logs”

The first SIGTERM or SIGINT starts a drain:

  1. The daemon stops admitting work: no new dispatch, no retry, no pending-input resume, no automation start (automations are queued and persisted), no reviewer. resume, reanalyze, ai-review and provide-input answer 409 draining.
  2. GET /api/v1/ready answers 503 with "draining": true, so a load balancer stops routing to the daemon. The dashboard and API stay up.
  3. In-flight agent turns keep running on an uncancelled context and finish normally; their tracker writes and handoffs happen as usual.
  4. When the last turn finishes — or --shutdown-grace (default 30s) expires, or a second SIGTERM/SIGINT arrives — the daemon cancels whatever is still running (the agent’s process group is killed), waits for the orchestrator to record the exits, flush its ledgers and join its workers (at most ~30 s more), removes daemon.pid / dashboard_url / HEARTBEAT.md, and exits 0. A third signal exits without waiting.

The log shows shutdown: draining …, then shutdown: drain complete / shutdown: drain grace expired, forcing stop / shutdown: received second signal, forcing immediate stop. Whether a turn cancelled at grace expiry resumes its agent session after the restart is not verified for either CLI — size --shutdown-grace to your longest normal turn. Under a service manager, its stop timeout must exceed --shutdown-grace (for systemd: TimeoutStopSec, with KillMode=mixed so only the daemon receives the signal), and itervox stop --grace must exceed it too, or the drain is cut short by a SIGKILL.

Settings saved from the dashboard or TUI never reload: each is written as the daemon’s own write and applied in memory. An operator edit (your editor, git pull, a deploy) is detected by the file watcher and applied with drain, then reload: admission stops exactly as for a shutdown, in-flight turns finish (bounded by the same --shutdown-grace; past it the remaining turns are cancelled), and then the new generation loads the file and starts dispatching again. /ready reports draining meanwhile. A signal during a reload drain turns it into a shutdown. Edits made while draining are picked up by the same reload.

Why not the alternatives: reload-and-cancel (the old behaviour) kills every running turn on any edit, and whether a killed turn resumes is unverified; apply without cancelling would need every load-time field to be hot-swappable, while only the fields on the runtime-mutable list can be changed safely under a running generation. Draining keeps the one-generation- at-a-time invariant and loses no work; the cost is that an edit takes effect only after the current turns end (or the grace expires), and no new work starts in between. An edit that makes the file invalid still drains first, then shows the “config invalid” banner and waits for a fix.

--log-format text|json (or ITERVOX_LOG_FORMAT; the flag wins) selects the format of the rotating log file and, when no terminal is attached (systemd, containers, CI), of stderr — one JSON object per log record, for platforms that collect stdout/stderr. With a terminal, stderr stays human-readable text (and once the TUI starts, logs go to the file only). Both sinks go through the same secret redaction (Bearer tokens, lin_api_…, ghp_…, sk-ant-… are written as ***). Lines printed before logging starts or outside it (usage text, fatal start-up errors, the TUI’s own “not starting” notice) are plain text in either format.

Deployment preflight: itervox doctor --deploy

Section titled “Deployment preflight: itervox doctor --deploy”

itervox doctor --deploy [--workflow PATH] runs the normal doctor checks plus non-mutating deployment probes and prints one [ok] / [warn] / [fail] / [skipped] line each; it exits 1 if any line is [fail]:

ProbeHow
claude credentialsANTHROPIC_API_KEY, CLAUDE_CODE_OAUTH_TOKEN or ANTHROPIC_AUTH_TOKEN set (Bedrock/Vertex env accepted), else ~/.claude/.credentials.json (or $CLAUDE_CONFIG_DIR) with its expiry; an expired access token with a refresh token is fine. On macOS the login lives in the Keychain, which doctor does not read ([warn]).
codex credentialsCODEX_API_KEY / OPENAI_API_KEY set, else the read-only codex login status.
per SSH hostWith agent.ssh_hosts, one [skipped] line per host and backend: only this machine is probed.
gh authgh auth status (missing gh is a [warn]: PR detection and merges will not work).
git push authgit push --dry-run origin HEAD:refs/heads/itervox-doctor-probe in the WORKFLOW.md directory, with prompts disabled. The dry run never creates the ref; it proves the credentials, not branch protection or receive hooks.
tracker APILinear: one { viewer { id name } } query. GitHub: GET /repos/{owner}/{repo}, reporting whether the token may push (write labels/states). Through the same client code as the daemon.
daemon /readyGET /api/v1/ready on the URL in .itervox/dashboard_url; [warn] when no daemon runs here.

Probes only backends the configuration uses, loads .itervox/.env first like the daemon, runs each probe with a 10 s timeout, and prints one redacted line of any CLI output — never raw tokens. It does not check HEARTBEAT.md: the heartbeat is rewritten only when state changes, so its age is not a liveness signal. Unknown doctor flags are now an error (exit 2) instead of being ignored.


Any string field value that matches the pattern $VAR_NAME (a dollar sign followed by a valid environment variable name) is resolved to the value of that environment variable at startup. This is the recommended way to supply API keys and tokens.

tracker:
api_key: $LINEAR_API_KEY # reads process env LINEAR_API_KEY
workspace:
root: $ITERVOX_WORKSPACE # reads process env ITERVOX_WORKSPACE

.env file loading: Itervox automatically loads environment variables from .env files before reading WORKFLOW.md. The search order is:

  1. .itervox/.env (relative to the current working directory)
  2. .env (relative to the current working directory)

Only the first file found is loaded. Existing environment variables (already set in the shell) are never overwritten. The .itervox/.env file is created and git-ignored automatically by itervox init.

A typical .itervox/.env looks like:

Terminal window
LINEAR_API_KEY=lin_api_xxxxxxxxxxxx
GITHUB_TOKEN=ghp_xxxxxxxxxxxx

itervox init creates .itervox/.gitignore, .itervox/.env, and starter profile files. Commit .itervox/agents/**: those files are project agent definitions. Do not commit .itervox/.env, .itervox/HEARTBEAT.md, logs, runtime queue files, or other generated daemon state.

On startup the daemon writes .itervox/HEARTBEAT.md atomically with the current workflow path, schema version, dashboard URL, tracker/project, capacity, automation queue pressure, dependency audit summary, input-required count, retry count, and last notable error. Agents can read it when they need current daemon state; it is generated runtime state, not prompt text.


The body of WORKFLOW.md (everything after the closing ---) is a Liquid template. Itervox renders it once per issue dispatch to produce the agent’s prompt.

VariableTypeDescription
issue.identifierstringTracker-specific issue ID (e.g. "ENG-42" for Linear, "#123" for GitHub).
issue.titlestringIssue title.
issue.descriptionstringIssue body/description. May be empty — use {% if issue.description %} to guard.
issue.urlstringFull URL to the issue in the tracker.
issue.branch_namestringSuggested git branch name derived from the issue (e.g. "eng-42-fix-login-bug"). May be empty for GitHub issues.
issue.labelsstring[]Labels attached to the issue. Iterable with {% for label in issue.labels %}.
issue.prioritystringPriority label (Linear: "urgent", "high", "medium", "low", "no priority"; GitHub: empty string).
issue.commentsobject[]Comments on the issue. Each comment has author_name, body, and created_at fields.
idstringInternal tracker issue ID.
statestringCurrent issue state (e.g. "Todo", "In Progress").
blocked_byobject[]List of blocking issues. Each entry has id, identifier, and state fields.
created_atstringIssue creation timestamp (ISO 8601).
updated_atstringIssue last update timestamp (ISO 8601).
attemptint|nullCurrent retry attempt number. null on the first attempt.
You are an expert engineer working on this project.
## Issue {{ issue.identifier }}: {{ issue.title }}
{% if issue.priority %}
Priority: {{ issue.priority }}
{% endif %}
{% if issue.description %}
{{ issue.description }}
{% endif %}
{% if issue.labels.size > 0 %}
Labels: {{ issue.labels | join: ", " }}
{% endif %}
{% if issue.comments %}
## Discussion
{% for comment in issue.comments %}
**{{ comment.author_name }}**: {{ comment.body }}
{% endfor %}
{% endif %}
Issue URL: {{ issue.url }}
---
Create a branch named `{{ issue.branch_name | default: issue.identifier | downcase }}`,
implement the change, run tests, then open a PR that closes {{ issue.url }}.

When the agent genuinely cannot continue without a human answer, the preferred contract is to end its final message with the literal marker <!-- itervox:needs-input --> on its own line, followed by the actual question or confirmation prompt. That marker is deterministic and lets Itervox pause the issue immediately without any ambiguity.

Itervox also has a deterministic fallback for successful turns that end in a real blocking question, such as a choice between two options or a confirmation request. That means plain-English messages like “Which option should I take?” or “Type discard to confirm” can still move the issue into input_required.

The fallback is backup behavior, not the recommended path:

  • The explicit marker is more reliable.
  • The fallback is heuristic and tuned for common English phrasing.
  • The explicit marker makes prompt and skill behavior easier to reason about.

When the user replies from the dashboard or tracker, Itervox resumes the agent with that reply as the next user message — normally in the same session; after a daemon restart that had to rebuild the entry from tracker comments, a fresh session starts with the question and the reply as context.