Deploying on a Cloud VM
Itervox is a long-running stateful daemon. It holds orchestrator state in
memory in a single goroutine, spawns claude/codex subprocesses that live for
minutes, keeps git worktrees on a local filesystem, and serves long-lived SSE
connections.
That shape points at one conclusion: one VM with a persistent disk and a systemd unit. Not serverless, not scale-to-zero. See the platform tier matrix.
The deployment is identical across GCP, AWS, and Azure — only the provisioning
commands and the private-access primitive differ, which is why the per-cloud
scripts are thin wrappers around one shared bootstrap.sh.
Quick start
Section titled “Quick start”# 1. Provision the VM (no public IP by default)./deploy/gcp/provision.sh # or aws/ or azure/
# 2. On the VM: install deps, itervox, the systemd unit and the data disk# (provision.sh prints the stable --data-disk path for your cloud)sudo ./bootstrap.sh --repo git@github.com:you/yourproject.git --version v0.2.1 \ --data-disk /dev/disk/by-id/google-itervox-data
# 3. Fill in secrets, re-run the read-only deploy doctor until it reports no [fail], startsudo -H -u itervox bash -c 'cd /srv/itervox/yourproject && itervox doctor --deploy --workflow WORKFLOW.md'sudo systemctl enable --now itervoxThen reach the dashboard through your cloud’s private tunnel.
Sizing
Section titled “Sizing”Size for concurrent agents, not for the daemon. The daemon itself is cheap;
each agent is a full Claude Code process running your project’s builds and
tests. agent.max_concurrent_agents is the real knob.
| Concurrency | Machine | Disk |
|---|---|---|
| 1–2 agents | 2 vCPU / 8GB | 50GB |
| 3–4 agents | 4 vCPU / 16GB | 100GB |
| 5+ agents | 8 vCPU / 32GB | 200GB |
Disk fills faster than you expect: every issue gets its own git worktree, plus a
node_modules (or equivalent) per worktree. Alert on disk usage.
Secrets
Section titled “Secrets”Everything goes in <repo>/.itervox/.env, which the daemon loads at startup and
injects into its own environment — so spawned agent subprocesses inherit it.
# trackerLINEAR_API_KEY=lin_api_xxxx # or GITHUB_TOKEN=ghp_xxxx
# dashboard auth — pin this on a VM, see belowITERVOX_API_TOKEN=<64 hex chars>
# agent credentials — without this every dispatch failsANTHROPIC_API_KEY=sk-ant-xxxx# ...or the token `claude setup-token` prints (shown once, saved nowhere)# CLAUDE_CODE_OAUTH_TOKEN=.itervox/.env is gitignored. Prefer materializing it from your cloud’s secret
manager in a systemd ExecStartPre over baking it into an image.
Agent credentials are the real hurdle
Section titled “Agent credentials are the real hurdle”The single most common failure mode for a headless deployment: the daemon
starts, the dashboard comes up, and every dispatch fails because claude
has no credentials.
claude needs either ANTHROPIC_API_KEY in the environment, or a long-lived
OAuth token minted once interactively via claude setup-token (run it anywhere
you have a browser). claude setup-token prints the token once and does not
save it anywhere — copy it into .itervox/.env as CLAUDE_CODE_OAUTH_TOKEN
(see Secrets above) so the daemon and the agents it spawns inherit
it. Only an interactive /login run on the VM as the service account persists
credentials to that account’s ~/.claude directory. Sort this out before
debugging anything else — the symptom looks like an orchestrator problem and
isn’t.
Dashboard authentication
Section titled “Dashboard authentication”Itervox auto-generates an ephemeral bearer token on every bind — including
loopback — unless you explicitly opt out with server.allow_unauthenticated: true. Bind address is deliberately not a signal: a loopback bind behind a
tunnel or reverse proxy is exactly as internet-exposed as a public bind, and the
daemon cannot tell the difference from inside the process.
So you are authenticated by default. The reason to still set a token explicitly on a VM is stability, not security — an auto-generated token is regenerated on every restart, which breaks bookmarks and any saved session each time systemd restarts the unit.
openssl rand -hex 32 # put in .itervox/.env as ITERVOX_API_TOKENReach the dashboard once at https://your-host/?token=<token>. The frontend
captures it, stores it, and strips it from the URL.
The ?token= URL is not in journalctl or any other log sink: under
systemd stderr is not a terminal, so the daemon logs only dashboard URL (token withheld from logs) with the token-free URL and a sha256: fingerprint. If you
did not pin ITERVOX_API_TOKEN, the auto-generated token is in
<logs-dir>/api-token (mode 0600, rewritten on every start; the log line names
the path — by default it sits next to itervox.log in
~/.itervox/logs/<kind>/<slug>-<project>/, or
~/.itervox/logs/workflow/<project>/ when no project_slug is set).
Set ITERVOX_PRINT_TOKEN=1 to put the tokenised URL on stderr anyway;
--no-print-token suppresses it everywhere.
Exposing the dashboard
Section titled “Exposing the dashboard”Private tunnel — recommended. No public IP, no firewall rule, no TLS to
manage, access gated by cloud IAM. Strictly stronger than a bearer token. Each
per-cloud provision.sh defaults to this and allocates no public IP.
| Cloud | Command |
|---|---|
| GCP | gcloud compute start-iap-tunnel itervox 8090 --local-host-port=localhost:8090 --zone=<zone> |
| AWS | aws ssm start-session --target <id> --document-name AWS-StartPortForwardingSession --parameters '{"portNumber":["8090"],"localPortNumber":["8090"]}' |
| Azure | az network bastion tunnel -g <rg> -n <bastion> --target-resource-id <vm-id> --resource-port 8090 --port 8090 |
Tailscale. Best if you want phone access without running a tunnel client per session. See the Remote Access guide.
Public HTTPS via Caddy. Only if you genuinely need a public URL. Itervox has
no TLS of its own, so something must terminate it. deploy/caddy/Caddyfile gets
an automatic Let’s Encrypt certificate; pass --public your.domain.com to
bootstrap.sh.
WORKFLOW.md settings that matter on a VM
Section titled “WORKFLOW.md settings that matter on a VM”server: port: 8090 # PIN THIS host: 127.0.0.1 # keep loopback; the tunnel or Caddy fronts itPin the port. Omitting server.port defaults to 8090, but itervox init scaffolds
port: 0 (OS-assigned, a random free port on every start), which makes any static
proxy or health-check configuration meaningless — pin it explicitly on a VM.
Running under systemd
Section titled “Running under systemd”itervox starts a Bubbletea TUI when run interactively. Headless this is fine:
a TTY-ownership guard detects the absence of a controlling terminal, logs
statusui: refusing to start TUI, and the daemon continues. No flag needed.
A few unit settings are load-bearing and worth understanding before you edit them:
--shutdown-grace 240s,TimeoutStopSec=300,KillMode=mixed—SIGTERMbegins a drain that lets in-flight agent turns finish; systemd must wait at least shutdown-grace + 35 s, and must signal only the daemon. See Stopping, restarts and upgrades.ExecStartPre=/usr/local/bin/fetch-secrets.sh— optional secret-manager fetch (GCP Secret Manager, AWS Secrets Manager, Azure Key Vault) into/run/itervox/env; configured in/etc/itervox/secrets.env, a no-op otherwise.RequiresMountsFor=/srv/itervox— the service account’s HOME (/srv/itervox/home) and the checkout live on the data disk; the unit never starts on an unmounted directory.StandardInput=null— gives the TUI guard the “no terminal” signal it needs to fall back to headless cleanly.- Deliberately modest hardening. Agents run arbitrary project build and test
commands inside git worktrees, so
ProtectSystem=strict,ReadOnlyPaths, orPrivateDevicesbreak them in ways that are tedious to diagnose. The unit usesProtectSystem=fullandNoNewPrivileges=trueinstead. LimitNOFILE=65536— worktrees plus per-worktree dependency trees open a lot of files.
deploy/monitoring/ has a heartbeat-metrics.sh that parses
.itervox/HEARTBEAT.md into cloud custom metrics, plus a logging and alerting
guide.
Stopping, restarts and upgrades
Section titled “Stopping, restarts and upgrades”SIGTERM starts a drain: the daemon stops admitting work, /api/v1/ready
answers 503 with "draining": true, and in-flight agent turns get up to
--shutdown-grace to finish. After that they are cancelled and the daemon
allows up to 30 s more for the generation to stop. A second SIGTERM forces an
immediate stop, so never send one early. An operator edit to WORKFLOW.md
drains the same way before it is applied.
Whatever supervises the process must wait at least shutdown-grace + 35 s
before it escalates to SIGKILL:
| Supervisor | Setting (with the shipped --shutdown-grace 240s) |
|---|---|
systemd (deploy/systemd/itervox.service) | TimeoutStopSec=300, plus KillMode=mixed so SIGTERM reaches the daemon only — the default control-group would signal every agent subprocess at once and abort the turns the drain protects |
docker run / docker stop | --stop-timeout 300 / docker stop -t 300 (Docker’s default is 10 s) |
| Docker Compose | stop_grace_period: 300s |
| Kubernetes | terminationGracePeriodSeconds: 300 |
itervox stop | itervox stop --grace 275s — its default --grace 30s SIGKILLs a daemon that is still draining |
Change --shutdown-grace and the supervisor timeout together.
Upgrades, cleanup and backups
Section titled “Upgrades, cleanup and backups”On a bootstrap.sh install, sudo itervox-upgrade.sh --version <tag> downloads the
release and checksums.txt, refuses a checksum mismatch before touching the service,
swaps the binary atomically (keeping itervox.prev), restarts through the drain and
rolls back automatically when /api/v1/ready does not answer 200. Exit 2 means it
rolled back and the service is ready on the previous binary.
bootstrap.sh also installs two timers it does not enable:
itervox-cleanup.timerprunes workspaces idle longer thanITERVOX_CLEANUP_WORKSPACE_DAYS(default 14) and rotated/session logs older thanITERVOX_CLEANUP_LOG_DAYS(default 30), set in/etc/itervox/cleanup.env. It keeps every workspace that is running, retrying, paused, waiting for input, queued or recorded in a continuation ledger, and removes nothing when the daemon is up but its state cannot be read.itervox-watchdog.timerrestarts the unit after three consecutive/api/v1/readyprobes report a wedged event loop ("loop_fresh": false) — never during a drain or a tracker outage.
Back up the data disk (snapshots are scheduled by the OpenTofu modules) or, with the
service stopped, tar everything under --root. The full procedure — what each state
root holds, restore onto a new VM, and git worktree repair after a restore to a
different path — is in
docs/deploy-runbook.md.
Infrastructure as code (OpenTofu)
Section titled “Infrastructure as code (OpenTofu)”deploy/terraform/
has gcp-vm, aws-ec2 and azure-vm modules that create what the provision.sh
scripts do: a private VM with no public dashboard port, a separate persistent data disk
(prevent_destroy, plus snapshot retention and — on Azure — a delete lock), NAT egress,
a least-privilege identity limited to the secrets you list, and a first-boot script that
verifies the release checksum and runs bootstrap.sh --data-disk. Each module has an
examples/basic root module. CI runs tofu fmt -check and tofu validate on all of
them; validation does not prove an apply, so plan against a sandbox account first.
Official container image
Section titled “Official container image”deploy/docker/Dockerfile builds the web bundle and the Go binary with it
embedded — the same thing make build does — and ships them on
debian:bookworm-slim with git, gh, ssh and tini. Release tags publish it to
ghcr.io/vnovick/itervox; you can also build it from the repo root:
docker build -f deploy/docker/Dockerfile -t itervox:dev .docker run -d --name itervox --stop-timeout 300 \ -p 127.0.0.1:8090:8090 -v itervox-data:/data \ -e ITERVOX_REPO=https://github.com/you/project.git \ -e ITERVOX_API_TOKEN -e GH_TOKEN -e LINEAR_API_KEY -e ANTHROPIC_API_KEY \ itervox:dev- Non-root, one volume. It runs as uid 10001 with
HOME=/data, so every state root —~/.itervox/workspaces,~/.itervox/logs/<project>/(history, queue, breakers, sessions,api-token),~/.claude,~/.codexand the checkout in/data/repo— is on the/datavolume. - Clone once. The entrypoint clones
ITERVOX_REPOon the first start and reuses the clone afterwards, fetching and fast-forwarding only (local commits and edits are left alone).GH_TOKENis handed to git through a credential helper that reads it from the environment; it is never written to disk or logged. - Reachable port. The image sets
ITERVOX_SERVER_HOST=0.0.0.0(the daemon’s own default is127.0.0.1). The port isserver.portfromWORKFLOW.md, overridden byPORT(Cloud Run, Heroku) orITERVOX_SERVER_PORT. - JSON logs.
ITERVOX_LOG_FORMAT=jsonmakes stderr one redacted JSON object per line. - Signals. tini is PID 1 and forwards
SIGTERM; the default command passes--shutdown-grace 240s(see the table above for the stop timeout). - Health. The image’s
HEALTHCHECKprobes/api/v1/health, not/ready— see the caution below. - Agent CLIs are not installed by default; build with
--build-arg INSTALL_AGENT_CLIS=true, or extend the image with your project’s toolchain and CLIs.
deploy/docker/docker-compose.yml wires the same settings for Compose.
Platform tier matrix
Section titled “Platform tier matrix”Itervox needs one process that stays up for days, a persistent filesystem for worktrees and state, long-lived SSE connections, and a stop timeout long enough to drain. Measured against that:
Supported
Section titled “Supported”| Platform | Notes |
|---|---|
| GCP Compute Engine | VM + systemd, persistent data disk. IAP TCP tunnelling for private access. deploy/gcp/provision.sh. |
| AWS EC2 | VM + systemd, persistent EBS data volume. SSM Session Manager (IAM-gated) for private access. deploy/aws/provision.sh. |
| Azure VM | VM + systemd, persistent managed data disk. Tailscale or Bastion for private access. deploy/azure/provision.sh. |
| Official container image on a host you control (Docker, Compose, a single-replica Kubernetes Deployment or StatefulSet with a persistent volume) | See Official container image. |
Experimental (untested by the maintainers)
Section titled “Experimental (untested by the maintainers)”| Platform | What has to hold | Known limits |
|---|---|---|
| Cloud Run services | Exactly one instance (min = max = 1), CPU always allocated, a mounted volume for /data (an in-memory or GCS FUSE mount is not a good home for git worktrees). | Cloud Run gives a container only 10 s between SIGTERM and SIGKILL, so a drain cannot finish: every revision rollout or instance recycle cancels in-flight agent turns. Requests, SSE included, are capped at 60 min; the dashboard reconnects. Cloud Run worker pools have no HTTP ingress, so the dashboard is unreachable there. |
| Azure Container Apps | min = max replicas = 1, a persistent Azure Files volume for /data, terminationGracePeriodSeconds ≥ shutdown-grace + 35 s (lower --shutdown-grace if the platform caps it). | Check the ingress request/idle timeouts against the SSE guidance below. |
Not supported
Section titled “Not supported”| Platform | Why |
|---|---|
| AWS Lambda, Azure Functions, GCP Cloud Functions | Request-scoped execution with a hard cap, no process reuse, ephemeral filesystem. |
| Vercel | See Why Vercel is not suitable. |
| Anything that scales to zero or runs several replicas | Itervox is a single writer: two replicas double-dispatch, and scale-to-zero stops polling. |
Front-door timeouts and SSE
Section titled “Front-door timeouts and SSE”Whatever sits in front of the daemon (a cloud load balancer, Caddy, a tunnel,
Cloudflare) sees long-lived text/event-stream responses and must not cut them:
- Idle timeout. The snapshot stream
GET /api/v1/eventssends a namedkeepaliveevent after 25 s of inactivity (keepaliveIntervalininternal/server/handlers.go), so an idle timeout of 60 s or more keeps it open. The log streams (/api/v1/logs,/api/v1/issues/{id}/log-stream,/sublog-stream) send no keepalive; a shorter idle timeout closes a quiet one and the dashboard reconnects it. - Response-duration cap. An SSE response lasts as long as the tab is open.
Where the front door caps total response time (for example the GCP Application
Load Balancer’s backend-service
timeoutSec, 30 s by default), raise it — 3600 s is a reasonable value; the dashboard reconnects when it is reached. - Host names. With a bearer token (the default) any Host is accepted. With
server.allow_unauthenticated: truethe Host guard answers403 host_not_allowedto every DNS name exceptlocalhost,server.hostand the names inserver.allowed_hosts(IP addresses always pass;/api/v1/healthand/api/v1/readyare exempt), so list the proxy’s public name, a tunnel name or a container service name there. - Draining. Take the daemon out of rotation on
/api/v1/readyso a drain (503,"draining": true) stops new traffic without killing the process.