Skip to main content

erun doctor

Inspect the local ERun configuration or the runtime pod state, report why a deploy may have failed, and offer recovery actions for any problems it finds.

Synopsis​

erun doctor [TENANT] [ENVIRONMENT] [flags]

What it checks​

erun doctor runs a different check set depending on context. From your laptop, it validates the per-user tenant + env config, the kubeconfig context, the runtime pod's reachability, and the project root, and reports why a deploy may have failed: the helm release status and the runtime namespace's pods (read-only — it never touches the release). Inside the runtime pod (detected via ERUN_REPO_REMOTE=true) it inspects the bootstrap marker, the in-pod project root, the git checkout, the SSH keypair, and the CodeCommit RSA key when applicable.

When the deploy diagnosis shows a stuck pending release or a failed image pull, the fix is to re-run erun deploy --force (rebuild and redeploy) or clear the pending release — the desktop app offers both as one-click buttons on the failed deploy in its Activities panel.

For the full per-check id catalogue and the offered recovery actions, see Agent reference · CLI flag spec · erun doctor.

For an environment carrying an AWS cloud alias, doctor also reports the host AWS credentials it acts with: whether the pod's erun-host profile exists, when it expires (or that it has already expired), and the AWS region the environment resolves — or that none does. Both failures otherwise surface far from their cause, as an SDK ExpiredToken or an image pull rejected with no basic auth credentials, so this is the check that names them. The fix in either case is erun cloud refresh; see Troubleshooting. Reading it requires the runtime pod: when the pod isn't reachable, doctor reports could not read for this check instead of aborting, pointing back at the helm release status and pod state it already reported above, and carries on to the rest of the run — retry the check once the pod is running again.

For a remote-agent or runtime environment — one whose project checkout lives inside the pod rather than on your own machine — doctor also reports git push access: the project's origin remote, whether an anonymous fetch against it succeeds, and whether any credential that could push (a gh session, GH_TOKEN/GITHUB_TOKEN, or an SSH key the remote host accepts) is configured. A public GitHub repository fetches anonymously for the entire life of a piece of work, so an environment can look completely healthy through a whole session of cloning, building, and testing, and only discover at the very end — trying to push a branch or open a PR — that it never had a credential. This check exists to surface that gap up front instead. It never runs gh auth login, gh auth refresh, or gh auth switch — only the read-only gh auth status — so reading it can never itself start gh's interactive device-code/browser flow, which cannot complete in a headless pod. See Troubleshooting. This check does not apply to a local-agent or host environment, whose checkout lives on your own machine and already carries your own git/gh credentials.

doctor also compares the environment's runtimeimage against its runtimeregistry and flags a mismatch by name — a runtime image pinned to a different registry than the one credentials refresh for is exactly the split that leaves a redeployed pod stuck in ImagePullBackOff. This check reads only local config, so it runs identically whether or not the pod is up. On a mismatch it names both registries and points to two fixes: confirm a credential for the image's registry resolves where erun deploy/erun open --deploy runs, or realign the two with erun init <tenant> <env> --runtime-registry <registry> so runtimeregistry matches the image. For example:

== Runtime image registry ==
runtimeimage resolves to registry 123456789012.dkr.ecr.eu-west-2.amazonaws.com, but runtimeregistry is ghcr.io/sophium.
The deploy that installs this env sets imageOverrides.erun-devops from 123456789012.dkr.ecr.eu-west-2.amazonaws.com while runtimeRegistry stays ghcr.io/sophium; the runtime pod can only pull if a credential for 123456789012.dkr.ecr.eu-west-2.amazonaws.com also resolves where `erun deploy`/`erun open --deploy` runs (the same AWS/docker session that can push to it). If the pod is failing to pull, confirm that credential is available there, or realign the two with `erun init team dev --runtime-registry 123456789012.dkr.ecr.eu-west-2.amazonaws.com` to match the image.

See Troubleshooting.

When any item is missing, doctor offers to run the corresponding recovery step.

doctor also reports the execution mode of every operation that can run through either a CLI subprocess or an equivalent Go library — today aws-sts (aws sts get-caller-identity), aws-sts-web-identity-token (aws sts get-web-identity-token), aws-export-credentials (aws configure export-credentials), kubectl-namespace-get (kubectl get namespace <name> -o name), kubectl-pvc-get (kubectl get pvc <claim> -o name), kubectl-secret-get (kubectl get secret <name> -o json), kubectl-pod-get (kubectl get pod <name> -o json), kubectl-deployment-get (kubectl get deployment <name> -o name), kubectl-deployment-wait (kubectl wait --for=condition=Available), kubectl-secret-apply (kubectl apply -f - of a Secret), kubectl-pod-watch (the kubectl get pods -o json poll erun deploy runs alongside every helm rollout), and kubectl-context-configure (the kubectl config set-cluster/set-credentials/set-context trio erun cloud context uses to point a cloud context's kubeconfig entry at its own cluster), with more to follow — so whether an install opted an operation into the library path, or left it on the default subprocess path, can be confirmed rather than guessed at:

== Execution modes ==
aws-sts: subprocess
aws-sts-web-identity-token: subprocess
aws-export-credentials: subprocess
kubectl-namespace-get: subprocess
kubectl-pvc-get: subprocess
kubectl-secret-get: subprocess
kubectl-pod-get: subprocess
kubectl-deployment-get: subprocess
kubectl-deployment-wait: subprocess
kubectl-secret-apply: subprocess
kubectl-pod-watch: subprocess
kubectl-context-configure: subprocess

See Configuration reference · Execution modes for the config key that controls it.

doctor also reports the environment's own resource pressure, under == Resources ==, together with the standing sizing verdict it implies:

== Resources ==
CPU: 49.8% of a 6.00-core quota (sampled over 1.0s)
Memory: 4.0GiB / 4.0GiB (100.0%), peak 4.0GiB, OOM kills 3
Warnings (3):
memory is at 100% of its 4096Mi limit (warns at 85%)
memory.peak reached 100% of the limit (warns at 95%) -- this environment came close to an OOM kill
the cgroup recorded 3 OOM kill(s)
Sizing recommendation:
sizing: memory raise to 6144Mi from 4096Mi (3 oom kill(s) at 4096Mi, high confidence); cpu raise to 9 from 6 (9.85% of scheduling periods throttled (9850 of 100000), high confidence)
sizing-evidence: 0m observed, 1 samples, 0 restarts, knob=runtimepod, from cgroup memory.peak, cgroup memory.events oom_kill, cgroup cpu.stat usage_usec/nr_throttled (not loadavg)
Apply with `erun resize --tenant team --environment dev --memory 6144Mi --cpu 9` (rolls the runtime pod and kills any live session in it), or `--apply-recommendation` from inside the environment, where the retained usage history lives. Doctor reports this and never resizes.

This is the same reading, the same thresholds and the same recommendation erun usage reports — doctor renders it through that command's own writers, so the two cannot report different numbers or a different verdict for one environment. It exists because an environment at 100% of its memory limit with OOM kills recorded is exactly the state an operator reaches for doctor to explain, and it used to say nothing about it: the signal was on a command nobody runs unless they already suspect sizing. On a build-capable environment the sidecar's own reading, and the disk figures beneath it, come along for the same reason they do in erun usage — every image build runs in erun-dind, a separate cgroup this reading would otherwise report as idle.

doctor never resizes: a resize rolls the pod and kills whatever is running, so the section names the remedy instead of applying it, with the --memory/--cpu values spelled out rather than a bare --apply-recommendation. That matters from a host: --apply-recommendation re-derives the verdict from the environment's retained usage history, which is readable only from inside the environment's own pod, so the explicit values are the form that works where doctor was run. A verdict that suggests nothing — a hold, or insufficient evidence — gets no next action, because there is nothing to apply.

The section always prints, because "nothing was observed" is not "nothing is wrong": a reading whose counters could not be read says so and names the reason, rather than staying silent and reading as a clean bill of health for the one environment least able to say otherwise. Under --dry-run the read is traced but not taken, like every other subprocess doctor would run, so a dry run shows the traced read and no resources section.

doctor's desktop-app check inspects every copy of the ERun.app bundle this host could hold, including the ones erun app itself launches from — beside the erun executable, which is where a packaged install and the dev wrapper's BIN_DIR (erun-cli/bin, or $ERUN_DEV_BIN_DIR) both put it, and the checkout's erun-ui/bin. It names which copy a launch would actually reach, so a bundle that matches this CLI no longer reads as "nothing to report" while the running app is a different build at a path the check never looked at. Copies that share the bundle id com.sophium.erun are all listed, because macOS (Finder, Spotlight, the Dock) can launch any of them regardless of which is current.

What it can repair​

Beyond reporting, doctor offers these fixes (each prompts first, or runs non-interactively with its flag). Without a TTY on stdin — an MCP client, an orchestrator, a CI step, or erun doctor … </dev/null — doctor skips the optional prune prompts instead of blocking on them, names each skipped step in the report, and still exits on the health of what it examined. Nothing is pruned without either an explicit --prune-* flag or an answer to the prompt, and a skipped optional step is never reported as a failed check: doctor is the command you reach for when a deploy has already failed, which is exactly when nobody is at a terminal to answer it.

  • Deploy recovery — when the diagnosis shows the runtime release is unhealthy, doctor recommends the one recovery that fits: clearing a stuck pending helm release when a deploy died mid-upgrade and left it locked, or rolling back to the last successful revision when the current one is bad. It prompts for that single action (never both — they are alternative fixes, and running both would roll the release back a revision too far). These mutate the live release and are offered only when the release looks unhealthy, never on a healthy env. To rebuild and roll out fresh images instead, re-run erun deploy --force.

  • Docker cleanup — prune the environment's unused images, build cache, or stopped containers, against the docker daemon that actually holds them. Which daemon that is follows the environment's type, and every line the report prints about docker storage names it:

    • a local-agent or remote-agent environment builds in the erun-dind sidecar inside its runtime pod, so the prune runs there — the same daemon the build's own disk-headroom preflight measures;
    • a host environment has no pod and builds against the docker daemon on the machine you run erun doctor from, so the prune runs there and says so;
    • a runtime environment installs published versions and never builds, so it carries no daemon holding build images. A prune requested against one fails naming that (and naming a build environment to prune instead) rather than printing a successful prune of a daemon that holds nothing.

    After a prune, the report states what it did to that daemon: the reclaimable space docker reported before and after it. A prune that freed nothing is reported as that, rather than resting on docker's own Total reclaimed space line alone.

  • Root config repair — restore the root erun config from a dated backup, or re-initialize orphaned cloud provider aliases.

  • Environment config restore — restore one environment's config.yaml from a dated backup when a setting was changed or corrupted (for example an environment type that resolved to the wrong value). Each save snapshots the previous config alongside it, so there is a daily backup to roll back to.

  • JetBrains Gateway — clear cached backend metadata for the environment when a Gateway connection is stuck.

  • Workspace sync (host mirror) — for a remote-agent env with workspace sync enabled, report the host mirror's SSH provisioning and, with --repair-workspace-sync, repair it without redeploying: resolve/persist the SSH key, write the local ~/.ssh/config alias, install the pod's authorized_keys through the runtime container (not a helm redeploy), and ensure the SSH port-forward. If SSH still can't reach the pod afterwards, it names erun sshd init as the remaining step (the redeploy this repair won't run).

  • Remote init — inside a runtime pod, finish an interrupted init (SSH keygen, repo clone).

The exact flags for running these non-interactively are on the CLI flag spec.

Flags​

FlagDescription
--dry-runRun the inspection and print the recovery plan without performing any recovery actions.
--sync-configReconcile the in-pod erun config with the helm-injected ERUN_* env vars. Only takes effect inside a runtime pod. The injected env wins: erun rebuilds the canonical projection (type, kubernetescontext, cloudprovideralias, managedcloud, the cloud-context/provider blocks, idle, runtimeregistry, containerregistries, disablebuildscript) and rewrites those keys, preserving every key the env does not carry (sshd, claude, runtimeversion, localrepopath, …). Drift is reported per key as missing, wrong, or legacy; with --dry-run the file writes are traced but not performed.
--restore-env-config-from-backup <date|path>Restore the target environment's config.yaml from a dated backup (YYYY-MM-DD) or an absolute path, before the rest of the inspection runs so a corrupted env config can be recovered first. Needs an explicit tenant and environment. With --dry-run, the copy is traced but not performed.
--repair-workspace-syncFor a remote-agent env with workspace sync enabled, repair the host mirror's SSH provisioning without redeploying the runtime: resolve/persist the SSH key, write the local ~/.ssh/config alias, install the pod's authorized_keys through the runtime container, and ensure the SSH port-forward. With --dry-run, every action is traced and nothing runs.

Examples​

Run from your laptop against the effective env:

erun doctor

Run inside a runtime pod after an interrupted init:

# (SSH'd into the pod)
erun doctor # see what's missing
erun doctor --dry-run # preview what doctor would do to fix it

Exit codes and the meaning of each are spec'd in Agent reference · CLI flag spec · erun doctor exit codes.

Sample output​

A healthy local-side run against an env named local on Docker Desktop:

erun doctor — my-tenant / local
config:
tenant config ok ~/Library/Application Support/erun/my-tenant/config.yaml
environment config ok ~/Library/Application Support/erun/my-tenant/local/config.yaml
project config ok /Users/you/code/my-project/.erun/config.yaml
cluster:
kubernetes context ok docker-desktop
runtime pod ok my-tenant-local/erun-devops-7c8b6d (running)
workspace:
project root ok /Users/you/code/my-project (git repo)

all checks passed

An unhealthy run after an interrupted init:

erun doctor — my-tenant / rihards-dev
config:
tenant config ok ~/Library/Application Support/erun/my-tenant/config.yaml
environment config ok ~/Library/Application Support/erun/my-tenant/rihards-dev/config.yaml
cluster:
kubernetes context ok erun-004-020362606330-eu-west-2
runtime pod missing no pod found in namespace my-tenant-rihards-dev

recovery actions:
[1] deploy runtime chart (erun open my-tenant rihards-dev)

run `erun doctor --dry-run` to preview the recovery, or `erun doctor` again with `-y` to apply.

The check format is fixed (<category>: <name> <status> <detail>); machine-readable consumers should prefer the MCP doctor tool which returns the same data as typed JSON.

Error behaviour​

FailureBehaviour
Root config missing or corrupted.Reports it; with --repair-config offers to restore from a dated backup, otherwise leaves it untouched.
--restore-env-config-from-backup without an explicit tenant and environment.Aborts with --restore-env-config-from-backup needs an explicit tenant and environment; exit code 1; nothing is changed.
--restore-env-config-from-backup <date> with no matching backup.Aborts naming the unmatched selector and the target env (no env config backup matches "<date>" for <tenant>/<env>); exit code 1; nothing is changed.
Helm release missing or cluster unreachable during deploy diagnosis.The helm probe's output (including the error) is shown as part of the diagnosis and doctor continues; the diagnosis is read-only, so nothing is changed. If the error confirms the Kubernetes API server itself is unreachable (not just a missing release or an RBAC denial), the pods probe that would normally follow is skipped rather than re-timing-out against the same unreachable cluster — see the next row.
Host AWS credentials, git push access, or docker-storage check needs the runtime pod, but it isn't reachable.If the deploy diagnosis already confirmed the cluster is unreachable, doctor reports skipped for that check immediately, pointing back at the helm release status shown above, instead of re-probing and paying its own timeout to rediscover the same fact. Otherwise doctor reports could not read the first time that check's own probe hits the failure. Either way doctor continues the rest of the run instead of aborting, and any requested prune action (--prune-images, --prune-build-cache, --prune-containers) is skipped with the same reason.
Deploy recovery (--clear-pending-helm / --rollback) fails.The failing helm/kubectl output is surfaced and doctor aborts with that error; the release is left as helm leaves it (a failed rollback does not partially apply). Re-run the diagnosis to see the new state, then retry or erun deploy --force.
--rollback with no prior successful revision.helm rollback reports it has no revision to roll back to; nothing changes. Use --clear-pending-helm then erun deploy --force instead.
Both --clear-pending-helm and --rollback passed.Aborts immediately with --clear-pending-helm and --rollback are alternative recoveries; pass only one; exit code 1; nothing runs.
Prune not confirmed and no --prune-* flag.No Docker state is touched — prunes run only on confirmation or with the matching flag.
A --prune-* action requested against an environment with no daemon holding build images (a runtime environment, or one whose type is unset).The prune is refused with the reason and the environment to prune instead; exit code 1. Nothing is dispatched: pruning some other daemon would report a reclaim this environment's builds cannot use. Without a prune flag the docker-storage section reports the same reason and the rest of the run continues.
No TTY on stdin, so the optional prune prompts cannot be answered.Each optional prune is reported as skipped: with its flag named as the way to run it explicitly. This is not a failed check: the run completes and the exit code reflects the health of what was actually examined, so an orchestrator or CI step gets the diagnosis instead of Doctor failed …: ^D.
Stdin reaches EOF at a prompt doctor cannot skip — a recovery that mutates the live release, or a repair you asked for with a flag.The step is not run: a missing answer is not consent. The report says the prompt went unconfirmed and names the flag that runs it without one (--clear-pending-helm, --rollback, --repair-config, …), and the rest of the diagnosis still runs. An unanswered prompt is never reported as a failed environment.
Run inside a runtime pod with a complete init.Reports "nothing to finish" and exits 0.
Cluster unreachable.Reports the pod check as failed; config and workspace checks still run.