erun list
List all configured tenants and environments, the effective target for the current directory, and configured cloud providers. list is read-only and never mutates state. (To list managed cloud contexts, use erun context list.)
Synopsis
erun list [flags]
Output
Sections print in order — configuration location, defaults, the effective target for the current directory, configured cloud providers, every tenant and its environments, any stale ssh aliases, then orchestrators. The ssh section is absent unless there is something to report:
Configuration:
directory: /Users/you/Library/Application Support/erun
Defaults:
tenant: my-tenant
environment: local
Current Directory:
path: /Users/you/code/my-project
repo: my-project
configured tenant: my-tenant
effective target: my-tenant/local
kubernetes context: docker-desktop
type: local-agent
snapshot: enabled
repo path: /Users/you/code/my-project
Cloud Providers:
none
Tenants:
- my-tenant (default)
...
Orchestrators:
none
The full per-env field set (local port allocations, API URL, SSH details, …) prints under each tenant; the example abbreviates. See Configuration for what each value means.
Each orchestrator lists its linked environments beside what that orchestrator uses each one for: role=code, role=build, role=runtime, or role=undeclared when nothing has been set. role and an environment's own type are independent fields shown in different places — a runtime-type environment linked with the runtime role shows type: runtime under its tenant entry and role=runtime under the orchestrator entry, and the two mean different things even though they share a spelling.
Stale ssh aliases
erun sshd init writes a Host erun-<tenant>-<env> block into ~/.ssh/config, and erun delete removes the one it wrote. Nothing covers the other ways a block goes stale: an environment renamed rather than deleted, or its sshd turned off. The block survives, still pointing at 127.0.0.1:<its old local port> — and local ports are reissued, so the stale alias starts resolving into whichever environment inherited that port:
SSH config (~/.ssh/config):
erun-erun-proxmox1: no environment claims it, and its port 17122 now belongs to erun/petios — `ssh erun-erun-proxmox1` reaches erun/petios, not the environment the alias names
erun-erun-local: no environment claims it, and its port 17099 belongs to no environment
remove each block above by hand; erun cannot tell its own blocks from hand-maintained ones
That first line is the one worth acting on. An alias that fails is legible — that is what the not in ~/.ssh/config annotation under each environment already tells you. An alias that succeeds against the wrong environment is not: ssh, scp, VS Code Remote-SSH and workspace sync all connect, authenticate and operate on a pod you did not name. Workspace sync is the sharpest case, because it writes. The block's HostKeyAlias is stale with it, so the known_hosts entry consulted is the dead environment's.
Two properties keep the report trustworthy:
- Only aliases erun derives are reported. A
Hostblock that is noterun-…is yours, anderun listnever mentions it. - An environment claims its alias only while its
sshdis enabled. Withsshdoff nothing on the host answers for that name, so a block still naming it is reported rather than treated as accounted for.
Removal is manual and deliberately so: erun cannot distinguish a block it wrote from one you hand-maintained under the same naming convention, and a prune that guessed would delete an alias you rely on. Delete the Host block from ~/.ssh/config yourself.
The same list is carried as an orphanedSSHAliases array on the MCP list tool's structured result, so an agent sees it too.
Release lines
runtime-version: is a bare number, and a number alone can't say which release it belongs to: a
tenant can publish its own <tenant>-devops image and version it on its own line, so two
environments of the same tenant can genuinely be running different lines from each other. erun list names the line beside the number:
runtime-version: 1.0.84 (frs line, ghcr.io/sophium/frs-devops)
runtime-version: 1.0.203 (erun line, ghcr.io/sophium/erun-devops — release name frs-devops disagrees with the image)
runtime-version: 1.0.226 (line undetermined — no resolved runtime image recorded; redeploy to record it)
The middle line is the case worth double-checking: a frs-devops release running the stock
erun-devops image is legitimate, but it means that environment moves on ERun's own release line,
not the tenant's — comparing its number against another environment's tenant-line number is
comparing two different things. The last line means exactly what it says: nothing has recorded which
image this environment's pod actually runs yet, so erun list says so rather than guessing from the
tenant's name.
The sizing recommendation
runtime-pod: prints the size an environment was given. Under it, when ERun has watched the
environment long enough to have an opinion, two more lines print what that size should be — derived
from what the environment has actually done, not from what anyone guessed when they created it:
runtime-pod: cpu=12 memory=23552Mi
sizing: memory lower to 18432Mi from 23552Mi (peak 12153Mi of 23552Mi (52%), no oom kills, keeping 1.5x headroom, low confidence); cpu lower to 10 from 12 (busiest interval 4.567 of 12, 0.00% of scheduling periods throttled (0 of 376556), keeping 2x headroom, low confidence)
sizing-evidence: 31h12m observed, 240 samples, 1 restarts, knob=runtimepod, from cgroup memory.peak, cgroup memory.events oom_kill, cgroup cpu.stat usage_usec/nr_throttled (not loadavg)
The lines are advisory: ERun never resizes an environment on its own. Acting on one is
erun resize --apply-recommendation, which applies the computed value without
retyping it and restarts the pod — a decision for you to make and time.
Where the figures come from
The runtime container's own cgroup counters, read from inside the container. No metrics-server is
required — that is deliberate, because kubectl top answers "Metrics API not available" on local
clusters, which are the ones you iterate in all day. The counters are sampled on the same tick the
idle monitor already runs, and retained per environment.
Retained inside the environment: the history is written to the runtime container's own cache by the
monitor running in that container. So the sizing: lines above appear when erun list runs in the
environment they describe, and not from a host, which holds no history to derive a verdict from. The
usage MCP tool runs in the environment and carries the same verdict as a
sizing field; erun usage and erun doctor print a sizing block too,
and reach the same verdict from the same evidence — a reading taken from a host has the live counters
behind it, so the raise direction still fires there while the shrink direction reads as insufficient
evidence.
What each signal says
| Signal | Direction | Why |
|---|---|---|
memory.events oom_kill above zero | raise memory, high confidence | Something was already killed. One kill is enough. The suggestion is sized from the limit that proved too small, not from the observed peak — the allocation that triggered the kill was refused, so it never reached memory.peak. |
| Observed memory peak at 85% of the limit or more | raise memory, high confidence | This is the same 85% a memory warning fires at, deliberately: the alarm and the advisory answer one question about one reading, so a peak that warns is never left without the size that would fix it. Sampling means the true peak is at least the peak observed, so an environment already this close has plausibly gone further between two reads. |
| Observed memory peak below about two-thirds of the limit, over a long quiet window, no kills | lower memory, low confidence | The suggestion keeps 1.5× the observed peak. |
| 5% or more of scheduling periods throttled | raise CPU, high confidence | nr_throttled/nr_periods is real starvation: the container wanted CPU and the quota refused it. |
| Any throttling below that threshold | hold CPU | Tolerable, but not unused — the quota does bind sometimes, so this is not grounds to shrink. |
| No throttling at all, over a long quiet window, busiest interval below half the quota | lower CPU, low confidence | The suggestion keeps 2× the busiest interval measured. |
insufficient-evidence is a distinct answer from hold. hold means ERun looked and the size is
right; insufficient-evidence means it has not watched long enough, or the counter was unavailable,
and the line says which.
Why raising and shrinking are not symmetric
Being wrong in the two directions does not cost the same. An over-provisioned environment quietly consumes cluster capacity. An under-provisioned one kills a running agent — the failure behind "was killed (exit 137) — likely out of memory". So ERun raises on modest evidence and shrinks only on a long, quiet window, never below 1.5× the peak it actually observed, and never at better than low confidence. A quiet window is an argument from silence, and it is labelled as one.
That asymmetry decides how much observation each direction needs. A raise needs one reading, not a
window. A peak at the limit, or a recorded OOM kill, is a fact about something that already
happened; waiting a day to confirm it would withhold the answer exactly when it is wanted. Only the
lower direction is gated on the 24-hour window and the sample count behind it — and the
insufficient-evidence line names whichever of the two fell short.
A warning always arrives with its recommendation
Wherever ERun reports a reading — erun usage, the usage MCP tool,
erun list — a crossed memory threshold and the sizing advice that answers it are derived together,
from the same counters, in the same call. You will not see "memory is at 99% of its limit" without
the line that says what to resize it to.
The two cannot disagree because there is only one of them: the reading is folded into the retained
history as one more observation before a single verdict is computed. A live erun usage on a host
with no history at all still answers, because the reading in front of you is evidence on its own.
A raise is also bounded by what the environment's namespace quota can admit, where one is
configured: a ResourceQuota counts every container in the pod, so the erun-dind sidecar's own
limit is spent before the runtime container gets anything, and a suggestion above the remainder is a
size Kubernetes would refuse to schedule. Raising the quota is a separate decision from resizing the
pod, so ERun clamps and says it clamped rather than silently recommending both.
Two limits worth knowing
Without these, the output reads wrong:
memory.peakresets when the container restarts. It is a high-water mark for the current container lifetime, not for the environment. This is exactly why ERun retains a history instead of reading the counter live —sizing-evidencereports how many restarts the history spans, and the peak it quotes survived them.- PSI is unavailable on some kernels.
memory.pressureandcpu.pressureare simply absent in the runtime container's cgroup on some hosts, so nothing here depends on pressure stall information.nr_throttled/nr_periodsis the CPU-starvation signal that is actually available.
And one thing the figures are not: a host loadavg. A loadavg counts the machine's runnable queue,
so a 12-core environment sitting at a load of 10 looks saturated while cpu.stat reports not one
throttled period — which is why the evidence line names its counters and says not loadavg.
When the lines do not print
The history is written by the container that produced it, so it exists on that environment's own
runtime pod. Running erun list from your laptop shows no sizing lines for a remote environment —
there is no history there to read. Ask the environment (over
MCP or an erun open shell) and it answers about itself; the MCP
list tool carries the same recommendation as a structured sizing field on each environment.
erun usage is the exception, and deliberately so: it reads the environment live, so
the reading it just took is itself the evidence, and it prints the recommendation whether or not a
history happens to exist on the host you ran it from.
A newly created environment prints nothing either, until its monitor has taken a sample.
Version drift across a tenant
Pass --tenant to switch list from the full listing into a focused check: which erun version every environment in that tenant is running, and the newest version observed among them.
erun list --tenant my-tenant
Version drift for tenant my-tenant:
max version: 1.0.247
environments:
- build version="1.0.246" [behind max]
- code4 version="1.0.247"
Add --gate-environment to name the environment that drives that tenant's merge-queue gate — erun doesn't track this on its own, so you say which one it is — and the report additionally flags whether that environment is running an older erun version than any environment it gates:
erun list --tenant my-tenant --gate-environment build
Version drift for tenant my-tenant:
max version: 1.0.247
environments:
- build version="1.0.246" [behind max]
- code4 version="1.0.247"
gate:
environment: build
version: 1.0.246
behind: yes -- outdated relative to code4
A gate older than the code it gates can pass a change that would fail on current code — this is what catches that before a bad merge slips through. See collaboration › merge queue for what the gate does. For the exact flag contract and JSON shape, see Agent reference › CLI flags.
Like the rest of list, this report exits 0 on its own — list reports, it doesn't gate. Add --fail-on-drift to make that one invocation exit non-zero when an environment is behind the tenant's max, or the named gate environment is itself behind (or its version couldn't be resolved), so a script or a schedule can act on it:
erun list --tenant my-tenant --fail-on-drift
Control plane versions
--tenant catches drift between your own environments. It has no baseline for a different, easy-to-miss gap: a control plane can simply never get rolled onto a release that has already shipped, and nothing about its own reported version says whether that release exists. Pass --control-planes for that check instead: every erun-hosted control plane you've configured (a cloud provider alias with provider: erun, e.g. the one erun cloud init erun creates), compared against the newest version erun's own registry has actually published.
erun list --control-planes
published version: 1.0.247
Control planes (1 backend, 1 alias):
- erun+api.erunpaas.com@erun api-url="https://api.erunpaas.com" reachable=yes version="1.0.245" [behind published -- roll it]
console: url="https://console.erunpaas.com" reachable=yes version="1.0.245" [behind published -- roll it]
docs site: url="https://docs.erunpaas.com" reachable=yes version="1.0.247"
Each plane is checked with its own unauthenticated GET /v1/platform — the same call erun cloud init erun uses to discover a plane in the first place — so a plane that doesn't answer prints reachable=no reason="..." instead of a version; an unreachable plane is never reported current. [behind published -- roll it] means the plane is running a real, older release than what's published; [ahead of published -- running an unpublished version] means the opposite and more unusual case — the plane is running something the registry has never published at all, which is worth investigating on its own rather than "just roll it".
That same GET /v1/platform response also names the plane's linked console (consoleUrl — a plane and its console are always deployed together, never configured as a separate alias), so each reachable plane's console is checked the same way, against the same published baseline, and printed nested under it as a console: line — a plane can be current while its console lags behind, or vice versa, and before this there was no way to tell. A plane whose response carries no consoleUrl prints no console: line at all, rather than guessing.
A surface that answers GET /version.json but doesn't serve the expected JSON document — for example an SPA fallback page served for a route that isn't wired up yet — is never folded into reachable=no: it did answer, so it prints reachable=yes version=unknown reason="..." instead, with the reason naming the HTTP status and content type it actually served.
The same response names the plane's documentation site (docsUrl), and its version is compared and printed exactly the same way, as a docs site: line under the plane. It is worth its own line rather than being folded into the console's because the two are published by different deploys: a docs site is normally uploaded on its own, so a console that is current says nothing about whether the docs you are reading come from the release you are running. As with the console, a plane whose response carries no docsUrl prints no docs site: line at all — an absent docs site reads as absent, never as one that is up to date.
That response also names the apiUrl the plane believes it is served at. When that resolves to a different backend than the one you configured — a plane advertising a different plane's api — it prints an [advertised apiUrl mismatch: ...] line beneath the plane's own line. An apiUrl that merely differs textually is never flagged: a vanity hostname that CNAMEs to the one erun dialed is the same backend under two names, so the two hostnames are resolved and compared instead of string-matched, and a hostname that doesn't resolve on either side prints nothing rather than a guess.
This makes real network calls (each plane and the surfaces it advertises, plus erun's registry), so add --dry-run to preview which planes, linked surfaces, and registry lookup would be checked without making any call.
This report also exits 0 on its own, same as --tenant's above. Add --fail-on-drift to make that one invocation exit non-zero when a plane, its console, or its docs site is behind or ahead of published, a plane advertises a foreign apiUrl, a plane or linked surface is unreachable, or the published baseline itself couldn't be resolved — none of those confirm a plane and its surfaces are running what erun actually published:
erun list --control-planes --fail-on-drift
--fail-on-drift never fires under --dry-run: nothing was actually probed, so there is nothing to fail on.
Common usages
erun list # full listing
erun list | grep -i "tenant" # quick scan of names
erun list | grep "effective" # what ERun targets right now
erun list is what every troubleshooting flow should start with — it tells you which environment ERun considers "effective" right now and what its resolved config looks like.
Error behaviour
| Failure | Behaviour |
|---|---|
| No config yet. | Prints the sections with none placeholders; not an error. |
| Current directory isn't a configured project. | effective target: none (or unavailable (…) with the reason); the rest still prints. |
| Config file unreadable. | Errors with the read failure; nothing is printed. |
--gate-environment passed without --tenant. | Errors --gate-environment requires --tenant; nothing is printed. |
--tenant names a tenant with no config. | Errors tenant "<name>" not found. |
--gate-environment names an environment not in that tenant. | Errors gate environment "<name>" not found in tenant "<tenant>". |
--control-planes combined with --tenant/--gate-environment. | Errors --control-planes cannot be combined with --tenant/--gate-environment; nothing is printed. |
--control-planes and a configured plane or a surface it advertises is unreachable, or the registry lookup fails. | Not an error — printed as a finding (reachable=no reason="...", or published version: unresolved (...)); exit code stays 0 unless --fail-on-drift is set. |
--fail-on-drift passed without --tenant or --control-planes. | Errors --fail-on-drift requires --tenant or --control-planes; nothing is printed. |
--fail-on-drift set and the report finds drift (an environment behind max, a behind gate, an unreachable/behind/ahead plane or linked surface, a plane advertising a foreign apiUrl, or an unresolved published baseline). | The full report still prints, then the command exits non-zero naming what it found. Never fires under --dry-run — nothing was probed, so there is nothing to fail on. |