Skip to main content

Observability

What can you see when something goes wrong (or right). ERun doesn't ship a built-in observability stack — Kubernetes is the substrate, and you reach for the cluster's normal tools.

Three layers​

LayerWhat you seeHow to read it
In-podStdout/stderr of every container, plus files under /var/log/erun/ (CLI audit traces).kubectl logs, erun open + shell, MCP raw.
ClusterAggregated logs / metrics / traces across all envs on the cluster.Whatever stack you've installed — Prometheus + Grafana + Loki, OpenTelemetry Collector, Datadog, ….
DurableReviews, comments, builds, audit events — anything posted to the erun API.erun API endpoints (GET /v1/audit-events for the audit trail; see Audit log · Query API).

The first layer is always there. The other two are admin-opt-in.

Logs​

From your laptop​

# Application service logs:
kubectl logs -n <tenant>-<env> -l app=<component> --tail=200 -f

# Runtime pod (Operator/Agent shared surface):
kubectl logs -n <tenant>-<env> -c erun-devops <runtime-pod> --tail=200

# All containers in a pod:
kubectl logs -n <tenant>-<env> <pod> --all-containers --prefix

The runtime pod has two containers (erun-devops, erun-dind); -c <container> selects between them.

From inside the env​

After erun open, you're in erun-devops. The CLI's audit trace lives at /var/log/erun/audit.log — JSON-lines, event shape documented separately.

tail -F /var/log/erun/audit.log | jq -r '.action + " " + .result'

Cluster-wide aggregation​

ERun doesn't deploy a log aggregator. The conventional setup is Loki + Promtail (or Vector / Fluent Bit / your existing aggregator) installed once per cluster. Every env's containers ship their stdout/stderr through it automatically.

In Grafana, filter by the Kubernetes namespace label to scope to one env: {namespace="<tenant>-<env>"}.

Is this environment busy?​

Logs tell you what happened; this tells you whether anything is happening right now. An environment reports its own condition, and the desktop's sidebar renders it — so an environment being driven hard by an Agent looks different from one nobody has touched in a week, and different again from one that is stopped.

Three things feed it:

  • Activity leases. Long work announces itself for its lifetime (erun activity lease take --name <what>, or the MCP activity_lease_take tool). This is the only signal that names the work, and the only one that survives a job making no calls for forty minutes. Idle-stop will not stop an environment holding one.
  • Sampled processes. The in-pod monitor notices resident build and agent processes that are actually burning CPU, so work nobody instrumented still registers.
  • Request activity. SSH, API, MCP, and CLI traffic, as before.

Read it from anywhere with the MCP idle tool or erun idle <tenant> <env> --json: leases says what is holding the environment, markers breaks the rest down per source, and stopBlockedReason names whichever one is deferring auto-stop.

For the full schema and the expiry / liveness rules, see Agent reference · Idle policy; for the operator-facing commands, erun idle · Activity leases.

Metrics​

ERun's runtime pod exposes a single Prometheus-format metrics endpoint on port 9100. The series cover idle-eligibility, terminal-input freshness, the network-traffic window, MCP tool-call counts, and audit-event counts — labelled by tenant + environment.

The endpoint carries no authentication, so reachability is scoped by a NetworkPolicy instead: always open within the env's own namespace, and open to another namespace only once that namespace is labelled network-policy/erun-metrics-scraper=true — label your cluster's Prometheus namespace once (kubectl label namespace <prometheus-namespace> network-policy/erun-metrics-scraper=true) to let it scrape every env.

For the full schema (every metric name, every label, type, source, cardinality envelope), see Agent reference · Metrics spec.

Application services use the metrics conventions of their own framework. ERun has no opinion — scrape with the cluster's normal Prometheus setup.

Traces​

For distributed traces across the env's services, deploy an OpenTelemetry Collector as a sidecar to your services (or run one as a DaemonSet). ERun has no built-in trace pipeline; the env's namespace is just a Kubernetes namespace, so any standard tracing setup works inside it.

The runtime pod itself does not emit traces. The audit log + MCP tool counts are the closest equivalent — they capture every Operator and Agent action with timestamps.

Step timing​

erun build, release, push, and deploy each print a duration-ordered table on completion — success or failure — breaking their single elapsed time down into every image build (per architecture), chart publish, release stage, and deploy target, so a slow or failed run points at what actually took the time instead of leaving you to guess:

==> Pushed in 6m54s
step timing (ordered by duration):
push [6m54s]
frs-docs (cache miss: fingerprint image is missing for platforms [linux/amd64, linux/arm64]) [6m20s]
linux/amd64 [3m20s]
linux/arm64 [3m0s]
chart frs-docs [30s]

Each command also writes the same breakdown as a JSON file under ~/.erun/timing/, so a slow run and a normal one can be diffed instead of compared by eye across terminal scrollback.

For the full JSON shape and the cache-hit/miss annotations, see Agent reference · Step timing.

The audit trail​

Distinct from logs / metrics / traces, the audit trail is ERun's primary record of who did what when. Three storage layers:

  • In-environment trace — /var/log/erun/audit.log in each runtime pod (lifetime of the pod).
  • MCP events — 30 days inside the pod.
  • erun API events — durable, lives as long as the tenant.

When a service incident needs an "actor and intent" reconstruction, the audit trail is the source of truth. Logs and metrics are for "what was the state of the system."

Quick reference​

You needWhere to look
Service crashedkubectl logs -n <tenant>-<env> -l app=<component> --previous
Idle-stop fired unexpectedlyMCP idle tool result + erun_idle_* Prometheus series
Is anything running in this env right nowSidebar row indicator; MCP idle tool leases + markers
Why an idle-looking env never stopsstopBlockedReason in the MCP idle tool result — a held lease names itself
An Agent's recent actions/var/log/erun/audit.log or the API audit-events table
A merge happened — who advanced the queueSecurity events: mergequeue.advance
The runtime pod restartedkubectl describe pod -n <tenant>-<env> <pod> (Kubernetes events)
Trace request across servicesOpenTelemetry collector you installed; ERun adds nothing here
A build/release/push/deploy took too longThe step-timing table it prints on completion, or its JSON record under ~/.erun/timing/

What ERun doesn't ship​

To be explicit:

  • No log aggregator (Loki / Fluentd / Vector — install once per cluster).
  • No metrics database (Prometheus / VictoriaMetrics — install once per cluster).
  • No trace pipeline (OTel Collector — install per service as needed).
  • No dashboards (Grafana / Kibana / Datadog — install once per cluster).

The cluster owns the observability stack. ERun owns the audit trail — that's the one piece it does ship, because it's tied to the platform's identity and trust model.