Skip to main content

Metrics spec

For the Operator view, see Observability.

ERun's runtime pod exposes a Prometheus-format metrics endpoint. The full schema — endpoint, every metric name, every label, every type — is below. Application services declared in .erun/config.yaml use their own framework's metrics conventions; ERun has no opinion on those.

Endpoint​

PropertyValue
Listenererun-devops container, port 9100, bound 0.0.0.0 (not loopback — a cluster Prometheus reaches it over the pod network).
Path/metrics.
FormatPrometheus text exposition format, v0.0.4.
AuthenticationNone — the listener carries no token or mTLS check, matching Prometheus scrape convention. Reachability is instead scoped by the runtime chart's NetworkPolicy, which selects this pod for Ingress and admits traffic one rule at a time: ssh (sshPort, default 17022) and mcp (mcpPort, default 17000, present only while mcpEnabled=true) are re-permitted by number from any source, and the metrics listener (metricsPort, default 9100) only from pods in the same namespace or from a pod in a namespace carrying the label network-policy/erun-metrics-scraper: "true". A pod selected by any Ingress policy is default-deny for every port no rule names, so the policy — not the port's own configuration — is what isolates this pod: ssh and mcp stay reachable only because they are named in the policy, and every other listener the pod runs, including the dind sidecar's unauthenticated daemon on 2375, is closed to every other pod. Adding a listener to this pod therefore means adding its port to this policy. Set metricsEnabled=false on the chart to remove the listener (and the NetworkPolicy) entirely; the pod then carries no policy at all and its ingress returns to unrestricted.
Scrape cadenceSet by the caller (Prometheus). The runtime pod does not enforce a minimum.

A typical Prometheus scrape configuration:

- job_name: erun-runtime-pod
kubernetes_sd_configs:
- role: pod
namespaces:
names: [] # all namespaces
relabel_configs:
- source_labels: [__meta_kubernetes_pod_container_name]
regex: erun-devops
action: keep
- source_labels: [__meta_kubernetes_namespace]
target_label: namespace
- source_labels: [__meta_kubernetes_pod_name]
target_label: pod

For Prometheus running outside the target namespace (the usual case — one cluster-wide Prometheus scraping every tenant's runtime pods), label Prometheus's own namespace once: kubectl label namespace <prometheus-namespace> network-policy/erun-metrics-scraper=true. Without that label, the scrape config above resolves targets fine (namespace listing is a Kubernetes API concern, not a network one) but the actual scrape connection is refused by the target pod's NetworkPolicy.

Metrics​

Every series is labelled with the env's tenant + environment, derived from ERUN_TENANT and ERUN_ENVIRONMENT at process start. Additional per-metric labels are listed below.

erun_idle_eligibility​

PropertyValue
Typegauge
Labelstenant, environment
Range0 (env not eligible for idle-stop) or 1 (eligible).
Sourceemcp runs its own 10-second ticker that calls the same read-only idle-status resolution the idle MCP tool exposes (eruncommon.ResolveStoredEnvironmentIdleStatus; see Idle policy), so this gauge can never disagree with what idle reports. Sampled once immediately at boot as well, so a scrape right after pod start does not read a stale zero-value gauge as "eligible".
Reset onProcess restart.

erun_terminal_input_seconds_since_last​

PropertyValue
Typegauge
Labelstenant, environment
Unitseconds.
SourceWall-clock now minus the latest of the ssh, cli, and mcp activity kinds' last-recorded timestamps — the same union Idle policy · last_terminal_input already specifies (a keystroke, a successful in-pod erun invocation, or a non-idle-probe MCP call), read from the same on-disk activity snapshots the idle MCP tool exposes rather than a metrics-only redefinition. Sampled on the same 10-second ticker as erun_idle_eligibility.
Initial value0 immediately after pod start (no activity of any of the three kinds observed yet).

erun_traffic_window_bytes​

PropertyValue
Typegauge
Labelstenant, environment
Unitbytes.
SourceThe delta between two samples taken 60 seconds apart, not a fixed byte counter of any single socket: MCP bytes are counted directly in emcp's own HTTP middleware (request Content-Length plus response bytes written); SSH bytes are read from the ssh-proxy's cumulative activity-snapshot counter (erun-common/activity.go) and diffed the same way.
Window definitionTumbling 60-second window, sampled on a fixed in-process ticker independent of scrape timing. Each scrape returns the last completed window's byte count, not the one still accumulating — a scrape landing early in an otherwise-busy window can therefore read 0 even though traffic flowed a few seconds earlier; the next tick picks it up.

erun_mcp_calls_total​

PropertyValue
Typecounter
Labelstenant, environment, tool, result
tool valuesThe MCP tool name as registered in erun-mcp/server.go — see MCP overview for the current list rather than duplicating it here; a prior version of this page listed 13 example names, several of which (logs, open) do not exist as tool names and several of which (most of the surface) it omitted.
result valuessuccess, error, or dry_run. Derived from guardTool, the one dispatch point every tool call passes through: an error return is error; absent an error, a tool whose result carries the shared CommandOutput.Executed field (the delivery tools — build, push, deploy, doctor, pin, terraform, and the rest) reports dry_run when Executed is false; every other tool (read-only tools with no preview concept, e.g. idle, list, version, usage) reports success whenever it returns no error.
Reset onProcess restart. (Counters are monotonic within a process.)

erun_audit_events_total​

PropertyValue
Typecounter
Labelstenant, environment, action, actor_kind, result
action valuesToday, only the mcp.<tool> namespace (the same dispatch point erun_mcp_calls_total reads, guardTool) — one increment per call, both allowed and denied. The erun.<command> and api.<resource>.<verb> namespaces Audit log format describes are not yet wired into this counter: the CLI has no in-pod audit writer of its own yet, and the durable, API-side audit_events table has no CLI/MCP caller yet either (see that page's "(Planned.)" note) — this counter is a separate, lighter-weight in-process tally of the MCP edge's own authorization decisions, not a read of that table.
actor_kind valuesagent for every sample today, because this counter's only source (the MCP edge) is Agent-only traffic by construction — an Operator's actions go through the CLI, which this counter does not read from yet. operator is reserved for when a CLI audit source is wired in.
result valuessuccess, error (including a capability refusal), or dry_run — same derivation as erun_mcp_calls_total, except a refusal (the call never reaches the tool handler) also counts here as error, since it is still an audited decision.

Cardinality envelope​

The label sets above are bounded:

LabelsCardinality bound
tenant × environmentOne per running env.
tool≤ 100 (the MCP-registered tool set, ~90 today including one-release deprecated aliases — see erun-mcp/server_test.go's wantRegisteredTools for the exact, tested count).
result3 (success, error, dry_run).
actionSame bound as tool today (only the mcp.<tool> namespace is wired in); will grow once erun.<command> and api.<resource>.<verb> are wired.
actor_kind1 today (agent only); 2 (operator, agent) once a CLI audit source is wired in.

A single-env Prometheus deployment emits at most ~600 series from this pod today: 3 gauges, erun_mcp_calls_total at ≤300 (tool × result), and erun_audit_events_total at ≤300 (action × actor_kind × result, currently 1 actor_kind value). Both counters' bounds double once the erun.<command> and api.<resource>.<verb> action namespaces and the operator actor kind are wired in — comfortably inside Prometheus's per-target defaults either way.

Conformance​

The runtime pod's metrics emitter conforms to:

  • Prometheus text exposition format v0.0.4.
  • OpenMetrics 1.0 is not declared (no # TYPE lines beyond gauge/counter; no # UNIT lines).
  • _total suffix convention is used on counters.

Application-service metrics​

The metrics above describe the runtime pod only. Application services declared in .erun/config.yaml use whatever metrics convention their framework ships with (Go prometheus/client_golang, Node prom-client, Java Micrometer, etc.). ERun's helm chart conventions do not standardise this; consult the cluster's Prometheus configuration for how application pods are scraped.

See also​

  • Observability — Operator-facing summary.
  • Idle policy — the predicate that drives erun_idle_eligibility.
  • Audit log format — the durable, API-side audit trail and its <namespace>.<verb> action-verb convention; erun_audit_events_total follows the same convention but is a separate, in-process count, not a read of that table (see the note on its action values above).