Skills spec
For the Operator view, see Skills.
A skill is a directory containing a SKILL.md plus optional reference content. ERun owns one canonical source of skills (erun-skills/skills/<name>/ in the sophium/erun repository) and publishes them through two paths: baked into the runtime image for in-pod use, and via a Claude Code plugin marketplace for laptop use. The Agent — Claude Code, Codex, or whatever the env's aitool selects — discovers them through its own skill-loading convention.
This page specifies: the on-disk format, the per-tool discovery paths, the deployment mechanism the runtime chart uses, the marketplace distribution contract, the layering rules, and the built-in skill catalogue.
SKILL.md format
A skill bundle is a directory:
<skill-name>/
├── SKILL.md # required; the entrypoint the Agent reads
├── references/ # optional; longer-form reference content the SKILL.md cites
│ └── ...
├── examples/ # optional; worked examples the Agent can use as starting points
│ └── ...
└── scripts/ # optional; helper scripts the Agent can invoke as part of applying the skill
└── ...
SKILL.md has YAML frontmatter followed by a markdown body:
---
name: go-service
description: Write a conformant Go HTTP service with the ERun multi-stage Dockerfile and helm chart layout. Use when the Operator asks to add a Go service, gRPC service, or background worker written in Go.
---
# Go service
When you're adding a Go service to an ERun project, follow this layout.
## Source module
Live at `<projectRoot>/<name>/`. Layout:
…
## Dockerfile
Live at `<projectRoot>/<tenant>-devops/docker/<name>/Dockerfile`. Multi-stage…
…
| Field | Required | Validation |
|---|---|---|
name | yes | Matches ^[a-z][a-z0-9-]*$. Must equal the parent directory name. Uniquely identifies the skill. |
description | yes | One sentence; under 200 characters. This is what the Agent reads to decide whether the skill applies — phrase it as when to use the skill, not what the skill contains. |
The body is plain markdown — the same content the Agent would read as instructions. Sections, lists, code blocks, links to reference files all work.
Per-tool discovery paths
The erun-devops container's entrypoint copies each skill into the conventional location for the env's configured Agent on every env start. The Agent picks them up automatically — no extra flag or config needed.
| Agent | Discovery path inside the runtime pod |
|---|---|
claude | ~/.claude/skills/<skill-name>/SKILL.md |
codex | ~/.codex/skills/<skill-name>/SKILL.md |
Both paths are installed in parallel from the same /etc/erun/skills/<name>/ source baked into the image. A generic canonical path under ~/.config/agent-skills/ is (Planned.) for future tools.
Source module
Skills live in exactly one place in the source tree: erun-skills/skills/<skill-name>/ in sophium/erun. Both the runtime image and the plugin marketplace vendor from this directory. Editing a skill is a single-file change in that module; the two distribution paths pick it up automatically.
Module guidance: erun-skills/AGENTS.md.
Deployment mechanism
Two paths deliver the skill set to the Agent:
In-pod (runtime image)
The runtime Dockerfile vendors the whole tree with one line:
COPY erun-skills/skills /etc/erun/skills
On every entrypoint run, initialize_claude_config and initialize_codex_config run skills-install.sh over every subdirectory under /etc/erun/skills/, installing each skill into ~/.claude/skills/<name>/ and ~/.codex/skills/<name>/. Supporting files (templates, helper scripts) inside the skill directory ship with the skill automatically.
The install both installs a skill when absent and refreshes it when the baked copy changed, so a rebuilt image's updated skill reaches existing envs — while preserving in-pod edits. Provenance is tracked per skill by recording the baked SKILL.md hash in a .erun-skill-baked-sha256 marker: a copy whose SKILL.md still matches its marker is unmodified since erun installed it and is refreshed to the baked version, while one that differs was edited in-pod and is left untouched (a legacy copy with no marker is treated as unmodified and adopted on the first refresh). So an un-edited skill tracks the image across upgrades, and a skill you edit inside a running env survives both pod restarts and image rebuilds.
Host orchestrator (desktop)
The desktop app installs the same canonical skills into the host's ~/.claude/skills/<name>/ for host-side orchestrator sessions, using the identical marker-based install-or-refresh — so a host orchestrator tracks the latest skill on each launch while preserving any host-side edits.
The source it installs from resolves in this order, first match wins:
ERUN_SKILLS_DIR, if set — taken verbatim, with no fallback, so pointing it at an empty directory installs nothing rather than silently resolving something else.- The
erun-skills/skillsdirectory of the checkout the desktop binary was built from, stamped into the binary byerun-ui/build.sh/build.ps1. This is what keeps a desktop that runs from outside its checkout — the usual case, since the built bundle is copied elsewhere to run — installing the skills its own build ships. erun-skills/skillsfound by walking up (max 8 levels) from the running executable, for a binary that was built without the stamp but sits inside a checkout.
If none resolves, the orchestrator still launches and the skills already installed are left alone — but the condition is reported, not silent: a warning notification naming the checkout that was expected, where the executable looked, and the two recoveries (set ERUN_SKILLS_DIR, or rebuild with erun-ui/build.sh / build.ps1) is posted once per desktop run, and every occurrence is logged. A build that silently stopped refreshing skills is indistinguishable from one where the skill had not changed. A desktop installed from a package manager carries no checkout, so ERUN_SKILLS_DIR is its only source.
The desktop also writes a SessionStart hook into the shared orchestrators workspace's .claude/settings.json that injects the operating contract — it prints the workspace CLAUDE.md, and then the orchestrator's own CLAUDE.<id>.md role file if it has one, to the session on every session boundary (Claude Code's SessionStart fires with source startup, resume, clear, and compact, and all four are re-injected) — so the contract and the standing role are always already in context rather than a erun-orchestrate skill the model is merely asked to load and could skip, and neither one silently drops out of context after a /clear or a compaction. The hook prints the files directly (plain stdout, so no additionalContext size cap), falling back to a short directive if the shared CLAUDE.md is ever missing.
Guidance is two layers, injected shared-then-specific so the ordering is the precedence rule. CLAUDE.md — described above — is the one shared contract every orchestrator obeys; erun rewrites it on every launch, so an edit there is discarded on the next one. CLAUDE.<id>.md, in the same shared orchestrators workspace, is this orchestrator's own standing role: erun seeds it once, with a short comment explaining the convention, and never rewrites it afterward — the deliberate inverse of the shared file. The hook prints it immediately after CLAUDE.md, so anything in it can add to the contract or override a line of it. <id> is the orchestrator's internal id, not its display name — an orchestrator the sidebar shows as erun-admin can carry id erun-issues, so its role file is CLAUDE.erun-issues.md — which is why the desktop's orchestrator management dialog resolves and opens both files by id rather than asking the operator to know the filename convention (see Desktop app · Orchestrators for the Operator-facing view).
Which conversation a launch resumes
For the Operator view, see Desktop app · The conversation an orchestrator comes back to.
Every launch of an orchestrator resumes a named conversation rather than --continue's most-recent one, which in a shared workspace collapses every orchestrator onto one session. The name is resolved from three sources, in this order.
The anchor (derived). uuid5(6f7e9c2a-1b3d-4e5f-8a9b-0c1d2e3f4a5b, <orchestrator id>). A pure function of the id, so it is identical on every launch and on every machine, needs nothing on disk, and cannot be written by another session. It answers which conversation is this orchestrator's by convention — and only that. A transient (Investigate) session has no id, so it has no anchor and starts unpinned.
The tracked conversation (live). The harness does not always adopt the id it is asked to resume; a launch that asks for the anchor can end up writing to a conversation of its own, after which the anchor's transcript stops growing while the work accumulates elsewhere. Only the session knows which one it is writing to, so the session reports it:
- The desktop mints a launch nonce (a v4 UUID) per launch, exports it as
ERUN_ORCHESTRATOR_LAUNCHbesideERUN_ORCHESTRATOR_ID, and writes it into its own durable open-set record (orchestrator-open.json,launchIdper entry) as the session is spawned. - A hook installed on
SessionStart,UserPromptSubmit,PreToolUse,PostToolUseandStopreadssession_idfrom the hook's own stdin JSON and writes{"conversationId":…,"launchId":…,"atUnix":…}to<UserConfigDir>/ERun/orchestrator-live/<id>.json. Session boundaries and turn boundaries, because the id can change mid-run; bare shell, so it works with erun offPATH; every failure path is|| true, since a hook that could wedge a session costs more than a missed record. - A record counts only while
launchIdequals thelaunchIdon that orchestrator's open-set entry. That makes the two halves independent: the desktop writes the nonce, the session writes the conversation, and neither is authoritative alone. A session that never saw this launch's nonce cannot claim the orchestrator's conversation however it came by the id, and a store whose writer is removed stops counting on the very next launch — the two failure modes of the record this replaces.
The attachment (operator). A conversation the Operator attached through the manage dialog. Durable on the open-set entry (attachedConversationId), so a later launch honours it instead of recomputing the anchor, and cleared by asking for the default back. It outranks both of the above.
Resolution is: attachment if usable → tracked record if confirmed → anchor. A candidate ahead of the anchor must clear two checks — its transcript is still on disk, and no other configured orchestrator has a claim on it (that orchestrator's anchor, its attachment, or its own tracked record) — because the cost of the wrong answer is another orchestrator's history presented as this one's. Recency is never an input: the transcript directory holds conversations belonging to several orchestrators and to none, so "the newest file" is usually somebody else's.
Nothing falls through silently. A resolution that is not the plain answer carries an operator-facing notice — surfaced beside the orchestrator list on a restore, and as a notification on an ordinary start:
| Outcome | Notice |
|---|---|
| Tracked conversation resumed | Names the conversation resumed and the anchor it beat. It is the good outcome; the two ids disagree, and only the Operator can tell whether the winner holds the work they expect. |
launchId missing or from another launch | Resumes the anchor, names the unconfirmed conversation, and says nothing has confirmed it since. |
| Tracked transcript no longer on disk | Resumes the anchor and says the transcript is gone. |
| Tracked conversation claimed by another orchestrator | Resumes the anchor and names the orchestrator that owns it. |
| Attachment unusable (either of the two checks above) | Resumes the anchor and says the Operator's own choice could not be honoured. |
The restart hand-off (orchestrator-restore/<id>.json) records the conversation the running session reports being on under its launch, not the id it was spawned with — a restart is the one path that must reach the session that asked for it. The crash respawn keeps the same nonce, since it is the same launch continuing.
Listing and attaching. ListOrchestratorConversations(<id>) reports what one orchestrator could resume, and is the surface that makes a wrong resume correctable rather than terminal:
| Field | Meaning |
|---|---|
resuming / resumingSource | The conversation a launch resolves right now, and which of attached / tracked / derived produced it. |
attached | The Operator's standing choice, empty when there is none. |
notice | The same notice the resolution would surface, so the picker explains the state it is showing. |
conversations[] | Per conversation: conversationId, folder, lastWrittenUnix, sizeBytes, excerpt, role, resuming. |
omittedNotMine / omittedForCap | How many rows were withheld because another orchestrator claims them, and how many older ones the 60-row cap dropped. Stated, never silently trimmed. |
transcriptsRoot | ~/.claude/projects, the directory the rows were read from. |
role is one of attached, live (tracked and confirmed), stranded (recorded as live by a launch nothing can confirm — the row most likely to hold unreachable work), derived, or unowned. Enumeration globs ~/.claude/projects/*/*.jsonl across every project directory, since a conversation id is globally unique and an orchestrator's own conversation may have been started elsewhere. folder comes from the transcript's own recorded cwd, never decoded from the project directory name — the harness encodes a path by replacing each separator with -, which no decode can round-trip for a path that already contained one. folder and excerpt are read from the first 128 KiB only, so listing a directory of multi-megabyte transcripts stays cheap; the anchor is always a row even before its first write.
AttachOrchestratorConversation(<id>, <conversationId>, cols, rows) records the choice and restarts the orchestrator in that conversation; DetachOrchestratorConversation(<id>, cols, rows) clears it and restarts on whatever resolves. Attaching a conversation another orchestrator claims is refused, with the owner named — recovering a mis-attached orchestrator is what this exists for, and crossing two orchestrators is what it exists to prevent.
Periodic pacing re-statement and crash auto-resume
For the Operator view, see Workflow · Level 3.
The SessionStart hook above closes the gap at session boundaries (startup, resume, clear, compact). It re-states nothing in the (often long) stretch between two boundaries, so the desktop's session-heartbeat poller — the same 15-second tick that reconciles each orchestrator's busy/shell-activity reports — also re-states the pacing contract directly into a live session whose own report has gone stale, and separately relaunches a session whose process exited from a failure.
Staleness read. Each orchestrator's turn-boundary activity report (the same file the busy-spinner reconciler reads, {"busy":…,"atUnix":…}) is read again here, independent of that reconciler's own busy/idle staleness bounds: a report not refreshed for 10 minutes (orchestratorPacingStaleAfter) reads as quiet regardless of whether its last write said busy or idle. The reference point is the later of the report's own write time and the session's launch time, so a session with no report yet (its first turn hasn't reached a boundary) is not nudged before it has had ten minutes to.
Nudge. For each live, non-transient orchestrator whose report is stale, the desktop writes this text into the session's pty, then — after a short settle, as a separate write — a bare carriage return to submit it:
Keep pacing yourself, on connection errors wait and resume, do not exit this loop. If the assigned task is already complete and verified, say so in one line and stop.
A background shell reported running for that orchestrator (orchestrator-shell-activity) no longer holds the nudge back. It used to: the reasoning was that a shell left running was evidence the orchestrator meant to be quiet, but a background shell is a fact about the shell, not about the turn behind it — a long-running build is the single most likely thing still running when a turn dies mid-response, so the old rule suppressed the nudge exactly when it was needed most. The shell-activity report is unchanged and still drives the shell indicator; pacing simply stopped reading it.
Each nudge appears in the pane as a dim marker line naming the attempt count and the measured quiet period — never the ten-minute constant — so a session quiet for twenty minutes reads as twenty minutes, not as the contract having been honoured on time:
── pacing nudge N/6 sent — no activity report for 19m42s ──
If the orchestrator's last report said busy: true and was never followed by an idle one, the marker names that too, since a report stuck on "busy" for a full staleness period usually means the harness never reached its own turn boundary (a dropped connection, a crash) rather than a turn that finished and simply stopped reporting:
── pacing nudge N/6 sent — no activity report for 19m42s — last report said mid-turn, so the turn may have died without one ──
Bounds.
- One nudge per orchestrator per tick, and never for a transient (Investigate) session — it has no
idand its own bounded lifecycle belongs to the investigation registry, not this one. - A cap of 6 consecutive un-answered nudges (
orchestratorPacingMaxNudges) — about an hour at the 10-minute interval. Crossing the cap posts a warning notification and a distinct pane marker instead of a 7th nudge, and the cap latches (no repeat notification on later ticks) until it is rearmed by either:- a fresh turn-boundary report whose own timestamp is later than the last nudge and says
busy: true— evidence the session resumed on its own, or - real operator input sent into that same pane (
SendSessionInput) — evidence the operator is now at the keyboard.
- a fresh turn-boundary report whose own timestamp is later than the last nudge and says
Every decision is logged, not just the ones that nudge. The reconciler names the reason it did or didn't nudge in the desktop's durable app log, once per transition rather than once per 15-second tick:
erun-app: orchestrator <name> pacing decision=<reason> quiet=<duration|unknown>
The reasons are fresh (it moved recently), not-alive (the desktop holds a session for it that is no longer live), already-capped, cap-crossed, nudge, env-busy (a linked environment it dispatched to is still busy on its own lease, so the silence is accounted for), and unreachable-from-transport. A stalled pane that never nudges (because the session itself is no longer alive) is therefore distinguishable from a healthy, ordinarily quiet one by reading that log, instead of both looking identical from outside.
Coverage is the whole configured population, not only the sessions that desktop launched. An orchestrator whose session was started outside it — in a terminal, or by a previous desktop instance — writes the same activity and live-conversation records a launched one does, so the desktop reads and shows it, but it holds no PTY to nudge and no session of its own to keep a nudge count on. It is logged as unreachable-from-transport, never as not-alive: the session is usually alive and reporting, and it is the transport that cannot reach it, not the session that is gone. quiet=unknown appears in place of a duration when nothing has ever reported for it, since there is then no reference point to measure one from.
This pass is neither the only way to trigger it nor fixed at these values. The message, the ten-minute threshold, and the cap of 6 are the user config's whip section (eruncommon.WhipConfigOverride) — unset by default, so an install that configures nothing keeps exactly the text and bounds above. erun whip is the manual, visible counterpart: an operator-triggered pass over the same decide/report core (eruncommon.DecideWhip) that pushes every environment's own AI session (over its whip MCP tool) plus every persisted orchestrator, reporting each target's outcome by name — pushed, or skipped and why — rather than only reporting that it ran. A session nudged this way sees the exact same text as the automatic pass; there is nothing in the message itself that distinguishes an automatic re-statement from an operator-triggered one, so treat any pacing nudge as erun restating a contract you already agreed to, not as a message from your operator.
Auto-resume after a crash. An orchestrator's managed session carries a respawn closure (the same tryReconnect mechanism environment tabs use to recover a dropped pod session), gated so it only ever fires for a genuine failure:
- A clean exit never respawns.
Wait()returningnil(terminalSessionExitReasonanswers"") means the operator quit the TUI from inside the pane, not a crash. - A torn-down registration never respawns, even carrying a real failure reason: the closure refuses unless the orchestrator's config-map entry and its managed-terminal registration are both still exactly what they were when the session was spawned.
StopOrchestratordeletes both under the same lock a Stop takes, so Stop always refuses its own respawn — including a respawn already in flight when Stop is called, since the same check runs again after the relaunch attempt returns. - A transient (Investigate) session never respawns — it carries no respawn closure at all; its lifecycle is the investigation registry's to end.
When it does fire, the closure relaunches through the same launch path a fresh start uses, resuming whatever conversation the crashed session last reported under this launch's nonce rather than only the id it was spawned with (see Which conversation a launch resumes), and hands the relaunched session a resume prompt distinct from the restart-hand-off one above (it names no return note, since nothing wrote one for an involuntary crash):
This session's process just exited unexpectedly and was relaunched automatically. Resume the conversation exactly where it left off and carry any in-progress task through to its verified end without waiting to be asked.
The crash-fallback shell command itself was also closed: buildOrchestratorLaunch's primary launch already resumes a pinned session id (--resume <id> if the conversation exists, --session-id <id> to create it) with a ||-chained (PowerShell: $LASTEXITCODE-checked) shell-level fallback for the same single invocation. That fallback now retries the identical pinned invocation rather than ever dropping to an unpinned claude, since an unpinned session is what the live-conversation recorder (above) would then report as this orchestrator's conversation — silently swapping it onto an amnesiac one on the very first in-shell retry. Only a transient/legacy launch (no pinned id) still falls back to plain claude.
Laptop (plugin marketplace)
The repo-root file .claude-plugin/marketplace.json publishes erun-skills/ as the erun-tools plugin via a git-subdir source. Users add the marketplace and install:
/plugin marketplace add sophium/erun
/plugin install erun-tools@sophium/erun
Skills become invocable as /erun-tools:<skill-name> (e.g. /erun-tools:erun-file-issue). See Marketplace distribution below for the full schema and update flow.
Layering (Planned.)
Tenant-level skills (<projectRoot>/<tenant>-devops/skills/<name>/) and project-level skills (<projectRoot>/.erun/skills/<name>/) are reserved as future layers above the in-pod baked set. The current implementation ships only the runtime-image baked set; tenant and project layers are not yet mounted.
Marketplace distribution
The plugin is published from sophium/erun itself — the repo is the marketplace. No second repository to maintain.
.claude-plugin/marketplace.json
At the repo root. Schema (current, as published):
{
"$schema": "https://anthropic.com/claude-code/marketplace.schema.json",
"name": "sophium/erun",
"description": "ERun skills for Claude Code …",
"plugins": [
{
"name": "erun-tools",
"description": "…",
"category": "development",
"source": {
"source": "git-subdir",
"url": "https://github.com/sophium/erun.git",
"path": "erun-skills",
"ref": "main",
"sha": "<commit-sha>"
},
"homepage": "https://github.com/sophium/erun/tree/main/erun-skills"
}
]
}
Fields:
| Field | Required | Notes |
|---|---|---|
name (marketplace) | yes | sophium/erun. Used by /plugin marketplace add and as the suffix in /plugin install <plugin>@<marketplace>. |
owner | recommended | Display info for the catalogue UI. |
plugins[].name | yes | erun-tools. Used as the namespace prefix for skill invocation (/erun-tools:<skill>). |
plugins[].source.source | yes | git-subdir so the plugin can live in a subdirectory of the marketplace repo. |
plugins[].source.url | yes | HTTPS git URL — must be cloneable without credentials for public marketplaces. |
plugins[].source.path | yes | erun-skills. The plugin root inside the marketplace repo. |
plugins[].source.ref | yes | Branch (main) used to resolve the SHA on update. |
plugins[].source.sha | yes | Pinned commit hash. Users only see updates when this changes. |
plugins[].homepage | recommended | Browseable URL for the plugin source. |
erun-skills/.claude-plugin/plugin.json
The plugin manifest. Skills are auto-discovered from erun-skills/skills/; they do not need to be enumerated.
{
"name": "erun-tools",
"version": "1.0.0",
"description": "…",
"homepage": "https://github.com/sophium/erun/tree/main/erun-skills",
"repository": "https://github.com/sophium/erun",
"license": "MIT"
}
Update flow
- Edit a skill in
erun-skills/skills/<name>/SKILL.md. - Commit and merge to
main. - The release flow bumps
source.shain.claude-plugin/marketplace.jsonto the merge commit. (Until release automation handles this end-to-end, the bump lands in the same PR as the skill edit.) - Users see the update on next
/plugin marketplace update sophium/erun.
Auto-update for third-party marketplaces is disabled by default in Claude Code; users either opt in via the /plugin UI or run /plugin marketplace update manually.
Install commands
/plugin marketplace add sophium/erun # add once
/plugin install erun-tools@sophium/erun # install plugin
/reload-plugins # pick up mid-session
/plugin marketplace update sophium/erun # refresh catalogue
/plugin uninstall erun-tools@sophium/erun # remove
Error behaviour
| Failure | What the user sees | Recovery |
|---|---|---|
| Marketplace repo unreachable (network, auth) | /plugin marketplace add fails with the upstream git clone error. | Check network / gh auth status; retry. |
marketplace.json malformed JSON | /plugin marketplace add fails with the parse error and the offending line. | File an issue against sophium/erun. |
source.sha no longer reachable (history rewritten) | Install fails with a "commit not found" error. | Re-run /plugin marketplace update sophium/erun to fetch the latest SHA. |
| Skill name collision with an existing user-installed skill of the same name | Claude Code namespaces the plugin-shipped skill as /erun-tools:<name>, so collisions are not possible at the invocation layer. | n/a. |
| Plugin install succeeds but no skills appear | /reload-plugins may be needed. If still missing, the plugin manifest may have rejected the install — check /plugin UI for an error entry. | Run /plugin and inspect the Errors tab. |
Codex distribution (Planned.)
Codex CLI does not have an analogous plugin marketplace yet. Inside a deployed env, Codex receives the same skills via the runtime-image baked install — no extra step. For laptop Codex use, copy erun-skills/skills/<name>/SKILL.md into ~/.codex/skills/<name>/ manually until upstream Codex ships plugin support.
Layering rules
| Conflict | Resolution |
|---|---|
A project skill has the same name as a built-in. | Project skill wins. The built-in is hidden in this env. |
A tenant skill has the same name as a built-in. | Tenant skill wins; project skill (if also present) wins over tenant. |
| Two skills in the same layer have the same name. | Sort lexicographically by source path; later wins. (This is a misconfiguration — flag in erun doctor.) |
A skill's name frontmatter doesn't match its directory. | The skill is skipped; erun doctor reports SKILL_NAME_MISMATCH. |
Built-in skill catalogue
The current v1 set, shipped both in the runtime image (/etc/erun/skills/) and via the plugin marketplace. Skills come in two semantic kinds:
- Blueprint skills — package ERun's accumulated best practices for building complex industry-strength solutions.
- Workflow skills — let users participate in ERun's processes (report problems, share improvements back so other users benefit).
erun-file-issue
| Field | Value |
|---|---|
| Kind | Workflow — participate in ERun's issue-reporting process. |
| Source | erun-skills/skills/erun-file-issue/SKILL.md |
| Description | "Register or file a bug or feature request for the ERun project itself on GitHub." |
| Triggers | "file an erun bug", "file an erun feature", "register erun bug", "register erun feature", "open an erun issue" |
| Inputs | Issue title; what-happened / what-expected / reproduction (or feature goal + acceptance criteria) |
| Outputs | gh issue create --repo sophium/erun --label bug (or --label enhancement) invocation with a templated body. Body adapts to context: inside an env it includes ${ERUN_TENANT}, ${ERUN_ENVIRONMENT}, and the ERUN_* env dump; on a laptop it omits those. |
| Error behaviour | gh not installed or unauthenticated → surfaces the gh auth status hint and stops. Title or body missing → re-prompts the user. |
erun-contribute
| Field | Value |
|---|---|
| Kind | Workflow — lets users share improvements back to the platform so other users benefit. |
| Source | erun-skills/skills/erun-contribute/SKILL.md |
| Description | "Contribute a change to the ERun platform itself — create a new GitHub issue against sophium/erun that captures the work, clone the repo, implement the change following its AGENTS.md rules, and submit a pull request back." |
| Triggers | "contribute to erun", "make a change to erun", "work on erun", "land a fix in erun", "submit a PR to erun", "propose an improvement to erun" |
| Inputs | One- or two-sentence description of the change; sentence-style title; issue type (bug / feature / enhancement); short kebab-case description for the branch name (defaulted from the title) |
| Outputs | A newly-filed issue against sophium/erun (Step 1), then a cloned repo at ~/git/erun, branch feature/<n>-… or bug/<n>-…, code change, make integration-test run, push, PR via gh pr create --repo sophium/erun --base main with Closes #<n> in the body |
| Error behaviour | gh issue create fails (auth, network, label not allowed) → stops; does not proceed to clone without an issue number to anchor the PR. make integration-test fails → does not push; surfaces the failure. PR title contains an agent marker ([claude], [codex]) → re-prompts the user per AGENTS.md § "Pull Request Titles". |
Semantic: erun-contribute is initiator-driven — the same person who runs it both files the issue and ships the PR. For reporting a problem without intent to follow up, use erun-file-issue instead. For picking up an issue someone else filed, no skill applies; the user clones, branches, implements, and PRs directly.
Key contract: the skill explicitly reads the cloned repo's AGENTS.md and every applicable subtree AGENTS.md each time it fires. Claude Code does not auto-reload CLAUDE.md after a cd mid-session, so this read step is binding.
erun-blueprint-agents
| Field | Value |
|---|---|
| Kind | Blueprint — packages ERun's orientation for a tenant repo (environment model, core commands, artifact locations, the one-version pinning contract, skill pointers). |
| Source | erun-skills/skills/erun-blueprint-agents/SKILL.md + templates/ |
| Description | "Blueprint the repo-root agent-guidance file for an erun tenant project — a canonical AGENTS.md plus a CLAUDE.md symlink (one file, one source of truth) pre-populated with orientation on working in the erun environment — the tenant/environment model, the core erun commands (build, deploy, terraform apply, list, doctor, open), where the deploy artifacts live and the one-version pinning contract, and pointers to the other skills. Idempotent — reconciles a missing or broken symlink and never clobbers hand-authored guidance." |
| Triggers | "scaffold root AGENTS.md", "add erun agent guidance to this repo", "orient this tenant repo for agents", "create the repo-root CLAUDE.md", "set up AGENTS.md for this erun project" |
| Inputs | The repo root (default: current working directory); the tenant + environment to name in the guidance — resolved from ${ERUN_TENANT}/${ERUN_ENVIRONMENT} in-pod or erun list on a laptop, else left as generic <tenant>/<env> pattern text. |
| Outputs | A repo-root canonical AGENTS.md (rendered from templates/AGENTS.md, with <tenant>/<env> substituted where resolved) plus a CLAUDE.md same-directory relative symlink to it (git mode 120000; blob content is the bare filename AGENTS.md) — matching erun's own repo convention that AGENTS.md is canonical and CLAUDE.md points at it, not the reverse. The content covers the tenant/environment model (agent env vs runtime env, working inside the pod via erun open, the runtimeversion pin), the core commands (erun list/open/build/deploy/terraform apply/doctor), the deploy-artifact locations (terraform-<tenant>/<env>/, <tenant>-devops/k8s/<tenant>-<component>/, <tenant>-devops/docker/<tenant>-devops/Dockerfile), the one-version pinning contract (the Terraform module ?ref, each Helm umbrella Chart.yaml version:, the build-env Dockerfile FROM, and the env runtimeversion, all bumped together), and pointers to the other skills. Both files are committed to git — never written to ${ERUN_OUTPUTS_DIR}. The other blueprint skills (erun-blueprint-api/-rls-db/-docs/-platform, erun-build-env) point at this skill so any scaffold path yields root guidance. The generated file documents the Windows symlink caveat (a symlink-less Windows checkout materializes CLAUDE.md as plain text containing AGENTS.md; read AGENTS.md directly there). Idempotent on re-run: a correct AGENTS.md + CLAUDE.md -> AGENTS.md symlink is left untouched, and a missing/broken symlink over a present canonical file is recreated. |
| Error behaviour | A hand-authored (regular, non-symlink) AGENTS.md/CLAUDE.md already at the root → not a stop; never overwrite — report it and offer to fold the erun orientation in with the user's confirmation. Tenant/env unresolvable → write the file with generic <tenant>/<env> pattern text (still valid). ln -s unavailable on a Windows shell without symlink support → the canonical AGENTS.md still works standalone; create the symlink via git on a platform that supports it. Not inside a git repository → write the files and surface that they must be committed once the repo is initialized. |
Do not confuse this skill with a reusable agent — a different artifact kind, specced at Reusable agents spec. This skill scaffolds a repo's AGENTS.md guidance file; a reusable agent (e.g. erun-builder, erun-reviewer) is a standing subagent role.
erun-blueprint-rls-db
| Field | Value |
|---|---|
| Kind | Blueprint — packages ERun's accumulated best practices for multi-tenant PostgreSQL. |
| Source | erun-skills/skills/erun-blueprint-rls-db/SKILL.md + templates/ |
| Description | "Build a multi-tenant PostgreSQL database module following ERun's blueprint — row-level security, Atlas migrations, UUIDv7 surrogate keys, shared timestamp trigger, separate erun_tenant / erun_operations PostgreSQL roles, and the canonical tenant/issuer/user bootstrap that erun-backend-db captures — and maintain, repair, and upgrade a module it previously produced by detecting existing artifacts and entering maintenance mode instead of stopping, filling blueprint gaps without clobbering the project's own tables or committed migrations, and re-pinning the module's own version axes — the PostgreSQL major and Atlas toolchain — to their targets (it has no erun-version coupling)." |
| Triggers | "build a multi-tenant postgres database", "create a tenant-scoped postgres schema with row-level security", "set up multi-tenant postgres migrations", "I need an erun-backend-db-shaped module", "build a multi-tenant rls db", "upgrade the multi-tenant postgres module", "repair the rls db module", "reconcile the tenant database schema to the blueprint", "bump the db module to <version>", "maintain the erun-backend-db-shaped module" |
| Inputs | Module name; target directory; list of tenant-owned tables; PostgreSQL major version (default 18) |
| Outputs | <module>/atlas.hcl, <module>/schema/{tables,indexes,triggers,rls,fks}/*.sql, <module>/schema/roles.sql, <module>/migrations/default/, <module>/AGENTS.md. Bootstrap tables (tenants, tenant_issuers, users, user_external_ids) plus one tables/indexes/triggers/rls set per user-supplied table. On an existing module (an atlas.hcl plus a schema/ tree) it enters maintenance mode instead of scaffolding: previews the plan, reconciles gaps against the current erun-backend-db blueprint (missing bootstrap tables, roles.sql, rls/context.sql, timestamp triggers, RLS ENABLE/FORCE + _tenant_policy/_operations_policy pairs, atlas.hcl src order), and re-pins the module's own version axes (the PostgreSQL major and Atlas toolchain) to their targets — never clobbering the project's own tables and never rewriting a committed migration (drift is corrected with a new forward atlas migrate diff). Cleanup removes only superseded scaffolding, never dropping a table or deleting a committed migration — a schema removal belongs in a reviewed forward migration, not a cleanup pass. |
| Error behaviour | Target dir already has atlas.hcl → not a stop; enter maintenance mode and reconcile gaps + re-pin in place, previewing before writing and never clobbering the project's tables or committed migrations. PostgreSQL < 18 detected → stop (native uuidv7() unavailable). atlas not installed → skip validate, surface install hint, continue. User-supplied table name collides with bootstrap names → stop and ask user to rename. |
erun-blueprint-api
| Field | Value |
|---|---|
| Kind | Blueprint — packages ERun's accumulated best practices for multi-tenant HTTP APIs. |
| Source | erun-skills/skills/erun-blueprint-api/SKILL.md + templates/ |
| Description | "Build or maintain a multi-tenant Go HTTP API service following ERun's blueprint — OIDC bearer authentication, tenant resolution from the token issuer, layered model / repository / service / routes structure, transaction-scoped PostgreSQL security context, identity resolution cache, and audit logging — and reconcile, repair, and upgrade a previously scaffolded service in place by realigning it to the current blueprint and refreshing the service's own dependency pins, without clobbering the project's own business logic (it is a standalone Go module with no erun-version coupling). Captures the patterns that erun-backend-api packages." |
| Triggers | "build a multi-tenant http api", "build a multi-tenant backend api", "create an erun-backend-api-shaped service", "I need a multi-tenant Go api with oidc auth and tenant rls", "upgrade the multi-tenant api", "repair the erun-backend-api-shaped service", "reconcile the api to the blueprint", "bump the api to <version>", "maintain the multi-tenant api" |
| Inputs | Module name; Go module path; target directory; OIDC issuers; initial entities (optional) |
| Outputs | <module>/go.mod, <module>/cmd/<module>/main.go, <module>/server.go, <module>/auth.go, <module>/oidc.go, <module>/identity_cache.go, <module>/api_path.go, <module>/audit.go, <module>/internal/{model,repository,routes}/..., <module>/AGENTS.md. Includes a working GET /v1/whoami endpoint; entity routes are produced per user-supplied entity. Its own Step 6 then composes erun-blueprint-service to add the <tenant>-devops/docker/<module>/Dockerfile and <tenant>-devops/k8s/<module>/ chart this skill's source-only output has no deploy artifacts for — without it, erun build/erun deploy have nothing to find. On an existing service (a go.mod/server.go/internal/repository/tx.go present) it enters maintenance mode instead of scaffolding: previews the plan, restores structural drift against the current blueprint (a missing layer, OIDC/authentication, authorization or audit middleware, tenant-from-issuer resolution, the TxManager.WithTx RLS security-context wiring, or the identity-resolution cache), and refreshes the service's own dependency require pins and go toolchain line, then re-proves with go mod tidy / go build / go test; it never clobbers the project's own domain entities or business logic. Cleanup removes only superseded generated files (preview-first), never the project's own code. |
| Error behaviour | Target dir already has go.mod → not a stop; enter maintenance mode and reconcile against the blueprint in place — fill structural drift and refresh the service's own dependency pins — without clobbering the project's own content. Empty OIDC issuer list → stop. Database side (matching erun-backend-db-shaped schema) missing → surface and offer to run erun-blueprint-rls-db first. go build fails after generation → surface compiler output; most common cause is module path mismatch. |
erun-blueprint-service
| Field | Value |
|---|---|
| Kind | Blueprint — packages ERun's component-naming and multi-stage-Dockerfile conventions for a deployable service. |
| Source | erun-skills/skills/erun-blueprint-service/SKILL.md + templates/ |
| Description | "Add a custom service's deploy artifacts — a multi-stage Dockerfile, a Helm chart, and per-env values overlays — in the exact layout erun build and erun deploy discover by convention (<tenant>-devops/docker/<component>/, <tenant>-devops/k8s/<component>/), so a hand-written or generated service becomes a component erun can build and ship without anyone reverse-engineering the convention. Also maintains, repairs, and upgrades deploy artifacts it previously produced in place, without clobbering the service's own source or hand-authored chart templates." |
| Triggers | "add deploy artifacts for this service", "scaffold a Dockerfile and helm chart for <component>", "make this service deployable with erun", "wire up build and deploy for <component>", "add a component chart", "this service has no Dockerfile/chart yet", "upgrade the <component> chart", "repair the <component> deploy wiring", "reconcile <component>'s deploy artifacts" |
| Inputs | Tenant; component name (validated against ^[a-z][a-z0-9-]*$, tenant-prefix recommended); the component's source location and language/toolchain; container port + health-check path (default 8080//healthz); the envs to generate values.<env>.yaml for (local always required); whether the component must be publicly reachable. |
| Outputs | <tenant>-devops/docker/<component>/Dockerfile (multi-stage: builder runs tests then builds, thin non-root runtime stage — Go skeleton shipped, swap the builder stage for another toolchain) and <tenant>-devops/docker/<component>/Dockerfile.dockerignore (this component’s own build context — BuildKit reads an ignore file named after its Dockerfile and sitting beside it, which is what lets each component under the shared repo-root context carry its own rather than share one root file). The Dockerfile’s layer order is part of the contract rather than house style: dependency manifests are copied and resolved before source, so a source-only edit reuses the dependency layer instead of re-resolving; and ARG TARGETOS/TARGETARCH are declared below the test step, so the per-architecture build invocations erun build makes share one cached, architecture-independent test layer instead of re-running the suite per arch — see the skill’s "Build speed" section for the full set and the measurements behind it. Also: <tenant>-devops/k8s/<component>/{Chart.yaml,values.local.yaml,values.<env>.yaml,templates/service.yaml} — a Deployment + Service named literally <component> (not tenant-templated the way erun's own published multi-tenant component charts are, since this chart is authored once for one tenant's own component), with the image defaulting to <containerRegistry>/<component>:<Chart.AppVersion> overridable via the same imageOverrides.<component> mechanism erun deploy already threads, and readiness/liveness probes on the configured health-check path. Before writing, checks whether */k8s/<component>/Chart.yaml already resolves anywhere else in the tree (componentHelmChartCandidate's matching is not scoped to the target <tenant>-devops/) and stops rather than creating a second chart erun would later refuse to disambiguate. Because erun push rewrites Chart.yaml version/appVersion to the resolved build version on every publish (overrideHelmChartVersion), the shipped placeholders need no hand-maintenance. Public reachability is not chart-side: pairing the tenant-prefixed component name with the literal Service name is what makes erun expose <tenant> <env> <service> (which targets the tenant-scoped Service <tenant>-<service>) resolve to the Service this chart already renders, with no separate Ingress in the chart itself. On an existing component (a Dockerfile or chart already at the conventional path) it enters maintenance mode instead of scaffolding: previews the plan, fills gaps against this skill's contract (a missing values.<env>.yaml, a Deployment/Service not literally named <component>, a missing probe, a runtime stage that isn't thin/non-root) without touching the Dockerfile's builder-stage toolchain commands or a hand-authored chart template, and re-validates with helm lint/helm template. There is no erun-version coupling to re-pin — this is the component's own release line, re-stamped by erun push as above. |
| Error behaviour | Component name fails ^[a-z][a-z0-9-]*$ → ask the user to rename (INVALID_COMPONENT_NAME). Name collides with a chart elsewhere in the tree → stop before writing; rename or reuse the existing chart. Name is the reserved <tenant>-devops → refuse (that's the runtime-pod chart's, owned by erun-build-env). Dockerfile/chart already exist → not a stop; enter maintenance mode and reconcile gaps in place, previewing first, never clobbering the service's own code or hand-authored templates. erun deploy fails values file not found for environment "<env>" → create the missing values.<env>.yaml (comment-only is valid); values.local.yaml is required too. erun deploy <component> fails multiple Helm charts found for component "<component>" → a second same-named chart was added by hand after scaffolding; rename one. helm lint/helm template unavailable locally → skip validation and say so; the runtime image ships helm. erun expose resolves but the Ingress 503s → the targeted Service <tenant>-<service> doesn't match what the chart rendered, usually a component scaffolded without the tenant prefix; rename to add the prefix, or pass the exact post-prefix role as <service>. A one-shot Job (migration/cron) requested instead of a long-running service → out of scope for the shipped templates/chart/templates/service.yaml shape; point at Conventions spec · Helm Job pattern for one-shots and hand-write the Job chart. |
| Compose | erun-blueprint-api produces a service's source only and has no deploy artifacts of its own; its Step 6 applies this skill to close that gap. Any other service-authoring skill or hand-written service can compose it the same way. |
erun-blueprint-docs
| Field | Value |
|---|---|
| Kind | Blueprint — packages ERun's docs-site pattern: a Docusaurus 3.x site published to Cloudflare Pages by a Kubernetes hook Job, the shape erun-docs captures. |
| Source | erun-skills/skills/erun-blueprint-docs/SKILL.md + templates/ |
| Description | "Scaffold a product documentation site following ERun's blueprint — a Docusaurus 3.x site published to Cloudflare Pages through a Kubernetes Job, the exact shape erun-docs captures — and also maintain, repair, and upgrade an already-scaffolded docs site in place, reconciling it with the current blueprint and re-pinning its versions without clobbering the project's own content pages." |
| Triggers | "set up product docs site", "scaffold a docusaurus docs site", "build erun-docs-shaped documentation", "create a docs site deployed to cloudflare pages", "add a documentation site for this project", "upgrade the docs site", "repair the docs deploy wiring", "reconcile the docusaurus site with the blueprint", "bump the docs site to <version>", "maintain the docs site" |
| Inputs | Module name (default <concern>-docs); target repo root; site title + tagline + production URL; Cloudflare Pages project name + branch alias; GitHub org/repo for editUrl |
| Outputs | <module>/ Docusaurus site (docusaurus.config.ts with onBrokenLinks: throw, sidebars.ts, docs/, src/css, static/img, package.json, tsconfig.json); erun-devops/docker/<module>/{Dockerfile,entrypoint.sh} (two-stage build → pinned wrangler); erun-devops/k8s/<module>/{Chart.yaml,values.local.yaml,values.prod.yaml,templates/docs.yaml} (ServiceAccount + post-install,post-upgrade hook Job that runs wrangler pages deploy). Both values.local.yaml (agent env, docs.enabled: false) and values.prod.yaml ship, because erun deploy requires a per-chart values.<env>.yaml for every env — including the <tenant>-local agent env the desktop deploys — with no fallback. On an existing site (a <module>/docusaurus.config.ts or the deploy plumbing present) it enters maintenance mode instead of scaffolding: previews the diff, reconciles the deploy wiring against the current erun-docs blueprint (a missing values.<env>.yaml — especially values.local.yaml — Chart.yaml/templates/docs.yaml/entrypoint.sh, onBrokenLinks: 'throw' turned off, a Git-connected Pages project, drifted plumbing), and re-pins two axes separately — the erun release (ERUN_VERSION, the Chart.yaml version/appVersion) to the target, and the docs toolchain (node/wrangler tags, @docusaurus/* pins) to current — before re-proving with yarn install/yarn build — never clobbering the operator's docs/ pages, sidebars.ts, or src/css/custom.css. Cleanup removes only superseded deploy-wiring, never the operator's content; a stale Cloudflare Pages project is the operator's to remove. |
| Error behaviour | Target dir already has <module>/docusaurus.config.ts → not a stop; enter maintenance mode and reconcile the deploy wiring against the blueprint + re-pin versions in place, preserving the existing docs/ content. yarn build fails on a broken link → fix the link, do not disable onBrokenLinks. npx create-docusaurus offline → fall back to bundled templates/. No Cloudflare alias (or its token lacks Pages:Edit) → scaffold still succeeds; the publish Secret and Direct-Upload Pages project are provisioned automatically from a Cloudflare alias, so surface that the first erun deploy Job fails until an alias whose token has Pages:Edit is attached (custom domain + DNS stay manual). User asks for a Git-connected Pages project → stop (Direct Upload only; a Git connection double-deploys). erun deploy fails values file not found for environment "<env>" → the chart is missing values.<env>.yaml; create it (an empty/comment-only file is valid), remembering the agent env needs values.local.yaml. |
erun-blueprint-platform
| Field | Value |
|---|---|
| Kind | Blueprint — packages ERun's accumulated best practices for hosted-platform deploy wiring. |
| Source | erun-skills/skills/erun-blueprint-platform/SKILL.md |
| Description | "Blueprint the deploy artifacts for a hosted erun platform — a per-env Terraform tree (terraform-<tenant>/) whose modules wrap erun's published Terraform modules, and the per-env Helm values overlays plus thin umbrella charts that reference erun's published OCI charts — all version-pinned to the erun release the environment runs; also maintains, repairs, and upgrades an existing terraform-<tenant>/ tree and its <tenant>-<component> umbrellas in place, re-pinning every erun reference to the target version and filling gaps against this contract." |
| Triggers | "blueprint the platform", "scaffold the platform terraform", "set up the platform helm charts and terraform", "create the terraform-<tenant> structure", "blueprint erun platform deploy", "set up platform deploy artifacts", "upgrade the platform terraform", "repair the platform charts", "reconcile the terraform-<tenant> tree", "bump the platform to <version>", "maintain the platform deploy artifacts" |
| Inputs | The env's tenant + short env name; the erun version to pin to (erun version in-pod, or the env's runtimeversion); the platform values (base_domain, services_zone, acme_email); the container registry (env containerregistry, default ghcr.io/sophium). |
| Outputs | terraform-<tenant>/{common.tf, variables.tf, .gitignore} (canonical providers + shared vars), terraform-<tenant>/modules/terraform-<tenant>-cluster-edge/ (wraps erun's terraform-erun-cluster-edge by ?ref=v<version>), and per env a terraform-<tenant>/<env>/ folder whose common.tf/variables.tf are symlinks to the root and that adds the env's services via its own main.tf + <env>.tfvars; plus — optional, the patch/override path (a runtime env deploys the published components by reference from config — see Deploy chart source — so no umbrella is needed for a normal deploy) — per platform component, a thin umbrella <tenant>-devops/k8s/<tenant>-<component>/Chart.yaml (directory name, chart name:, and Helm release all <tenant>-<component>, e.g. acme-docs) depending on erun's published erun-<component> OCI chart, with a per-chart values.<env>.yaml for every env it deploys to — including values.local.yaml. erun deploy <tenant> <env> reads <tenant>-<component>/values.<env>.yaml from each chart dir (required, no fallback, no config-dir overlay) and keys the component name off the directory, and the desktop deploys the <tenant>-local agent env, so a missing values.local.yaml fails the deploy. Each umbrella's resolved dependency is tracked as Chart.lock (committed) with charts/*.tgz gitignored (**/charts/*.tgz); erun deploy runs helm dependency build before install, rebuilding charts/ from Chart.lock, so the tgz is never committed (vendor it only for an air-gapped install). No run.tf, no per-env shell scripts — erun terraform apply owns the apply workflow. Where the tree produces a credential (an SES SMTP password, a generated DB user, a provisioned API key) rather than consuming one, the skill lays down the produce → Secrets Manager → Kubernetes Secret path instead: an aws_secretsmanager_secret + aws_secretsmanager_secret_version pair at the tenant/env-scoped path <tenant>/<env>/<name> (so a secret's owner is legible from its name and an IAM policy can scope to the prefix), outputs that carry the secret's name and ARN but never its value, and a sync that materialises it as a Kubernetes Secret in the tenant's namespace. IRSA is unavailable — erun's clusters are not EKS (a self-managed node, or OrbStack behind Tailscale), there is no OIDC provider to associate a ServiceAccount with, and the API server's discovery document is not publicly reachable, so the cluster cannot be registered as one either; the pod's own erun-host AWS profile is short-lived and is not a substitute for a controller that re-reads later. erun deploys no sync: no erun chart or published erun Terraform module installs External Secrets Operator or the Secrets Store CSI driver (terraform-erun-cluster-edge installs Traefik, cert-manager and the DNS-01 shim only, and erun deploy's <release>-cloudflare Secret is Cloudflare-specific and deploy-time), so the tenant module installs one by helm_release — with its CRs applied through a tiny local chart with depends_on, not kubernetes_manifest, whose plan-time CRD requirement fails on a first apply (the same reason erun's cluster-edge module ships chart-issuer/). Auth is a bootstrap credential: one long-lived IAM key held in a single Kubernetes Secret, read by a namespaced SecretStore (never a ClusterSecretStore), policy limited to secretsmanager:GetSecretValue + DescribeSecret on arn:…:secret:<tenant>/<env>/* — the trailing * absorbs the six-character suffix AWS appends to every secret ARN — and never ListSecrets, which is not resource-scopable. Rotation has three owners: Terraform (apply -replace= on the key, which destroys before it recreates), the sync (refreshInterval, the upper bound on propagation), and the workload (env-var Secrets are captured at container start and need kubectl rollout restart; a mounted volume updates in place only if the process re-reads it). The pattern solves distribution and rotation, not state exposure — a credential Terraform creates is in state regardless, and erun's default backend "local" {} keeps that state unencrypted on the env's home PVC — so the scaffold forces one of two recorded mitigations per credential: Terraform manages the aws_iam_user identity while the key is minted and put out of band (preferred wherever the credential allows it; the stance the blueprint already takes for the Zitadel masterkey, at the cost of manual rotation), or the state backend is encrypted and access-controlled and treated as secret material in its own right. The first mitigation is unavailable for a provider-derived credential — one the provider computes as a local transform on the managed resource itself rather than returning from an API — because no data source exists to read it back for a key minted out of band; aws_iam_access_key.ses_smtp_password_v4 (a SigV4 derivation the AWS provider computes on the resource, not an IAM API value — the plural aws_iam_access_keys data source returns key metadata only, never the secret or its SMTP derivation) is the worked example, and such a credential must fall back to the second mitigation regardless of preference. The choice is per credential, not per tree — a tree commonly applies the first mitigation to every credential that admits it and falls back only for the one(s), like the SMTP password, that structurally cannot — and a fallback forced this way must be justified with a comment at the resource naming why the preferred path was unavailable, not just recording which mitigation applies. This skill wraps only the erun platform's own component charts; it never emits a runtime erun-devops/<tenant>-devops umbrella from here — the runtime chart is erun-build-env's. A tenant that ships these component charts runs them on its own version line, so erun-build-env must also publish a <tenant>-devops chart at the tenant version (its Step 6, required for such a tenant, and what erun deploy demands when a deploy includes tenant components); a bootstrap/erun-only env with no components of its own may still ride the shared erun-devops chart via imageOverrides. On an existing tree (a terraform-<tenant>/ or any <tenant>-<component> umbrella present) it enters maintenance mode instead of stopping: previews the diff/plan, then reconciles in place — re-pins every erun reference on both sides (each tenant module's ?ref=v<version> and each umbrella Chart.yaml dependency version:) to one target, fills contract gaps (absent common.tf/variables.tf symlinks, a missing per-env values.<env>.yaml including values.local.yaml, a missing **/charts/*.tgz gitignore entry, an uncommitted Chart.lock), and refreshes derived artifacts (helm dependency update to regenerate the committed Chart.lock, then erun terraform apply) — never clobbering the project's own tfvars, values overrides, or extra tenant modules, and confirming the tenant first on a loose match. Cleanup removes only what the reconcile supersedes — a dropped umbrella dir or a stale/mis-named component release the new set replaces — preview-first, and never helm uninstalls a stateful release (postgres / a data PVC) or drops data as a side effect (it stops and flags instead). |
| Error behaviour | terraform-<tenant>/ already exists → not a stop; enter maintenance mode and reconcile the tree in place (re-pin every erun reference to the target version, fill contract gaps, refresh derived artifacts) after previewing the diff, offering to add a new <env>/ folder if that's the ask and confirming the tenant only on a loose match. erun version unresolvable → stop and ask (never default to main for production wiring). ?ref=v<version> doesn't resolve on terraform init → pin to a released vX.Y.Z. helm dependency build 404s → that version's chart isn't published; pin to a pushed version. erun deploy fails values file not found for environment "<env>" → the umbrella chart lacks values.<env>.yaml; create it (empty/comment-only is valid), including values.local.yaml for the agent env. erun deploy fails tenant is required (or environment is required) → the wrapped subchart reads those in its own scope; author them nested under the dependency name in the umbrella's values.<env>.yaml. A by-reference deploy re-scopes deploy's --sets under the subchart key and helm pulls that file to apply it, so this now surfaces only on a worktree deploy whose values.<env>.yaml omits the nested keys. A component that can't run in an env (erun-powerdns needs :53/hostNetwork + a private-image pull secret; erun-zitadel needs a public auth host, its cert, and an existing masterkey Secret) → omit it from that env's --components, don't force it. erun-zitadel render fails zitadel.masterkeySecretName is required → create the 32-character masterkey Secret out of band and name it in the umbrella's values.<env>.yaml; never generate one into the chart or the repo. erun-powerdns CrashLoops binding :53 → it bound 0.0.0.0, which collides with the node's systemd-resolved 127.0.0.53:53 stub; set erun-powerdns.powerdns.localAddress in the umbrella's values.<env>.yaml to the node's interface IP (empty binds the node IP by default on current erun; the override is honored on every version) rather than hand-patching the live Deployment. Operator asks to put the Cloudflare token in <env>.tfvars → refuse; it is injected as TF_VAR_cloudflare_api_token at apply time. Operator asks to put a Terraform-produced credential in <env>.tfvars, a values.<env>.yaml, or a chart value (or to emit it as an output, or to paste it into a console field) → refuse and name where it goes: Secrets Manager at <tenant>/<env>/<name>, synced into a Kubernetes Secret; the tree carries the secret's name, never its material. Operator asks for IRSA / an eks.amazonaws.com/role-arn ServiceAccount annotation so the sync can reach Secrets Manager → refuse; it cannot work on a non-EKS cluster with no reachable discovery document — use the namespaced SecretStore with a prefix-scoped bootstrap credential. An erun-devops/<tenant>-devops umbrella under <tenant>-devops/k8s/ → that is the runtime chart (owned by erun-build-env), legitimate and required once the tenant ships its own components; don't create or edit it from this skill, only remove a stray one this skill created by mistake. |
erun-build-env
| Field | Value |
|---|---|
| Kind | Workflow — extend the environment's runtime image through ERun's supported extension path. |
| Source | erun-skills/skills/erun-build-env/SKILL.md |
| Description | "Create a custom build environment by extending ERun's published runtime image with the project's own toolchain, then pointing the environment at the result, and maintain, repair, or upgrade an existing custom build environment in place by re-pinning it to the target runtime version and filling any gaps against this skill's contract." |
| Triggers | "init build environment", "init erun build environment", "create a custom build environment", "customize the runtime image", "upgrade the build environment", "upgrade the custom runtime image", "repair the build environment", "reconcile the <tenant>-devops module", "bump the runtime image to <version>", "maintain the build environment" |
| Surfaced by | erun build — and the build it runs for --release / --deploy, plus the MCP build tool — prints a one-line advisory recommending this skill whenever it runs in a project with no <tenant>-devops build module (the module that would hold the custom runtime image). The advisory fires regardless of whether the build itself succeeds. |
| Inputs | The tooling to add (packages, toolchains, CLIs); the target tenant + environment. The module and image names are fixed by convention: <tenant>-devops for both (see Outputs). |
| Outputs | A <tenant>-devops module (outer directory name must end in -devops — erun build discovers the runtime build module by that suffix) containing a starter Dockerfile at <tenant>-devops/docker/<tenant>-devops/Dockerfile (inner directory name becomes the image name) with FROM <registry>/erun-devops:<runtime-version> (version read from erun version in-pod, or the env's runtimeversion / erun list on a laptop); a VERSION file at the module root (<tenant>-devops/VERSION, e.g. 1.0.0) — erun build mints the image version from it; an erun build run that builds both architectures and pushes to the env's registry; the env's runtimeimage field set to <tenant>-devops via erun init --runtime-image <ref> or a direct config edit. On the next deploy/open the image rides into the published chart as imageOverrides.erun-devops (Advanced chart values). For files a sourceless runtime env must carry on disk (a platform Terraform tree, seed data, fixtures), the skill directs baking them into /opt/erun/release/ — never under /home/erun, which the runtime pod's home PVC shadows — laid out relative to the repo root; on a runtime env the entrypoint symlinks the git folder (~/git/<tenant>) at /opt/erun/release, so COPY … /opt/erun/release/<tenant>-devops/terraform-<tenant>/ surfaces at ~/git/<tenant>/<tenant>-devops/terraform-<tenant>/ where erun terraform resolves it. A <tenant>-devops/k8s/<tenant>-devops/ umbrella chart depending on the published erun-devops chart — required once the tenant ships its own component charts (so the runtime deploys on the tenant's own version line, which erun deploy demands when a deploy includes tenant components), optional otherwise for pod shape the image can't express, and published by erun push/erun release at the tenant version — supplying extraContainers/extraVolumes/extraVolumeMounts/extraEnv/extraRules nested under the erun-devops subchart key in per-env values.<env>.yaml; erun deploy installs it as the runtime chart, helm dependency builds it, and re-scopes every runtime value (incl. the image override) under the subchart key (pod shape extensions). Because erun push publishes the umbrella and its <tenant>-devops image together, deploy defaults imageOverrides.erun-devops to the umbrella's own image (default runtime image), so once the umbrella is published runtimeimage is optional — set it only to pin a different image; an image-only env riding the shared erun-devops chart still needs it. On an existing <tenant>-devops module it enters maintenance mode instead of re-scaffolding: previews the diff, then re-pins to one target runtime version (FROM ghcr.io/sophium/erun-devops:<version>, the module VERSION, and any Step 6 umbrella erun-devops dependency version:), fills contract gaps (a missing VERSION, wrong <tenant>-devops module/image naming, a FROM that isn't erun-devops, or under the umbrella a missing per-env values.<env>.yaml/Chart.lock/charts/*.tgz gitignore), then rebuilds and pushes both arches — never clobbering the project's own toolchain layers. Cleanup removes a renamed/relocated old module (preview-first) but never prunes pushed images — those are the operator's. For a project that can't nest the module at the conventional repo-root <tenant>-devops/docker/<tenant>-devops/ depth, the skill also points at the .erun/config.yaml paths: escape hatch — paths.docker to relocate discovery and paths.dockercontext: repo-root so a deeper Dockerfile's repo-relative COPYs still resolve. |
| Error behaviour | Runtime version unresolvable (no runtimeversion in the env config and not inside a pod) → asks the Operator before writing the Dockerfile. Existing <tenant>-devops module (Dockerfile + VERSION present) → not a stop; enter maintenance mode and reconcile in place — re-pin the FROM/VERSION (and any Step 6 umbrella dependency) to the target runtime version and fill contract gaps, previewing the diff first and keeping the project's own toolchain layers. Module directory not ending in -devops, or no VERSION file at the module root → erun build fails (dockerfile not found in current directory / version file not found for current module); the skill's Steps 2–3 produce the layout that avoids both. erun build fails (e.g. BINFMT_MISSING, registry push rejected) → surfaces the build error and does not touch the env config. Base image other than erun-devops requested → refuses; the entrypoint, the Agent tooling, and the in-pod erun live in that image. |
erun-browser-session-rest
| Field | Value |
|---|---|
| Kind | Workflow — authenticated REST against a host that blocks API tokens, via a reused browser session. |
| Source | erun-skills/skills/erun-browser-session-rest/SKILL.md + save-session.mjs + request.mjs |
| Description | "Make authenticated REST calls to a host whose org blocks API tokens and admin-gates OAuth, by reusing a saved browser login session (Playwright storageState)." |
| Triggers | "authenticated REST via a browser session", "call an API that blocks API tokens", "reuse my browser login for API calls", "hit the <host> API without a token" |
| Inputs | Host base URL (ERUN_REST_BASE_URL / --base); login URL (ERUN_REST_LOGIN_URL / --login); session-file path (ERUN_REST_SESSION / --session, default ./session.json); per call: HTTP method, path, optional JSON body. No host, credentials, or IdP are baked in. |
| Outputs | save-session.mjs opens a real browser for manual login (SSO + MFA) and writes a Playwright storageState session file (cookies only — never a password). request.mjs makes the authenticated call, prints the response body to stdout, and rolls the session forward (re-saves it) so refreshed cookies persist. |
| Error behaviour | Missing base/login URL → usage error, exit 2. Session file missing/unreadable → "run save-session.mjs first", exit 2. HTTP error status → response body to stdout, status to stderr, exit 1. Expired session (401 / login redirect) → re-run save-session.mjs. Requires Node 18+ and Playwright (npx playwright install chromium). |
| Security | The session file holds live session cookies — treat it as a secret, keep it out of git, keep it short-lived. The login is intentionally manual; the skill never stores a plaintext password. This is a fallback — prefer an API token or an approved OAuth app whenever the host allows one. |
erun-enable-hosting-edge
| Field | Value |
|---|---|
| Kind | Workflow — applies the public hosting edge to a cluster through ERun's published Terraform module. |
| Source | erun-skills/skills/erun-enable-hosting-edge/SKILL.md |
| Description | "Stand up the public hosting edge for an erun cluster — a Traefik ingress controller, cert-manager, and a namespaced DNS-01 Issuer that issues wildcard TLS for the services zone — by applying the terraform-erun-cluster-edge module, and maintain, repair, and upgrade that edge afterwards by re-pinning the module ?ref to the env's erun version and re-applying to reconcile drift." |
| Triggers | "enable the hosting edge", "enable public hosting", "set up TLS ingress for erun", "apply the cluster edge", "set up cert-manager and traefik", "issue wildcard TLS for the services zone", "upgrade the hosting edge", "repair the cluster edge", "reconcile cert-manager and traefik", "bump the cluster edge to <version>", "maintain the public hosting edge" |
| Inputs | The env's CLOUDFLARE_API_TOKEN (injected by a Cloudflare alias) passed as TF_VAR_cloudflare_api_token; the services zone (platform.serviceszone) and ACME email (platform.acmeemail); the erun version the module ?ref pins to (erun version, else main off-pod). The apply runs in-pod as the runtime ServiceAccount and creates cluster-scoped resources (namespaces, CRDs), so the env must be a platform account (platformaccount: true, erun init --platform-account) — otherwise the SA has namespaced admin only and the apply is denied. |
| Outputs | A Terraform root (in a temp dir) that references terraform-erun-cluster-edge from erun's GitHub by ?ref=v<version> and applies it: a Traefik ingress controller, cert-manager + CRDs, a namespaced DNS-01 Issuer (erun-cloudflare, in the env namespace), and a wildcard Certificate for *.<services-zone>. Idempotent — re-running reconciles. Maintenance is the same re-apply: when the edge already exists (kubectl get issuer -n <issuer-namespace> erun-cloudflare succeeds) it re-pins the module ?ref to the env's erun version and re-applies to reconcile drift, previewing with terraform plan first — no separate scaffold artifacts, no clobbering operator-owned cluster content. There are no local artifacts to clean; tearing the edge down is a deliberate erun terraform destroy, never a maintenance side effect. The platform's own wildcard Issuer's DNS-01 solver has three modes via dns01_provider: cloudflare (default), powerdns-rfc2136 (DNS UPDATE + TSIG direct to the self-hosted PowerDNS — single-tenant platform cluster only), and powerdns-broker: a per-tenant namespaced Issuer whose challenges route through a per-cluster cert-manager webhook shim to the DNS-01 broker, which authorizes each challenge against the env's own subzone. Installing that webhook shim is a separate switch, install_dns01_webhook — it defaults to matching dns01_provider == "powerdns-broker" (back-compat) but can be set independently, so a platform can keep its own wildcard on cloudflare/powerdns-rfc2136 while still installing the shim for per-tenant brokered Issuers (the ones erun expose provisions). Once installed, per env it mints a DNS-01 token (POST …/dns01-token) landed as the Issuer's token Secret; it needs no Cloudflare token or TSIG key. A separate opt-in switch, install_coredns_forward (default false, so an already-applied cluster's DNS behavior never changes on a module upgrade), declares a coredns-custom ConfigMap server block that forwards base_domain_name to coredns_forward_upstreams (default the public resolvers 1.1.1.1/1.0.0.1/8.8.8.8, overridable for an air-gapped or policy-constrained cluster) — k3s's bundled CoreDNS already mounts and imports that ConfigMap, so declaring the block needs no change to the CoreDNS Deployment. A further opt-in switch, install_local_path_helper_pod_resilience (also default false), keeps storage reclaimable on a node under DiskPressure: local-path is the only storage class these clusters ship and its provisioner does both provisioning and reclamation through a helper pod it pins onto the node holding the volume, where kubelet's eviction manager — which admits and evicts by whether a pod is critical, and treats a pod with no priorityClassName as ordinary priority 0 — turns it away, so a new PVC never binds and a deleted PVC frees no bytes on exactly the node that needs the space back. Enabling it merges priorityClassName: system-node-critical into the distribution's own helper pod template (image and every other field preserved) and states the disk-pressure toleration explicitly, because any declared toleration suppresses the provisioner's own default one; it also annotates the provisioner Deployment's pod template with a digest of the template written, since the provisioner reads that template once at startup and k3s's 30-second refresh is a no-op without CONFIG_MOUNT_PATH, so a ConfigMap write alone would apply cleanly and change nothing. The ConfigMap key and the annotation both live in the distribution's manifest, which k3s re-applies on every server start, so a k3s upgrade rewrites them and the module must be re-applied if reclamation on a pressured node stops working. This makes in-cluster resolution of the platform's own published names (which the HTTP-01 self-check depends on, at issuance and every unattended renewal) independent of whatever DNS the node happens to fall through to. The same module owns the edge's transport policy — the plaintext entrypoint redirected to https (301) and Strict-Transport-Security served on the secure one, declared once at the only layer that sees every public host, since Traefik answers :80 for every rule it routes and a relative Location issued behind the edge inherits whatever scheme the browser started on. http_redirect_enabled and hsts_enabled are independent switches, both defaulting on; hsts_max_age_seconds defaults to one day with hsts_include_subdomains and hsts_preload off. That policy is carried on the controller the module installs, so install_ingress_controller=false leaves it nowhere to land: that combination is refused while manage_transport_policy is left at its default true, and a bring-your-own-controller cluster instead sets manage_transport_policy=false and reads what the module would have carried from the edge_transport_policy output — the redirect and HSTS entrypoint arguments, the HSTS Middleware object, redirect_objects when the ACME exemption below is on, and the resolved http01_acme_challenges_present the policy was built on — the one fact the module cannot observe, echoed because a caller that never sees the module's own resources has no other way to read it back. Because that redirect is entrypoint-wide and has no path predicate, it also covers /.well-known/acme-challenge/: safe where every challenge is DNS-01 (the module's own Issuer always is), but a certificate the platform brings on an HTTP-01 Issuer is starved by it. A cluster in that position declares it with http01_acme_challenges_present=true and sets acme_challenge_path_exempt=true, which carries the redirect as a Middleware plus a catch-all router whose rule excludes the challenge prefix. Both default to false. |
| Error behaviour | No CLOUDFLARE_API_TOKEN → stops, points at erun cloud init cloudflare + erun cloud set … --alias <name>@cloudflare. terraform/kubectl missing, or kubectl not pointed at a reachable cluster → stops. apply fails with namespaces is forbidden / customresourcedefinitions … is forbidden for the runtime SA → the env is not a platform account; set platformaccount: true (erun init --platform-account) and redeploy from an admin-capable context so the chart binds the SA to cluster-admin, then re-apply. Issuer/Certificate stalls → kubectl describe the ACME order/challenge; usual causes are a token missing Zone:Read+DNS:Edit or the services zone not yet delegated to Cloudflare. A self-check failure specifically (dial tcp: lookup <host> ...: no such host in the challenge, while the name resolves fine from outside the cluster) points at the node's resolver rather than the zone — confirm with an in-cluster nslookup and apply install_coredns_forward=true with base_domain_name set. install_coredns_forward = true with base_domain_name unset → rejected at plan time by a Terraform precondition. A cluster already carrying a hand-applied coredns-custom ConfigMap needs it reconciled once (terraform import, or delete the hand-applied copy) before the first apply with install_coredns_forward = true, because the module creates that object when nothing else has. It owns only its own key inside it (server-side apply under its own field manager), so importing an existing ConfigMap keeps every other key through subsequent applies, and a destroy removes just the module's key. Set manage_coredns_custom_configmap = false on a cluster where something else owns that object's lifecycle and the module will manage its key without ever creating or deleting the ConfigMap. The forward also refuses to apply when CoreDNS's Corefile does not import /etc/coredns/custom/*.server, rather than writing an entry nothing will read, and validates each coredns_forward_upstreams entry at plan time — a malformed one would otherwise write an invalid server block that only fails at CoreDNS's next restart. While validating a fresh zone, -var acme_server=<staging> avoids Let's Encrypt production rate limits. install_ingress_controller=false with manage_transport_policy left at its default → rejected at plan time by an output precondition, because no controller would carry the redirect or the HSTS header and every public host would serve cleartext while the plan reported success. http01_acme_challenges_present=true with the redirect still entrypoint-wide → rejected at plan time by a second precondition: Traefik applies that redirect to /.well-known/acme-challenge/ too, so the solver is redirected to https where its Ingress has no TLS block and the certificate fails to renew — weeks later, as an expiry rather than as a plan error. Resolve it with acme_challenge_path_exempt=true, or http_redirect_enabled=false, or by moving those hosts to a DNS-01 Issuer and clearing the declaration, or — where the HTTP-01 challenges never reach this edge's plaintext entrypoint in the first place, because another controller fronts those hosts or the solver answers where this entrypoint does not see it — by clearing the declaration, which is the correct remedy there rather than a way around the refusal. |
erun-onboard-service
| Field | Value |
|---|---|
| Kind | Workflow — adopts an existing repository's own layout into erun, then builds, deploys, and serves one of its services over TLS. |
| Source | erun-skills/skills/erun-onboard-service/SKILL.md |
| Description | "Adopt a repository that already has its own layout into erun — discover where its Dockerfiles and charts actually live, wire .erun/config.yaml to that layout without moving a single file, preflight the environment for the failures that surface far from their cause, then build, deploy, and expose one of its services at HTTPS with a valid certificate (public hostname or localhost)." |
| Triggers | "start using erun in this repo", "onboard this service to erun", "this repo has its own structure", "expose this service with a valid cert", "make this service reachable over https", "serve it on localhost with a valid certificate", "wire this repo up to erun", "deploy a service from a custom repo layout" |
| Inputs | The repository (any layout); the tenant/environment to onboard into; which candidate service to bring up when the repo holds several; the hostname the service should be served at; the DNS-01 path to use (cloudflare with the tenant's own zone and token, or powerdns-broker with a platform-minted per-tenant token). |
| Outputs | An .erun/config.yaml committed to the repo whose paths.docker/paths.k8s/paths.dockercontext point at the repo's existing directories — no repo file is moved or renamed — plus a k8s.deployments entry for the chosen component; when the roll-out set names more than one independent docker/k8s root, a components: map (one entry per root, keyed by component name — see Configuration spec · components: block) instead of paths:, so --component <name> selects which root erun build/erun push resolve and every root stays committed rather than being swapped in and out of paths:; a .gitignore correction when the repo ignored that config (.erun/* + !.erun/config.yaml, since git cannot re-include a file whose parent directory is excluded) and a git status --porcelain check proving nothing the onboarding produced is left untracked; a preflight report covering cluster-registry SA rights, the dind insecure-registry flag, anonymous pullability of every image the plan references at its pinned version, whether cert-manager/an ingress controller already exist, and whether the committed chart values are a production config that cannot boot without secrets; a built and deployed component (with a non-prod values override when the preflight found one is needed); and the service served over HTTPS with a certificate verified by fetching it. |
| Error behaviour | Several candidate services discovered → report all, onboard one, never pick silently (more than one may still be wired via components: in the same change if the operator asks for it; onboarding still builds and deploys one at a time). The project config is gitignored (git check-ignore -v .erun/config.yaml names the rule) → fix the ignore and commit it, rather than reporting the situation and leaving a per-checkout hack behind a green "onboarded"; an ignored config means the next environment's build resolves the fallback registry it cannot push to, and that surfaces as an authorization error at push time, nowhere near its cause. An untracked values.<env>.yaml is the same defect in a second file — the next person deploys the chart's production defaults. COPY fails during build → paths.dockercontext (or the entry's dockercontext under components:) must be repo-root. kubectl get svc kube-system/erun-registry: Forbidden → the runtime SA lacks the kube-system Role (get/list services, get/list pods, create pods/portforward); add it rather than switching registries around it. An image manifest returning 403/404 at the pinned version → stop; mirror it into the in-cluster registry or make the package pullable, because an unpullable image cannot be deployed around and its symptom (no certificate) appears three layers from its cause. First deploy CrashLoopBackOff → the committed production values; apply the non-prod override (an empty-stub secret is worse than none — an empty DATABASE_URL breaks a driver at connect time). A builder-stage test gate failing → the repo team's fix; report it rather than bypassing the gate for a green deploy. erun expose resolving but the Ingress 503ing → the chart renders a repo-native Service name that is not <tenant>-<service>; add an Ingress to the repo's own chart instead of renaming a Service the repo owns. Certificate never Ready → walk Certificate → CertificateRequest → Order → Challenge; the usual causes are an unpullable webhook shim, a token that cannot write the subzone, or a hostname outside the issued zone. ssl_verify_result non-zero → the chain is a staging-ACME or local-CA one; say so rather than reporting "HTTPS works". |
| Compose | The mirror image of erun-blueprint-service: that one writes a service's missing deploy artifacts into erun's conventional <tenant>-devops/{docker,k8s}/<component>/ layout, this one leaves the repo's existing artifacts exactly where they are and teaches erun to find them. Delegates the edge itself to erun-enable-hosting-edge when the cluster has no cert-manager or ingress controller yet. |
erun-pin-version
| Field | Value |
|---|---|
| Kind | Workflow — the agent-facing wrapper around erun pin for moving the erun version an environment records, and the operation the desktop's Change erun version dialog and the MCP pin tool drive. |
| Source | erun-skills/skills/erun-pin-version/SKILL.md |
| Description | "Change or try the erun version an environment uses, by re-pinning every place that version is recorded — the Terraform module refs, an erun image reference set directly in Terraform variables (e.g. the cluster-edge module's dns01_webhook_image), each umbrella chart's erun dependencies, the build-env image tag, a stated runtime chart or runtime image naming erun's own stock release, and the environment's own runtime version — in one verified, idempotent motion, then reverting just as easily if it doesn't work out." |
| Triggers | "change the erun version", "try a newer erun", "pin erun to <version>", "upgrade this environment to erun <version>", "what erun versions can I pin to", "the terraform ref and the charts disagree", "realign the erun pins", "revert the erun version", "roll back the pin" |
| Inputs | The tenant + environment; a target version, or --list / --revert instead of one. The project root is resolved by the command — the environment's own checkout, then a sibling environment of the same tenant that has one, and only as a last resort the caller's working directory, which the plan flags as unverified. |
| Outputs | A resolved plan naming every pin site (erun-common's PinSiteKind) with its current and target value, printed before anything is written; the rewritten references; the version being left, recorded so a later --revert reaches it; and the regenerated Chart.lock of every chart that moved. Only erun's own references move — a tenant's own Terraform sources, its own chart dependencies, and the umbrella chart's own version: are versioned independently and are never touched. Nothing is deployed: realizing the version stays a separate erun terraform apply and erun deploy. The same operation is available as the desktop's Change erun version dialog (erun pin) and the MCP pin tool, whose preview returns the plan without writing. |
| Error behaviour | Target not published → refused before writing, naming the registry it checked; report that and offer erun pin --list, never work around the refusal, and never pin to main — a tree pinned to something unpublished fails much later, at a terraform init or a chart pull. Registry unreadable with an explicit --version → pins anyway and traces that it could not check, because "could not verify" is not "not published"; unreadable with no version → errors, since resolving the latest genuinely needs the registry — do not substitute a guess. --revert with nothing recorded → errors; ask for an explicit version instead. Already aligned → reports no changes and writes nothing. A host env, or an environment with no local checkout anywhere → refused before writing, naming the tenant and environment. helm dependency update fails on a chart → the command names the chart and stops rather than leaving some locks refreshed and others stale; fix that chart's access to the registry and re-run, the re-pin being idempotent. A site the operator expects but the plan does not show → report the gap (erun-file-issue); never hand-edit the pin site, which is the drift this exists to end. |
erun-setup-k3s-cluster
| Field | Value |
|---|---|
| Kind | Workflow — provisions a durable local cluster on Windows for erun to build and deploy to. |
| Source | erun-skills/skills/erun-setup-k3s-cluster/SKILL.md |
| Description | "Stand up a durable local Kubernetes cluster on Windows that erun builds and deploys to — real k3s running inside WSL2 (with an in-cluster image registry and a WSL-hosted Docker engine, no Docker Desktop), wired to an erun local-agent environment — and maintain, repair, or tear it down afterwards." |
| Triggers | "set up a local erun cluster on Windows", "set up k3s for erun", "create a local k3s cluster", "install k3s on Windows for erun", "run erun locally on Windows", "wire erun to a local cluster", "give me a local cluster to deploy erun to", "repair the local k3s cluster", "tear down the local erun cluster" |
| Inputs | Windows 11 22H2+ (WSL mirrored networking + hostAddressLoopback); kubectl, helm, and the docker client on Windows (scoop install kubectl helm docker); the erun CLI on PATH; the target tenant + environment name (env named local by convention). No Docker Desktop — the Docker daemon runs inside the distro. The registry (localhost:5000), kube-context (erun-k3s), and Docker endpoint (tcp://127.0.0.1:2375) are fixed by the skill. |
| Outputs | A WSL2 Ubuntu distro (installed via wsl --install, or the WSL MSI from GitHub releases when the inbox stub misbehaves) with systemd enabled and %USERPROFILE%\.wslconfig set to networkingMode=mirrored and [experimental] hostAddressLoopback=true — both required so localhost resolves host↔WSL in each direction (mirrored alone leaves the Hyper-V firewall blocking host→WSL). Inside the distro: iptables installed and a dev-kmsg.service that provides /dev/kmsg (both needed by k3s on WSL2); a k3s server (--write-kubeconfig-mode=644) with /etc/rancher/k3s/registries.yaml mirroring localhost:5000 over HTTP; an in-cluster registry (registry:2 as a kube-system Deployment on hostNetwork with REGISTRY_HTTP_ADDR=0.0.0.0:5000 and a hostPath volume, dropped into /var/lib/rancher/k3s/server/manifests/ so k3s reconciles it — hostNetwork binds the node's :5000 directly, avoiding the CNI portmap/hostPort path that is unreliable on WSL2); a docker.io daemon exposed on unix:// + tcp://127.0.0.1:2375 via a systemd drop-in, with binfmt registered (tonistiigi/binfmt) for erun's mandatory multi-arch (linux/amd64+linux/arm64) builds. On Windows: DOCKER_HOST=tcp://127.0.0.1:2375 (User env); an erun-k3s context written to %USERPROFILE%\.kube\config from /etc/rancher/k3s/k3s.yaml (server https://127.0.0.1:6443) via a rename regex that handles both name: default and k3s's - name: default; an erun local-agent environment created with erun init <tenant> local --type local-agent --kubernetes-context erun-k3s --container-registry localhost:5000 run inside the project (writes containerregistries: [{registry: localhost:5000, roles: [build, deploy]}] to the project .erun/config.yaml and kubernetescontext: erun-k3s to %LOCALAPPDATA%\erun\<tenant>\local\config.yaml); and an erun-k3s-boot logon Scheduled Task that boots the distro so systemd restarts k3s + registry + dockerd across reboots. Validated end-to-end on Windows 11 24H2 (k3s v1.36, Ubuntu 26.04): kubectl get nodes Ready from Windows, docker push localhost:5000/... from Windows then a k3s pod pulling that image to completion, and erun deploy <tenant> local --version <v> --dry-run resolving kubectl --context erun-k3s, --kube-context erun-k3s, and containerRegistry=localhost:5000 / oci://localhost:5000/charts/erun-devops. |
| Error behaviour | Windows older than 11 22H2 (no mirrored networking/hostAddressLoopback) → stop; the host↔WSL localhost wiring cannot work. dism/wsl --install returns Error 740 → the shell has Administrators-group membership but only Medium integrity; run the elevated step from a Scheduled Task as NT AUTHORITY\SYSTEM -RunLevel Highest (use msiexec for the WSL MSI; the inbox wsl --install stub needs an interactive session and fails as SYSTEM). Inbox wsl.exe only reprints "…is not installed" → install the WSL MSI directly. .wslconfig missing hostAddressLoopback=true → kubectl gets error: EOF / registry unreachable from Windows though the distro is fine internally; add it and wsl --shutdown. kubectl → Please enter Username → the kubeconfig rename missed k3s's - name: default, so the context's user reference has no matching user; re-run the rename regex. k3s pull http: server gave HTTP response to HTTPS client → the registries.yaml mirror is missing; re-write it and systemctl restart k3s. Registry :5000 not listening on the host / iptables absent → the CNI hostPort path is unreliable on WSL2; keep the registry on hostNetwork (don't switch back to hostPort). Invoke-WebRequest http://localhost:5000/... times out → a .NET HttpClient quirk over the WSL loopback; use curl.exe --noproxy '*' (docker push/pull over the same address work). erun deploy pushes to ghcr.io/... instead of localhost:5000 → the project's .erun/config.yaml local env pins another registry; add the containerregistries override. First command after a cold start fails but later ones succeed → the distro was booting and k3s wasn't Ready; wait for kubectl --context erun-k3s get nodes Ready. Teardown (k3s-uninstall.sh + unregister the task, optionally wsl --unregister) is destructive and operator-initiated, never a maintenance side effect. |
erun-orchestrate
| Field | Value |
|---|---|
| Kind | Workflow — host-side coordination of authorized work across linked environments. |
| Source | erun-skills/skills/erun-orchestrate/SKILL.md |
| Description | Operate as a host-side erun orchestrator that drives and reviews work across the environments it links — delegating pod work to the in-pod Agent, and working a host environment's own directory directly. Use when asked to "orchestrate erun environments", "drive the remote agents", "coordinate work across environments", "review what the agents changed", "review changes across envs", "run the built app to verify", or "delegate this to the environment's agent". |
| Triggers | "orchestrate erun environments", "drive the remote agents", "coordinate work across environments", "review what the agents changed", "review changes across envs", "run the built app to verify", "delegate this to the environment's agent" |
| Inputs | ERUN_ORCHESTRATOR_ID, its matching configured links, the directories its own definition names, each environment's type and declared role, host/pod paths, and the authorized task. Config owns scope; directory contents and names do not. Code/build/runtime roles select implementation, regression/release, or operation respectively. Runtime links have no worktree/Agent; missing roles are reported rather than guessed. |
| Outputs | In-pod development over erun MCP, and read-only host review of a pod-backed environment's worktree; a host environment has no pod, so it is worked in its own directory directly — as are the orchestrator's own directories, which belong to no environment at all and are named by path in its definition. Host-native verification of cross-built outputs. Remote mirrors deliver artifacts under .erun-outputs/; local-agent artifacts need download. Native GUI compilation is the host-build exception, using a separate owned copy outside all review directories — a host environment is already exactly such a copy. Restarts use erun app restart and an ID-specific RESUME-NOTE.<id>.md naming task, delivered work, in-flight job IDs, and first checks. Final reports include evidence, gaps, and assumptions. |
| Ownership | Issue claims use wip:<orchestrator-id>: confirm open, add, then re-read state/labels immediately before dispatch. On a concurrent double claim, the lower-sorting ID yields. Release at completion/handback; a later return is a fresh claim. Worktree mutations require exclusive ownership; heavy gates require an exclusive environment scope through terminal verdict, even with separate clones. |
| Supervision | Long work uses tracked jobs, lifetime leases, bounded awaits, and actual terminal records. Recheck job/channel/HEAD/resources regularly and promptly observe fast lanes. Automatic reinvocation is bounded recovery, not guaranteed completion. An exit-zero wrapper with orphaned work is not a verified final result. |
| Authority | Carry the requested outcome through authorized steps, not an automatic release/deploy expansion. Respect shorter requested flows and stop for genuine missing authority. General contribution, engineering, and direct in-pod interaction guidance stays in repository AGENTS.md; an Operator need not use an orchestrator to work with an in-pod Agent. |
| Error behaviour | Channel recovery uses open --reconnect, not bare open that wakes a deliberately stopped environment. CLI channel exit 126 and wait timeout 124/Make 2 are not job verdicts; inspect the actual job. A stopped environment requires a separately authorized host lifecycle action. Lease refusal reports its holder without overriding it. Missing scope/role/access or unverified host-native behavior is reported explicitly. After restart, re-discover channels and confirm the running process/version; do not duplicate existing jobs. |
See activity leases and exclusive claims for lifecycle semantics, desktop orchestrators for configuration, and restart recovery for the restart operation. Release orchestration verifies publication before public tag push; release does not itself deploy.
erun-merge
| Field | Value |
|---|---|
| Kind | Workflow — takes a finished change from "done" to a review sitting at READY on the erun platform, the same rungs the desktop's diff-panel Merge action drives. |
| Source | erun-skills/skills/erun-merge/SKILL.md |
| Description | "Take the current branch from "the work is done" to a review sitting at READY on the erun platform — resolve or accept a target branch, merge it in, commit and push, open or reuse the review, build and record the result. Stops at READY/FAILED and never advances the merge queue. Needs a machine with a configured erun platform cloud alias; an agent environment can hold one only if erun init provisioned it from a signed-in host, so where it has none it stops after the push and hands the review rungs to a credentialed host." |
| Triggers | "merge this branch", "land this change", "merge onto main", "advance the merge queue for this branch", "run erun-merge" |
| Inputs | <targetBranch>, optional — given, it is the target; omitted, resolved from erun exec diff --json --scope all's reviewBase.branch (the same fork-point detection the desktop diff panel uses), stated before acting. A commit message from the operator only if the working tree has uncommitted changes to fold in. A review name only when opening a fresh review and the branch's latest commit subject doesn't already describe it. |
| Outputs | The target branch merged into the current branch with an explicit merge commit (erun exec merge, never a rebase); the branch committed if dirty, checked for a declared regression reproduction, then pushed (erun exec commit / scripts/check-regression-coverage.mjs / erun exec push); a review opened or reused for the source→target pair (erun review list then erun review create); the pushed commit built (erun build, a plain non-release build — narrowed with --platform to the architecture the machine actually runs when the checked-out branch's .erun/config.yaml pin does not already do so) and the result recorded against the review (erun review record-build), which is what actually moves it to READY or FAILED — there is no separate status-setting call. Reports the review id, the build outcome, and the resulting status, then stops. |
| Error behaviour | erun not on PATH → stops before touching git, naming erun cloud init erun / erun cloud login as laptop setup. No usable erun platform alias → stops before touching git, on a probe of erun review list --dry-run (a local alias resolution that never reaches the network) whose exit 127 is the CLI's own "cannot resolve a usable platform alias" code; names the build/platform split — this environment builds, a credentialed host makes every erun review call — and exits 127 without attempting erun cloud login, whose device and PKCE flows both need a human at a browser. Any other nonzero probe result is left to the real call that reports it. Never guesses a URL or alias. A conflicted merge → stops, names every conflicted file, leaves the worktree mid-merge for the operator to resolve or git merge --abort; never resolves a conflict itself. Pushing to a branch whose review already reads MERGED/CLOSED → stops before the push (a merge queue squash-lands under a new SHA, so a push to the old branch would succeed and land nowhere — erun#2007); names the branch and directs starting a fresh one from the target instead. A bug/ branch that names no regression reproduction → stops before the push, printing the Reproduces:/Regression-Test: trailer block to amend onto a commit in the range (or the closed set of Regression-Test-Exemption: kinds when the fix genuinely cannot carry one); the check is skipped only outside a checkout of sophium/erun, where the script does not exist, and renaming the branch to dodge it is not a supported path. A failed erun build → records the failure against the review (--failed, moving it to FAILED) rather than retrying on its own, taking the version erun build --dry-run --output json mints since a failed run printed no result; re-running the skill re-runs the build. |
Never advances the merge queue (erun review queue advance) or its override (erun review queue override-advance) — both are a separate operator decision once a READY review's threads are resolved; see Review loop topology's READY, all resolved row and Merge queue for the gate mechanics this skill deliberately does not restate. Every rung checks the state it would produce before acting, so re-running after a partial failure resumes rather than repeating a merge, a push, or opening a second review for the same pair.
erun-review
| Field | Value |
|---|---|
| Kind | Workflow — the reviewer-side counterpart to erun-merge: reads a review's diff, leaves line-anchored comments for what should block the merge, and pushes a proposal branch where it has a concrete fix, the same rungs the erun-reviewer reusable agent drives. |
| Source | erun-skills/skills/erun-review/SKILL.md |
| Description | "Review someone else's branch on the erun platform — read the diff, leave line-anchored comments only for what should actually block the merge, and where there is a concrete fix, push it as a proposal branch the author can take. Stops at "reported"; never advances the merge queue, overrides it, closes the review, or resolves a thread it did not open." |
| Triggers | "review this branch", "review the change", "leave review comments", "review PR", "run erun-review" |
| Inputs | <reviewId>, optional — given, it is the review; omitted, resolved from erun review list --source-branch <currentBranch> when exactly one open review matches. |
| Outputs | Existing comment threads read first (erun review show) so no point already made is re-raised; own OPEN threads with a reply that genuinely addresses the point resolved (erun review resolve); the source fetched and diffed against the target (git fetch + erun exec raw git diff) without checking out the source branch; line-anchored comments posted (erun review comment, body on stdin, paced against the write-endpoint budget) only for findings that should block the merge, with anything advisory folded into one summary comment or dropped; where a finding has a concrete fix, a proposal/<reviewId>/<slug> branch pushed from the review's head (never the source branch) and named in a comment with the exact command to take it; the worktree restored to the branch it started on, or the left-checked-out state named plainly if it can't be. Reports the review id, every thread opened or resolved, every proposal branch pushed, and what wasn't reviewed and why, then stops. |
| Error behaviour | erun not on PATH → stops before any platform call, naming laptop setup. No erun-type cloud alias configured → surfaces erun review show's own error naming erun cloud init erun --api-url <url>; never guesses one. Zero or more than one open review matching the current branch when <reviewId> is omitted → stops, asks for it explicitly. A dirty worktree before pushing a proposal → stops and reports the uncommitted state rather than stashing, discarding, or checking out over it. A write call failing mid-batch → stops and reports exactly which comments posted and which didn't; comments are immutable, so there is no retry-in-place. |
The single policy this skill exists to enforce: a thread is for something that should stop the merge. Every open thread blocks erun review queue advance's 409 gate until its own root author resolves it or an operator burns an audited override-advance — see Merge queue § The unresolved-thread check. Volume is not thoroughness here; it is a denial of service on the author, so anything advisory goes in one summary comment or isn't raised at all. This skill never advances the queue, never overrides it, never closes the review, and never resolves a thread it did not open — see Review loop topology § The reviewer must come back for why only the thread's own author can unblock it.
erun-merge-queue-drive
| Field | Value |
|---|---|
| Kind | Workflow — execute an explicitly invoked gate for already-promoted reviews; no automatic drainer. |
| Source | erun-skills/skills/erun-merge-queue-drive/SKILL.md |
| Description | Drive one or more reviews already promoted to MERGE through the merge-queue gate — batch their sources into one prospective merge with erun exec gate-merge (skipping, per branch, any that conflict), gate the landed stack with one real erun build, and push and report MERGED only for branches that actually landed and passed. Reports each actual outcome, including reviews left at MERGE after an inconclusive gate, and never advances, overrides, or promotes the queue itself. Requires a machine with a configured erun platform cloud alias, since every rung is a platform call; an agent environment can hold one only if erun init provisioned it from a signed-in host, so where it has none it stops before claiming the environment and hands the drive to a credentialed host. Use when the user says "drive the merge queue", "batch these reviews through the gate", "run the merge gate", "gate this promoted review", "build and push the merge queue head", or any similar request to execute the gate for one or more reviews that are already at MERGE. |
| Triggers | "drive the merge queue", "batch these reviews through the gate", "run the merge gate", "gate this promoted review", "build and push the merge queue head" |
| Inputs | One or more supplied review IDs already at MERGE, a common target, verified remote source SHAs, platform/git/build access, and the owning environment. Platform access is a hard precondition of the machine running the drive, not something the drive can arrange: an agent environment cannot sign itself in to an erun platform alias (only erun init on a signed-in host can provision one), so where it has none the drive stops before claiming the environment instead of reserving it for work it could never record. Non-MERGE reviews are dropped; no survivors, mixed targets, or missing refs stop the drive. The normal API permits only one MERGE review per tenant/target: multi-source git composition does not establish general multi-review acceptance. |
| Outputs | An explicit environment-wide lease, ordered gate-merge composition with landed/skipped evidence, one real non-release build and one gate-run per batch attempt, eligible per-review GATE evidence, then a single target push and verified report-merged calls. Confirmed accepted sources use close-pr with both gated source and landing SHAs. Final output names records, commits, actual review states, logs, closures, and anomalies. The claim is renewed and released on every exit. |
| Error behaviour | Missing platform access stops before the environment claim: a probe of erun review list --dry-run (a local alias resolution that never reaches the network) exiting 127 — the CLI's own "cannot resolve a usable platform alias" code — ends the drive with the build/platform split named, and never attempts erun cloud login, whose device and PKCE flows both need a human at a browser; stopping before the claim is deliberate, so a drive that could record nothing does not reserve the environment and refuse the gate job a credentialed host could run. Ownership refusal stops before mutation. Nothing landed means no build/push/pass: classify the captured error rather than treating every empty stdout as a code failure. Conflicts remain for the branch owner. Capture stderr and report the gate-run first; known infrastructure failures or genuinely unknown outcomes are INCONCLUSIVE and leave reviews at MERGE. Real gate failures may record failed GATE builds against real commits. Wrapper timeout is not a verdict. Successful-build recording refusal is inconclusive, never a fabricated failure. Push/acceptance/PR-head-moved refusals stop their dependent steps without erasing earlier successful evidence. |
| Acceptance | MERGED requires a successful GATE build for the exact review, target reachability, and ancestry of the recorded gated target tip — not immediate-parent equality. Stop at a refusal; never report acceptance before push or close an unaccepted/moved-head PR. General batch-to-review verification remains an unresolved API design, not a shortcut authorized by this skill. |
| Resumption | Re-read reviews, refs, saved composition, and records. Before successful evidence, an authorized new attempt gets a new gate-run. After a partial successful build/push/report, resume only missing verified steps using original IDs; do not blindly rebuild, duplicate evidence, or repeat promotion. |
| Limits | Never advances/overrides/promotes the queue, resolves skipped conflicts, automatically retries failed pushes, or triggers a release. Ruleset migration and release-cadence guidance are separate, explicitly authorized operational handoffs, not extra execution rungs. |
See merge queue for acceptance and bypass reconciliation for ruleset evidence. The desktop gate runs real headless tests, not a manual attestation; native-OS gaps remain distinct.
Infrastructure-failure classification
- Gate-run start/report examines only caller-reported
FAILEDstatus (trimmed, case-insensitive). Other requested statuses are unchanged. - It combines
failingStep, thelogRefvalue itself, and the contents oflogRefwhen it names a readable local regular file no larger than 1 MiB. Missing/unreadable files, larger files, job IDs, and URLs contribute no file content; the classifier does not fetch remote logs. - It matches this text case-insensitively against these signatures, any one of
which is sufficient:
failed to resolve source metadata,failed to fetch oauth token, andtls handshake timeout. - A match changes the gate-run report to
INCONCLUSIVEand emits a trace in both preview and real execution. Unmatched failures remainFAILED. - A failed review
GATEbuild applies the same signature matcher to itsfailureDetail, but refuses a match rather than writing a false failed build: that record has only a boolean success field. Report the gate-run inconclusive and preserve the review for a separately authorized recovery attempt.
This classifier recognizes only the listed signatures. An unknown outcome or another established infrastructure fault still needs an explicit inconclusive report; a wrapper timeout alone is not evidence that the underlying job ended.
Catalogue evolution
The catalogue is open — new skills land in erun-skills/skills/ and ship through both distribution paths automatically. Each skill's description is what the Agent matches on, so additions don't require coordinated client changes.
Adding a custom skill
- Create
<projectRoot>/.erun/skills/<skill-name>/SKILL.mdwith the frontmatter above plus your guidance body. - Commit it. Anyone who opens an env on this project picks it up on the next
erun open. erun doctorvalidates skill bundles on startup — frontmatter parse failures, name mismatches, and missingSKILL.mdshow up asskill.<name>check failures.
Inspecting deployed skills
The MCP doctor tool reports the resolved skill set per env:
{
"checks": [
{
"name": "skills",
"status": "ok",
"detail": "10 built-in + 2 project skills loaded",
"skills": [
{ "name": "go-service", "source": "builtin" },
{ "name": "house-style", "source": "project" }
]
}
]
}
A raw listing also works: ls -la /etc/erun/skills/ ~/.claude/skills/ ~/.codex/skills/.
See also
- Skills — Operator-facing summary.
- Marketplace distribution — the plugin manifest and update flow.
- Conventions — what the skills teach.
- Conventions spec — the underlying layout the skills target.