CLI Reference
All @sensigo/realm-cli commands. Run realm <group> <command> --help for full option details.
Project extensions across commands
Section titled “Project extensions across commands”Workflows may declare custom adapters/handlers/processors via a top-level extensions: key —
see the Project extensions guide. Command behavior:
| Command | Behavior |
|---|---|
run, agent (fresh) |
Extensions load before the run is created — a broken module means no run. |
agent --run-id |
Extensions load before the run is claimed. Pre-execution failure marks the run terminal_reason: 'extensions_load_failed' (see recovery below). |
listen |
Loads every routed workflow’s extensions at startup (fail-fast); registers each workflow once at startup. Children re-resolve at spawn. |
serve, mcp |
Per-definition registries via a process-lifetime cache — restart the process to pick up module content changes. |
workflow register, watch |
Full module load + validation + config_schema two-pass before persisting (each watch reload re-validates). |
workflow test |
Extension handlers run real; unmocked extension adapters fail the fixture (tripwire). |
workflow validate |
Declaring workflows get two-pass config_schema validation against the resolved registry. |
run replay/inspect/list/diff |
Never load extension code. |
--project <dir> (on agent, run, serve, mcp) — the CONFIG anchor: the
deployment root whose realm.yaml applies to definitions without a
stored trust_root (agent-created, from-string). Defaults to the current directory for the
operator-launched agent/run/serve; realm mcp has NO default — its stdio cwd is
client-controlled, so the manifest loads only when --project is explicitly configured
(recorded security decision).
Sentinel secret mode: validate and test resolve ${secret:NAME} references to
labeled sentinels (no real credentials; constructor failures become warnings);
register/watch resolve real secrets when available and degrade with a loud WARN when
not. Execution paths always require real resolution.
--extensions-module <path> (on agent, run, serve, mcp, validate, test)
REPLACES the workflow’s declared modules for that invocation — a loudly-logged repair/override
tool, not the primary mechanism.
Recovering extensions_load_failed: fix the module (or pass --extensions-module), then
re-run realm agent --run-id <id> — attaching to a run whose terminal reason is exactly
extensions_load_failed clears that marker and retries (nothing executed before the failure).
All other terminal reasons keep the normal refusal.
Drift evidence: runs of extension-declaring workflows carry an append-on-change
extension_identity history (see the
Project extensions guide). realm agent --run-id
WARNs on stderr when the freshly loaded extension code differs from the run’s last recorded
identity — advisory only, never a gate; the new identity is appended at the next executed
step. realm run inspect <run-id> renders the history; realm run inspect <run-id> --check-drift recomputes the last entry against current disk state with pure hashing
(read-only commands never load project code).
Restart semantics for long-lived processes: listen, serve, and mcp cache extension
registries for the process lifetime and never re-import changed module content — restart the
process after rebuilding extension modules. Also restart realm listen after re-registering a
workflow: it serves the definitions snapshotted at startup (and if two listen instances mount
the same workflow from different directories, the last registration wins for source_dir).
Workflow commands
Section titled “Workflow commands”Operations on workflow definitions and YAML files.
realm workflow init <name>
Section titled “realm workflow init <name>”Scaffolds a new workflow project directory.
realm workflow init my-workflowCreates my-workflow/ containing workflow.yaml, schema.json, .env.example, and README.md.
realm workflow list
Section titled “realm workflow list”Lists the workflow definitions in the local registry — what realm workflow register has put
there, which is otherwise a black box.
realm workflow listrealm workflow list --jsonColumns: ID NAME VERSION ORIGIN SCHEMA. SCHEMA reads 1 (current) for a definition
registered under the current schema, and legacy (re-register) for an older one — those can no
longer run at all: every runtime call (starting a run, executing a step, responding to a gate)
refuses them with the same re-register message. Re-register from source to bring them back. Rows
are sorted by id.
Registry entries realm could not read are reported rather than skipped silently, one stderr
⚠ line per class — a file the OS refused (could not be read (<errno>)), a directory sitting
where a file belongs, an empty file, a file that is not parseable JSON — grouped by class and
errno and naming the files; an unreadable registry directory is one line and no table. The stdout
count says what it could not count (0 workflows registered; 1 entry in the registry could not be read.), so a stdout-only reader is never told the registry is clean when it is not. Files whose
stored id differs from their filename are reported too — --registered resolves by filename.
--json carries workflows, unreadable[] (file, class, errno where the OS gave one,
reason, and repair — the act for the class) and mismatched[]. Exit code is always 0: this is
a read surface, and reporting the broken entry IS the signal; validate --registered <id> is the
per-workflow gate.
realm workflow validate [path]
Section titled “realm workflow validate [path]”Validates a workflow YAML without registering it. Reports schema errors, duplicate step IDs,
and invalid depends_on references.
realm workflow validate ./my-workflowrealm workflow validate ./my-workflow --strict # fail (exit 1) if any warning is presentrealm workflow validate ./my-workflow --explain # full per-step structured_output detailrealm workflow validate --registered my-workflow # audit the stored copyrealm workflow validate ./my-workflow --json # emit the result as JSON, and nothing elseOne pass reports every warning beside the error (issue #424). A workflow can be wrong in more than one way at
once, and a hard error no longer hides the rest: when validation fails, the loader warnings that
were live at the moment it failed — a typo and its did-you-mean, a retry advisory — print above
the error, so a single run gives you the whole defect set. Previously the warnings unwound with
the failure and only surfaced once the error was fixed, one layer per round trip. The same holds
for realm workflow register and realm workflow watch.
validate runs the same admission path as register (issue #553): the file loader (agent-profile
resolution, the context_wrapper/workflow_context rules), the project-extensions pass (extension
modules, the deployment manifest, config_schema two-pass) and real-then-sentinel secret
resolution — so what validate blesses register accepts, and what register refuses validate
refuses, with the same message. --registered runs the same rules on the stored copy and, per
check, resolves agent profiles against the recorded source_dir and the manifest against the
recorded trust_root when the recorded paths still exist; otherwise it prints exactly which checks
it could not run and why (N check(s) not run: <check> (<reason>)[; <check> (<reason>)] — each
check beside its own reason — checks_not_run in --json; never a --strict failure).
--strict (issue #169): by default, a non-fatal loader warning (an unknown workflow/step key,
a retry-without-timeout advisory, a sentinel-credential fallback) is printed but the command still
exits 0 — the workflow is Valid: regardless. --strict changes that: if any warning is
present, the summary line names the count and the command exits 1 instead. Nothing is checked
that isn’t already checked without the flag — --strict only changes what counts as success, so
it’s meant for CI gates that want to catch a typo before it reaches register.
--registered <id> (issue #427): audits the STORED copy of a registered workflow instead of
a file — the pre-upgrade check. It strips the keys the loader stamps at registration, feeds the
rest back through the real loader, and reports what re-registering that definition today would
say. Current-schema copies stay grandfathered at runtime against LOADER changes only — the
audit changes nothing about your runs; it tells you what you would hit if you re-registered. A
legacy (re-register) copy is already unreachable, and the audit tells you so.
Grandfathering does not cover every kind of change (issue #508). A new LOADER rule only
applies going forward — a definition that was clean when it registered keeps running, and this
audit is how you find out what re-registering it today would now refuse. A new engine-side
dispatch check is different: it reads the CURRENT stored definition on every dispatch, so it
applies immediately to every registered copy, upgrade or not — this release’s trust-value refusal
is one (an invalid trust on an already-registered step starts refusing at that step’s next
dispatch, with no re-registration involved). This audit still tells you about it (the loader now
validates trust too), but “the audit changes nothing about your runs” is only true for the
loader-only class.
Two things worth knowing. It audits the loader you have INSTALLED, so the pre-upgrade journey is
upgrade the CLI first, then audit — which is safe for the loader-only class precisely because that
grandfathering holds; it does NOT make an already-registered engine-side defect (like an invalid
trust) safe to leave unaudited. And it is SUPPLY-OR-DECLARE, not blind to the source tree:
agent-profile resolution and the project-extensions pass (modules, manifest, config_schema) each
need a piece of it, and per check the audit uses the RECORDED source_dir/trust_root when that
path still exists — running the check for real — or, when it does not, says so on its own line
instead of guessing (N check(s) not run: <check> (<reason>)[; …], each check beside its own
reason). --extensions-module applies wherever the extensions pass actually runs, and its reason
gains ; --extensions-module not applied when the pass itself could not run. checks_not_run (in
--json) and the line above both list only checks skipped for a missing recorded path — a refusal
ends the audit outright and is its own reason, so a valid: false result always shows
checks_not_run: [] even though those checks were never reached.
extension_keys (issue #559) has the identical shape. It lists the top-level x- keys the
definition carries — realm’s author-extension namespace, carried verbatim and never read — and is
present on every arm: the accepted list on a success arm, [] on every refusal or load-failure
arm. As with checks_not_run, [] on a valid: false result means “nothing was ACCEPTED”, not
“the file has none” — a refused file may well carry x- keys the audit never got far enough to
report.
--explain (issue #422): prints the full per-step structured_output adoption detail
described below, in place of the one-line summary a default run prints. It changes nothing about
what is validated and nothing about the exit code.
--json (issue #454): stdout carries only the JSON object — every other stdout channel (the
description line, the structured_output nudge, the Extensions: manifest line, and
--registered’s audit headers) is suppressed. Advisories printed WHILE the checks ran — the
manifest-secrets ⚠ block, the sentinel-credentials line — may still reach stderr; they are not
part of the JSON contract and --json does not silence them. Works in both modes; exit codes are
unchanged either way.
{ "valid": true, "mode": "file", "path": "./my-workflow", "workflow_id": "my-workflow", "loader_version": "0.41.0", "schema_version": null, "error_count": 0, "warning_count": 1, "strict": { "requested": false, "failed": false }, "diagnostics": [ { "code": "RETRY_NO_TIMEOUT", "severity": "warn", "scope": "step", "step": "sync_data", "message": "Step 'sync_data': declares 'retry' but no 'timeout_seconds' — …" } ], "errors": [], "checks_not_run": [], "extension_keys": []}A file that also carried a top-level x-category-enum key would show "extension_keys": ["x-category-enum"] on this same success arm — nothing else in the object changes shape.
valid is the boundary truth — would a plain register accept this file — reported SEPARATELY
from strict: a warning-only workflow is valid: true even when --strict --json fails and exits
1, because --strict is a run mode you asked for, not a property of the file. mode is "file"
or "registered"; path and schema_version are null in whichever mode does not apply.
workflow_id is the parsed definition’s id once one exists, otherwise the id you asked for
(--registered <id>) or null if nothing parsed far enough to have one. diagnostics is issue
#169’s structured loader-warning channel, one entry per warning, with severity always the
EFFECTIVE severity under the default policy — never --strict’s all-error mode, which stays its
own strict.failed bit. errors holds one string per hard failure (issue #402’s per-step
boundaries survive as separate entries); channel prefixes like Error:/Invalid: are stripped,
except two sentences that ship whole because the prefix IS the composed message (an unresolvable
extensions module, and a warning escalated to an error by policy). extension_keys is the accepted top-level x- keys, in authored order (issue #559). Not
represented at all: the structured_output adoption nudge, --explain’s per-step detail (inert
under --json), and --registered’s audit headers — human-informational, not part of the
contract. For a SCHEMA_STRICT_ADVISORY entry (issue #586) key names the schema BLOCK (params_schema,
input_schema, …) while line/column/endLine/endColumn span the KEYWORD the advisory is
about inside it — two granularities in one object, by design; a jump-to goes to the position, a
“which block” filter reads key.
$ realm workflow validate ./my-workflow⚠ step 'sync_data': unknown key 'dependson' (line 14) — ignored (did you mean 'depends_on'?)Invalid: 1 warning, 1 escalated to an error by policy: UNKNOWN_STEP_KEY 'dependson'$ echo $?1Every unknown-key warning line reads — ignored (…) — a true statement about what the parse did
with the key, whether or not this run goes on to refuse the workflow over it (issue #540) — except the three targeted x-
messages (a step-level x- key, the reserved x-realm- prefix, and a case-only miss such as
X-Foo), which replace that tail with the remedy itself (issue #559).
The line below the warnings is what says whether — and which — warnings the run actually refuses
over: an unrecognized key is refused with or without --strict (issue #170),
and the policy check runs first — so this class never reaches the failing due to --strict line
any more. --strict still does its job for everything that is genuinely a warning under the
default policy: a retry advisory, a dead config block, an unrecognized key inside retry: or
gate:. Those warnings print the identical — ignored line above them; only the Invalid: … escalated to an error by policy: … line — present only when something actually escalates — tells
you which. The one exception: a top-level x- key mints no warning at all — it is the
author’s own extension namespace, never refused (issue #559; see
yaml-schema.md’s “Extension namespace” section); validate, register and
watch instead name every one they accept, on the verdict line itself.
realm run, realm agent, and realm listen are unaffected — they load leniently, so a workflow
already deployed with an unknown key keeps running.
The (line 14) is the line the offending key sits on (issue #392).
Two forms appear, and they mean different things. (line N) always points AT the offending key.
(step at line N) points at the step’s own line — used where the message is about the step rather
than one key, and as the fallback when a key cannot be placed exactly. A step that trips several
step-scoped checks reports the same (step at line N) on all of them rather than a different line
per message. Workflow-scoped errors — a dependency cycle, a missing required field, a bad
services entry — do not carry a line yet; some of them, like a missing field, have no key in the
file to point at. A position is omitted rather than approximated whenever it cannot be resolved
exactly.
The structured_output adoption nudge (issues #236, #422): validate looks at every
execution: agent step with an effective schema and works out whether it could adopt
structured_output: strict, and what — if anything — stands in the way. All of it is pure
information on its own line (ℹ, never a ⚠ warning); none of it affects --strict’s exit code.
A default run says it in one line, graded by how far each step has to go:
$ realm workflow validate examples/02-ticket-classifierValid: ticket-classifier v1 (4 steps)ℹ 2 steps ready for structured_output: strict (2 with caveats) — run 'realm workflow validate --explain' for detail (REALM_NO_NUDGE=1 to silence).--explain swaps that line for the per-step detail — which step, which caveat, which one line it
is short. REALM_NO_NUDGE=1 silences the summary for good; an explicit --explain still prints,
because asking for detail beats a standing preference not to be told.
A step that has already opted in is different, and always loud. Its caveats print on every
run — --explain does not gate them and REALM_NO_NUDGE does not silence them — and it is left
out of the summary’s counts entirely. The rule behind that split: advice about config you
DECLARED is a diagnostic, and advice about config you COULD adopt is one line you can turn off.
See structured_output for the full
gate table.
realm workflow register <path>
Section titled “realm workflow register <path>”Registers a workflow in the local store (~/.realm/workflows/). The stored copy carries the
version: the workflow file declares — re-registering the same file overwrites the copy and keeps
that number (the version is yours to bump, never incremented for you). Fails immediately if any
agent profile declared in the workflow is not found in profiles_dir.
realm workflow register ./my-workflowrealm workflow register ./my-workflow --strict # refuse to persist if any warning is presentRegistering mints the trust decision for project extensions: when the workflow declares
extensions:, the modules are fully loaded and duck-validated and step config gets the
config_schema two-pass — all before anything is persisted. See the
Project extensions guide.
--strict (issue #169): refuses to persist a warning-bearing workflow — the warnings are
printed, store.register is not called, and the command exits 1 with
Error: '<id>' v<version> has N warning(s); refusing to register due to --strict. The warning
population is the loader’s — an unknown key inside a block such as retry: (unknown keys at the
STEP level are refused by policy with or without the flag, so --strict never sees them as a mere
warning; a top-level x- key mints no warning at all and so is never counted here either — issue
#559) and the per-entry sentinel-credential fallback; register does not run validate’s
retry-without-timeout advisory, so the two flags do not count the same set (issue #464). Without
--strict, registration proceeds as always: the workflow is persisted and every warning is printed
alongside the Registered: line.
realm workflow watch <path>
Section titled “realm workflow watch <path>”Watches a workflow YAML file and re-registers it into the local store on every change.
Performs an initial registration immediately on startup, then re-registers whenever the
file is modified — no manual realm workflow register required during active development.
realm workflow watch ./my-workflowrealm workflow watch ./my-workflow/workflow.yaml # or point directly at the filePress Ctrl+C to stop watching.
Errors from an invalid YAML edit are logged (with a timestamp) to stderr but do not crash the watcher — fix the file and save again to recover.
The watched file, and the watched directory. Editing workflow.yaml itself is always
survived, including an editor that deletes and recreates it on every save (a vim-style atomic
write) — the invalid moment in between is logged and the recreated file is picked back up on
its own. The directory holding it can also change: a fast delete-and-recreate (a build step, a
git checkout) is survived too — the watch re-arms itself on the recreated directory and keeps
going. What the watch cannot survive is the directory itself being genuinely gone — deleted with
nothing to replace it, or moved somewhere the watch cannot follow — in which case it prints
Error: The watched directory no longer exists … and exits 1 rather than sitting silently
alive on a directory that no longer means anything; there is no waiting for it to come back,
restart the command once the path exists again. One residual case goes unobserved: if an
ancestor of the watched directory is moved (the workflow’s own directory keeps its name, just
relocated under a renamed parent) and nothing inside it changes afterward, the watch has nothing
to react to and stays silently pointed at the old location — moving the watched directory
itself, by contrast, is detected. The profiles_dir (agent profiles alongside the
workflow, watched for changes too) follows the same replace/re-arm rule but is never fatal on
its own: its disappearance is reported on stderr and profile edits stop being watched, while the
workflow YAML keeps being watched as before.
Development inner loop:
- Start
realm workflow watch ./my-workflowin one terminal. - Edit
workflow.yamlfreely — every save auto-registers. - Run
realm workflow run ./my-workflowor start your MCP session in another terminal.
realm workflow run <path>
Section titled “realm workflow run <path>”Runs a workflow interactively in development mode. For each agent step, prompts for JSON output. For human gates, prompts for approval. Use this to exercise the full workflow without an AI agent.
realm workflow run ./my-workflowrealm workflow run ./my-workflow --params '{"company_name":"Acme"}'--params is validated against the workflow’s params_schema (issue #586); a violation refuses
the run before any run record is created — the same check start_run, start_run_batch,
realm agent and listen apply. It runs AFTER the terminal check below, so a piped invocation
with bad params is told about the terminal first: that refusal holds for the whole wiring, while
the params one depends on the values.
It needs a real terminal (issue #426). Because it prompts for every step and every gate, a
piped or scripted invocation has nothing to answer with — so one is refused up front, before any
run record is created. For scripted flows use realm workflow test (fixture-driven) or
realm listen / realm agent (the production drives).
Cancelling a prompt detaches cleanly (issue #447). Pressing Ctrl-D or Ctrl-C at a prompt in the
terminal used to dump a raw Node stack, which looked like a crash even though nothing was lost. It now exits with a short
map: the run id, the step it was waiting on, the run’s derived phase, and the exact next commands
for that run’s state — respond with the real gate id and choices when a gate is waiting,
realm agent --run-id to re-attach otherwise, and inspect either way. The run is saved; every
settled step persisted before the prompt appeared.
Prompts follow the screen, not the pipe (issue #458). With stdout redirected —
realm workflow run ./my-workflow | tee log — prompts and what you type move to stderr, joining
the cancel map above (which always printed there): the log keeps the run’s narrative, the screen
keeps the conversation. Ctrl-C and Ctrl-D still detach cleanly with the map and exit 1 in this
wiring too — previously Ctrl-C died silently and Ctrl-D exited 0 with a wedged run behind it,
neither one saying so. One corner, noted honestly: redirect BOTH stdout and stderr and you type
blind, but cancelling and the exit code stay truthful either way.
A typo at a prompt re-prompts (issue #459). An answer that is not valid JSON used to crash the
run with an uncaught stack, the run wedged mid-step; it now says why (Not valid JSON: …) and asks
again. An answer that is valid JSON but not an object — 42, null, a list — asks again too: 42
or a list used to be accepted and sealed into the run’s evidence as the step’s output, and null
crashed the run. Enter still means {}. Cancelling is unchanged (above).
The exit code tells the truth (issue #468). Every terminal phase but completed exits 1 —
previously all four exited 0. A stalled workflow prints Workflow stalled — detached from run …
with the same remedy map as cancellation and exits 1. Completed-with-failed-steps still exits 0
(outcome-keyed; the run-health finding is the disclosure).
Id-addressed follow-ups need the workflow registered (issue #456). A dev-mode run never
registers its workflow, so the id-addressed commands that come after it — respond, resume,
drain, and realm agent --run-id — say so the first time one runs into it: Workflow not found: … — most often this run was created from a file without --register. Register the workflow (realm workflow register <file>) and <verb> again.
The remedy is per CLASS, not always “register it again” (issue #558). A run’s registered workflow copy can be unreachable in more than one way, and registering is the right answer for only one of them. Every run-context refusal now names which:
| what is wrong | what the refusal says | the way out |
|---|---|---|
| the copy is missing | Workflow not found: <id> — most often … |
register the workflow from its file |
| the copy exists but cannot be read | the registered copy of '<id>' could not be read (<errno>: <path>) |
make that path readable (chmod u+r; a permission, a mount) |
| a directory sits where the file belongs | <path> is a directory, not a workflow file |
remove the directory (rm -r), then re-register from source |
| the copy is empty or corrupt | … is empty (0 bytes) — not a workflow / … is not parseable JSON: … |
re-register from source; if the file was never registered from a source, remove it (rm) — for a registered copy, removing it repairs nothing (the next attempt says “not found”) |
| the copy is a legacy record | This workflow was registered with an older version of Realm. … |
re-register it (the message already says how) |
| the registry directory cannot be read | the workflow registry at <dir> cannot be read (<errno>) |
make that directory readable and searchable (chmod u+rx) |
| an AGENT created it and its copy is gone | Workflow '<id>' not found — this run's workflow was created by an agent (create_workflow) and its stored copy is gone; there is no source file to register. |
there is nothing to register — end the run |
Each of those also carries the way to end the run today. For a run with no open gate that is
To end the run: realm run abandon <run-id>. For a run waiting on a human gate it says instead that
the gate cannot be answered until the copy reads and that abandon refuses a gate-waiting run,
and its repair clause ends in the answer command: This run is waiting on human gate '<step>', which cannot be answered until its workflow can be read; realm run abandon refuses a run that is waiting on a gate. To repair: make <path> readable (chmod u+r <path>), then answer the gate (realm run respond <run-id> --gate <gate-id> --choice <one of: …>). On a run that is already terminal the refusal says so —
The run is terminal (<phase>); there is nothing to <verb>. — and, when the copy itself is
broken, still names the repair, because that copy is shared with every other run of that workflow.
realm run list --stuck selects every run whose workflow copy cannot be resolved, so the class is
visible without opening each run.
realm workflow test <path>
Section titled “realm workflow test <path>”Runs fixture-based tests against a workflow. Loads fixtures from the specified directory, executes each scenario with mocked services and pre-built agent responses, and checks expected final states and step outputs.
realm workflow test ./my-workflow --fixtures ./my-workflow/fixtures/Fixture format (fixtures/happy-path.yaml):
workflow: my-workflowdescription: 'Complete happy-path run'params: {}steps: gather_input: output: summary: 'the collected information'expected_final_state: completedrealm workflow migrate
Section titled “realm workflow migrate”Back-fills the origin field on workflow definition files created before provenance tracking
was introduced. Run this once after upgrading from an older version of Realm if
realm run inspect shows missing origin values in workflow metadata.
realm workflow migrateScans ~/.realm/workflows/, adds origin: human to any file that does not already have an
origin field, and reports how many files were migrated vs. already up to date. Safe to run
more than once — files that already have origin are skipped.
Standalone agent
Section titled “Standalone agent”realm agent
Section titled “realm agent”Runs a workflow end-to-end using a real LLM — no MCP client, no IDE, no running server. Auto steps execute immediately; agent steps are driven by the LLM; human gates pause until a choice is submitted.
realm agent \ --workflow ./my-workflow \ --params '{"key":"value"}'| Option | Default | Description |
|---|---|---|
--workflow <path> |
(required) | Path to workflow.yaml or its containing directory |
--params <json> |
{} |
Run params as a JSON string — validated against the workflow’s params_schema; a violation refuses the run |
--provider <name> |
auto | LLM provider. Values: openai, anthropic. Auto-detected from whichever API key is set. |
--model <name> |
provider default | Model name override. Default: gpt-4o for OpenAI, claude-sonnet-4-5 for Anthropic. |
--base-url <url> |
— | Base URL for OpenAI-compatible endpoints (DeepSeek, Qwen, Groq, etc.). Only valid with --provider openai or when OPENAI_API_KEY is set. |
--strict-base-url |
off | Attest that the --base-url endpoint genuinely enforces structured-output strict mode (issue #313). Without it, strict is never sent to a compat endpoint and every strict-declared step records compat_endpoint. An attestation, not a verification — realm cannot detect an endpoint that accepts strict and ignores it. Affects opted-in steps only. |
--provider-module <path> |
— | Path to a custom provider module. Cannot be combined with --provider, --model, --base-url, or --strict-base-url. See Custom providers below. |
--register |
off | Persist the workflow to ~/.realm/workflows/ so realm run inspect resolves it by ID |
--run-id <id> |
— | Attach to an existing run instead of creating a new one. Mutually exclusive with --workflow. The run must exist, must not be in a terminal state, and its workflow must be registered (--register at first run, or realm workflow register <file>) — a run created from an unregistered file cannot be re-attached, because the definition was never persisted. |
--schema-retries <n> |
2 |
In-drive repair attempts when an agent step’s output/input is rejected by schema validation (issue #217) — the drive re-prompts with the validator’s errors appended. Non-negative integer; 0 disables. |
--llm-timeout <seconds> |
600 |
Per-ATTEMPT ceiling for each model request (issue #401). A step’s own llm_timeout_seconds wins; this fills in for every step that authored nothing. Positive integer. When a request exceeds the derived ceiling the drive stops, records aborted_by_budget with the derived ceiling, plus the declared per-attempt value when a step or the flag set one, and names this lever. |
LLM key: set OPENAI_API_KEY or ANTHROPIC_API_KEY in your shell or .env file. The
CLI loads .env automatically on startup.
Gate handling: when the run reaches a human gate, realm agent prints the gate message
and a realm run respond command. Optionally configure Slack so the message is delivered
there and the gate can be resolved from a Slack thread reply. The active Slack mode is
selected automatically from the env vars that are present:
| Env vars set | Gate mode |
|---|---|
| (none) | Terminal-only. Gate message printed to stdout. Respond with realm run respond. |
SLACK_WEBHOOK_URL |
Mode 1 — gate notification posted to Slack. Terminal command required to respond. |
SLACK_BOT_TOKEN + SLACK_CHANNEL_ID + SLACK_APP_TOKEN |
Mode 2 — Socket Mode. Reply in Slack thread; resolves in < 1 s. |
SLACK_BOT_TOKEN + SLACK_CHANNEL_ID + SLACK_SIGNING_SECRET |
Mode 3 — Events API. Reply in Slack thread; resolves in < 1 s. |
For step-by-step Slack app setup and the full env var reference, see Slack Gate Modes.
OpenAI-compatible endpoints: use --base-url to point realm agent at any OpenAI-compatible
API, including cost-efficient alternatives:
# DeepSeekOPENAI_API_KEY=sk-... realm agent --workflow ./my-workflow \ --model deepseek-chat \ --base-url https://api.deepseek.com/v1
# GroqOPENAI_API_KEY=gsk_... realm agent --workflow ./my-workflow \ --model llama-3.3-70b-versatile \ --base-url https://api.groq.com/openai/v1OpenAI reasoning models: o1-series models (o1, o1-mini, o1-preview) do not support tool
calling. If your workflow uses MCP tool-enabled steps, realm agent exits at startup with an
error when an o1-series model is selected. Use --model gpt-4o (or any non-o1 model) or
--provider anthropic to run tool-enabled workflows. All other OpenAI models — including o3,
o3-mini, and o4-mini — support tool-enabled steps.
Custom providers: use --provider-module to supply a fully custom LLM implementation.
The module must export an instance (not a class) extending LlmProvider from
@sensigo/realm-cli/agent:
import { LlmProvider } from '@sensigo/realm-cli/agent';
class MyProvider extends LlmProvider { async callStep(prompt: string, inputSchema?: Record<string, unknown>) { // call your LLM here return { result: '...' }; }}
export default new MyProvider();realm agent --workflow ./my-workflow --provider-module ./my-provider.js--provider-module cannot be combined with --provider, --model, --base-url, or --strict-base-url.
If the default export is not an instance of LlmProvider, realm agent exits with a
descriptive error before the run starts.
Run commands
Section titled “Run commands”Operations on workflow run instances stored in ~/.realm/runs/.
realm run list
Section titled “realm run list”Lists all runs, sorted by most recent first.
realm run listrealm run list --workflow <workflow-id> # filter by workflowrealm run list --status <phase> # filter by run phaserealm run list --stuck # only wedged/idle runs (typed run-health classification)realm run list --stuck --older-than 6h # override the idle-age threshold (default 24h)Valid --status values: running, gate_waiting, completed, failed, abandoned, aborted.
When filtering by gate_waiting, each line also shows the gate step name and gate age (time since the gate opened).
--stuck (mutually exclusive with --status) shows only runs flagged by the same health
classification get_run_state and inspect use. Nine finding kinds select a run onto this
list: a stale or
unknown-age claim, a wedged non-gated sibling on a gate_waiting run, a capability block, a
running run with no claimed step idle for at least the active threshold (default 24h — printed
in the header as (threshold 24h)), a terminal run with an undrained finalizer, a failing drive,
an expired gate, a corrupted gate record, and a terminal run still carrying a pending gate.
Two more kinds are structurally absent here: resolved_gate_with_eligible_guard and
trust_value_invalid (issue #508). Both producers require a workflow definition and list
classifies definition-free, so neither ever fires on this surface. Two further kinds never
select: completed_with_failed_steps (issue #302) and structured_output_downgraded (issue
#316) — a completed run and a degraded-assurance disclosure are not “stuck” symptoms. Either can
still appear on a run selected by one of the nine. The never-claimed check is age-gated: a
run simply between agent drives is no longer flagged the instant its last claim settles (a
disclosed behavior change from the prior unconditional check — see the CHANGELOG). Each flagged
line appends its idle age plus finding labels.
The labels — one row per label form; the claim row serves two kinds:
| label | means |
|---|---|
<step>=claim_stale / <step>=claim_unknown_age |
the step’s claim is past its deadline, or carries none to judge by. On a running run the kind behind it is stale_claim; behind an open gate it is wedged_gate_sibling — the distinction lives in the run’s shape and in inspect’s finding text, never in the label |
<step>: needs <kind> '<name>' |
the step is blocked on a capability the registry does not have |
<step>=<reason> (realm run drain) |
a terminal run with an undrained finalizer — the pointer is the fix. <reason> is never_leased, lease_expired, or lease_held |
<step>=drive_failing(<class>) |
the drive died on this step; <class> is the failure class (fuller vocabulary in mcp-protocol.md) |
<step>=gate_expired(<disposition>), and <step>=gate_expired(finding_only) (realm run respond) |
the gate is past expires_at. abort and settle_default enact themselves at the next enactment point (realm run drain --expired can force them). finding_only means nothing will enact itself — a human response ends it: realm run respond <run-id> --gate <gate-id> --choice <choice>, with the gate id and choices shown by realm run inspect. A wrong choice is safe: the run refuses with Choice '<x>' is not valid. Expected one of: … and no state changes |
<step>=gate_corruption |
a settled gate entry coexists with a live pending gate of the same id — a store that diverged, not an engine path |
<step>=stale_gate (realm run purge) |
a terminal run still carrying a pending gate — one written by an older realm, before terminal seals stripped their gates. Purge DISPOSES of the record; a failed or abandoned run WITH a failed step can instead continue via realm run resume --from <step> (the phase alone is not enough — the step must be in failed_steps; purge’s own dry-run counts how many selected runs qualify). For a grandfathered record disposal is the usual intent, which is why the label points at purge |
definition_unresolvable (<class>) (realm run inspect <run-id>) |
this run’s registered workflow copy cannot be resolved (issue #558) — produced for every LIVE run (a terminal run’s copy matters only to replay/drain, which say so themselves). The pointer is inspect — the surface that names the repair, the consequence and the way out (for a gate-less run, realm run abandon; for a gate-waiting run, the answer command — its copy must be repaired before the gate can be answered, and abandon refuses it) — never the destructive act alone. <class> is missing, unreadable, not_a_file, empty, registry_broken, corrupt (not JSON, or JSON that is not a workflow object), legacy, or unknown (a failure realm could not classify — never an error code). realm run inspect <run-id> carries the full sentence, the way out and the repair |
never_claimed_idle carries no label of its own: it IS the reason the run is listed, and the
header’s threshold already says so. --older-than <duration> overrides the idle-age
threshold (e.g. 30m, 6h, 7d; requires --stuck); --older-than 0m restores the old
unconditional breadth (bare 0 is rejected — use 0m).
A registry entry larger than 4 MiB is never parsed by the listing — a listing must not be the
surface that detonates an expansion bomb (issue #557) — and an unread copy is never judged: it is
not a finding but a check that did not run, stated on one stderr line per definition —
⚠ workflow definition <id> (<size> MiB, <n> run(s)) was not inspected by --stuck (over the 4 MiB listing cap): these runs were not checked for a broken definition; realm run list --workflow <id> lists them, realm workflow validate --registered <id> reads the copy.
The line prints before the rows, and the stdout verdict carries the count and the ids (No stuck runs found (threshold 24h; 1 definition not inspected: <id>). / Stuck runs (threshold 0m; 2 definitions not inspected: <id>, <id>):), so a reader of stdout alone knows the sweep was partial and where. A gate-waiting run on such a copy has no finding of its own and is
not listed — realm run list --workflow <id> shows it. The cap applies to the listing only; get()
and every other surface read the copy in full.
On a corrupt record a label may render (unknown) in place of its enum value — that is the armor,
not a vocabulary member. Labels of the SAME cause family join comma-separated with the step
repeated per label (x=stale_gate (realm run purge), x=gate_corruption); different families
render as separate two-space-separated segments. The timestamp column is updated_at.
This is a different flag from realm run reclaim --older-than, which is a deadline-margin add-on for --all auto-reclaim selection, not
an idle-age threshold. See Operating & recovering runs.
Output per run: run-id workflow-id vN run_phase timestamp N step(s)
N step(s) is the count of distinct steps that produced evidence (retried steps count once;
gate responses are excluded). This is the same count shown in
realm run inspect under Evidence (N steps):.
State colors: green = completed, red = failed/abandoned, cyan = gate_waiting, yellow = in-progress.
realm run inspect <run-id>
Section titled “realm run inspect <run-id>”Prints the full run record and evidence chain for a workflow run. This is the primary debugging tool — use it whenever a run fails, gets stuck, or produces unexpected output.
realm run inspect abc12345-0000-0000-0000-000000000000realm run inspect abc12345-0000-0000-0000-000000000000 --check-driftFor runs of extension-declaring workflows, the output includes the Extension Identity
history (drift evidence: per-module entry hashes, the dir_tree_v1 fingerprint, signals,
override/error flags, and the coverage sentence). --check-drift recomputes the LAST
recorded entry against current disk state under its recorded rules — pure hashing, no
code loading — and prints same/DIFFERS/MISSING per component. See the
Project extensions guide.
Output format
Section titled “Output format”Run: abc12345-0000-0000-0000-000000000000Workflow: incident-response v3Phase: completed ✓
Created: 2026-01-15T10:30:00.000ZUpdated: 2026-01-15T10:30:42.000Z
Evidence (3 steps):
1. read_alert success 12ms hash: f3a9b2c1 Input: {} Resolved: {"path":"/tmp/alert.md"} Output: {"content":"## SEV-2 Alert\nDisk usage on prod-db-1 at 94%","line_count":14} Diagnostics: ~22 tokens | no preconditions
2. analyze_cause [profile: senior-sre] success 8432ms hash: 2d7e4f81 Input: {"content":"## SEV-2 Alert\nDisk usage on prod-db-1 at 94%"} Output: {"root_cause":"log_rotation_disabled","severity":"sev2","affected_system":"prod-db-1"} Diagnostics: ~1840 tokens | preconditions: analyze_cause.result.content != "" → true ("")
3. draft_response [profile: senior-sre] success 5211ms hash: 9c3b1a0f Input: {"root_cause":"log_rotation_disabled","severity":"sev2","affected_system":"prod-db-1"} Output: {"message":"SEV-2 on prod-db-1: log rotation was disabled causing disk accumulation…"} Diagnostics: ~2100 tokens | preconditions: draft_response.result.root_cause != "" → true (log_rot…)The seal, and any ruling on it (issue #367)
Section titled “The seal, and any ruling on it (issue #367)”A terminal run renders the recorded fact that ended it, directly above the evidence:
Sealed by: guard_abort (g)Cause: Guard 'g' aborted the run.The step in parentheses appears only where the step IS the seal’s identity — a guard, a gate, or a handler abort. A run that failed in several places does not name one of them here, because the step recorded on that kind of seal is whichever one settled last, and printing it directly above the cause line would read as the culprit.
(recovered by classifier) after the arm means the run predates the seal substrate and the engine
recovered its arm on read rather than reading a stamped one. Both are the truth about the run; the
marker tells you which kind of truth it is.
When an operator has adjudicated the seal, one further line records who and when:
Sealed by: completeRuled: mihai at 2026-08-21T00:00:00.000Z (was step_failure) — the retry succeeded; the earlier failure is not the outcome(was <arm>) names the arm the ruling replaced; a first stamp on a record that never had one reads
(first stamp — no prior arm existed) instead. The reason, when the ruling carries one, is printed
verbatim and never shortened.
by is a recorded CLAIM of identity, not a verified one: there is no auth model behind it, so
treat it as attribution-by-assertion. The run’s own prose is never rewritten to match a ruling — it
stays as the historical evidence of what the engine said at the time.
State colors: green = completed, red = failed or abandoned, yellow = anything else (including gate_waiting and in-progress).
Output truncation: Input and Output fields are truncated at 120 characters. A … suffix
indicates truncation — use realm run replay to re-evaluate with modified values.
Human gate steps: A gate step appears in the evidence chain as a single entry — the same
step ID covers both the gate opening (when the engine paused for human input) and the gate
response (when a choice was submitted via realm run respond). The step count does not
increase when a gate is responded to. When gate.message is configured, the inspect output
shows a Message: line under Choice: with the exact text the human saw at decision time.
Field reference
Section titled “Field reference”| Field | What it tells you |
|---|---|
| Phase | Current run phase (run_phase). Terminal runs show ✓ (completed) or no suffix (failed/abandoned). |
| Rerun of | Present only when this run SUPERSEDED a terminal run under the same idempotency key (on_terminal_match: 'rerun'/'rerun_if_failed') — the id of the run it replaced. Absent on a first run and on a reuse. |
| Evidence (N steps) | Number of distinct steps that produced evidence. Steps with multiple attempts are counted once. Human gate steps are counted once regardless of whether the gate has been responded to. |
| Step number | Execution order, 1-based. |
| Step name | The id of the step in your workflow YAML. |
[profile: ...] |
Which agent profile handled this step. Present on agent steps only; absent on auto steps and human gates. |
| Status | success (green), error (red), or other engine-assigned state (yellow). |
| Duration | Wall-clock time the step took to complete. High values on agent steps are normal. |
hash: XXXXXXXX |
First 8 characters of the SHA-256 chain hash. The hash covers all evidence up to and including this step — it changes if any prior step’s output changes. Use it to detect replay divergence. |
| Input | What the caller passed to the step. For the first step: the run params. For subsequent steps: the output of the prior step. For input_map steps this is typically {} — see Resolved for the actual adapter params. |
| Resolved | The params the service adapter actually received, assembled by the engine from input_map dot-paths. Only present on execution: auto steps that declare input_map. The token estimate in Diagnostics is computed from this value, not Input. |
| Output | What the step produced. For agent steps: the JSON the AI returned. For auto steps: the handler return value. For adapter steps: the raw adapter response injected by the engine. |
Diagnostics: ~N tokens |
Estimated token count of the context window passed to the agent for this step. Useful for spotting steps that approach model context limits. |
| Diagnostics: preconditions | Each precondition expression, whether it passed (→ true) or failed (→ false), and the resolved value in parentheses. If a step ran unexpectedly or was blocked, this is where you look. |
What to look for
Section titled “What to look for”Run failed — step shows error:
Read the failed step’s Output field. For handler steps, the error message is in the output.
For agent steps, the output may be missing required fields — compare its shape against the
input_schema of the step that consumes it in your workflow YAML.
Run stuck at a gate (gate_waiting state):
The state line will show gate_waiting in yellow. Look at the last entry in the evidence
chain — it will be the step that produced the gate. Submit the gate response with
realm run respond <run-id> --gate <gate-id> --choice <choice> — the gate id is on the Gate:
line, the choices on the Choices: line below it.
Precondition blocked a step unexpectedly:
Find the step that was supposed to run and look at its precondition_trace. Each expression
shows the actual resolved value in parentheses. The value will tell you whether the prior step
produced the right field name, type, or content. Cross-reference with the prior step’s
Output field to see what was actually returned.
Agent returned wrong output shape:
Find the agent step in the evidence chain and read its Output. Compare the field names and
types against the input_schema of the next step in your YAML. The mismatch will be visible
— missing keys, wrong types, or extra nesting are common causes of precondition failures.
realm run resume <run-id>
Section titled “realm run resume <run-id>”Resets a failed or abandoned run so a specific FAILED step can re-execute — the step named by --from must be in the run’s failed_steps, so a run with none has no resume path.
realm run resume abc123 --from <step-name>realm run abandon <run-id>
Section titled “realm run abandon <run-id>”Explicitly abandons a non-terminal run — marks it terminal with phase abandoned. Use it on a run
that is parked and not coming back (see realm run list --stuck).
realm run abandon abc123realm run abandon abc123 --reason "stale runner, no longer processing"Idempotent (abandoning an already-abandoned run succeeds as a no-op — the one terminal state it accepts). Refuses any other terminal run, and
refuses a run with an OPEN gate (pending_gate; resolve it via realm run respond first) — the refusal prints
the exact command, with this run’s id, gate id and the gate’s own choices. The refusal keys on the gate
itself, never the persisted run_phase label: a record whose label says gate_waiting but carries no gate
has nothing to answer and abandons like the running run it derives to. It releases every
claim in the same write, so a just-abandoned run is purgeable as soon as it passes
--older-than. To run the same work again: realm workflow run <source_dir> — the directory the registered copy records it was registered from, printed when the copy reads, with this run’s own params (--params '<json>') when it had any; a blank to fill (<the workflow.yaml you registered '<workflow_id>' from>) when it does not (a fresh
run; this run’s evidence stays at realm run inspect <run-id>). Over MCP, start_run with
on_terminal_match: 'rerun' supersedes it instead, and the fresh run is stamped rerun_of. See
Operating & recovering runs.
There is no operator abort verb: a workflow’s own guard abort (abort_unless) is the graceful
path that runs finalizers; abandon is a kill and runs nothing. Every successful abandon prints
this reminder.
realm run respond <run-id>
Section titled “realm run respond <run-id>”Submits a human gate response. Both flags are REQUIRED — nothing is prompted for.
realm run respond abc123 --gate <gate-id> --choice approveBoth arguments come from realm run inspect: the id on its Gate: line, the choices on the
Choices: line below it. Should you mistype one anyway, the refusal is safe — Choice '<x>' is not valid. Expected one of: …, and nothing is recorded.
Expiry-WINS (issue #291): if the gate carries an authored gate.timeout_seconds and has
already expired unresolved when this is called, the response is refused with an honest
explanation of what actually happened instead — never silently recorded as if it arrived in time.
See gate.timeout_seconds for
the full disposition/disclosure table.
realm run reclaim <run-id>
Section titled “realm run reclaim <run-id>”Recovers a wedged claim (issue #101) — see Operating & recovering runs.
Never enacts an expired gate itself, even when the same run also carries one: if it does, the
result carries an advisory pointing at realm run drain --expired as the enactment lever.
realm run drain [<run-id>] [--all]
Section titled “realm run drain [<run-id>] [--all]”Delivers post-commit finalizers for a terminal run (crash-window recovery, issue #279). Dry-run
by default; --force actually drains. --all batch-drains every terminal run with an actionable
pending finalizer.
realm run drain abc123 # dry-run: report what would runrealm run drain abc123 --force # actually drainrealm run drain --all --force # batch-drain every actionable terminal runrealm run drain abc123 --void <finalizer> # void one pending finalizer instead--expired (issue #291, opt-in): without this flag, drain is byte-stable and
terminal-only — a non-terminal run with an expired gate is completely invisible to it, on both
per-run and --all. With --expired:
realm run drain abc123 --expired # dry-run: also reports an expired gaterealm run drain abc123 --expired --force # also enacts it (settle_default/abort per the frozen disposition)realm run drain --all --expired --force # batch: enacts every expired, enactable gate store-wideA settle_default disposition may or may not terminalize the run (depends on the workflow’s own
remaining steps); an abort disposition always does, and its finalizer terminalization flows
into the SAME drain pass. A finding-only gate (timeout_seconds with no on_expiry) is listed
as expired — finding-only and is never enacted, even under --force.
When a finalizer cannot be run here (issue #558): a --force pass halts at the first finalizer
it cannot run — one the workflow does not declare (left pending — not declared by the workflow definition) or one whose handler is not available on this surface (left pending — handler not available on this surface) — so that one and every finalizer behind it stay pending. drain says
so (nothing drained; N finalizer(s) left pending: … when the pass ran none, Drained run '<id>' (M ran) — N finalizer(s) left pending: … when it ran some), names --void <finalizer> --force
once per pending finalizer (each line carrying this run’s id), and exits 1 — it does not
report Drained run and exit 0, which made realm run list --stuck re-offer the same drain
command forever. A pass that had nothing to run at all (an abandoned run mints no finalizer ledger)
prints has no pending finalizers. Nothing to drain. and exits 0 — Drained run means a drain
happened. The dry run predicts the same per finalizer from the registered copy: a pending finalizer
the workflow does not declare is listed as NOT declared by the workflow definition — --force would leave it pending; to void it: …; a declared one as would lease and run on --force, if its handler resolves on this surface; and when the copy cannot be read, the dry run says so (⚠ could not read this run's workflow copy …), marks every entry unknown — the workflow copy could not be read, so --force will refuse until it is repaired … with its void command, and never invites a --force
that would refuse.
A non-terminal run is not drainable; the message says so with the run’s DERIVED phase and names the
way out from the record — To end the run: realm run abandon <run-id>. for a running run, and for a
run waiting on a human gate the answer command first (Answer its gate first: realm run respond <run-id> --gate <gate-id> --choice <one of: …>; then realm run abandon <run-id>. — abandon
refuses a gate-waiting run, so naming it alone named a command that would fail) — dry-run exits 0,
--force exits 1.
realm run replay <run-id>
Section titled “realm run replay <run-id>”Re-evaluates workflow preconditions with modified step outputs. Shows what would change without executing anything. Useful for tuning extraction schemas and precondition expressions.
realm run replay abc123realm run replay abc123 --with "step_id.field=value"--with syntax: step_id.field_path=literal_value where step_id is the step name,
field_path is a dot-separated path into the step’s output, and literal_value is one of:
trueorfalse— boolean- A number (e.g.
0.85,3) — numeric - A quoted string (e.g.
"high") — string - An unquoted string (any other value) — treated as a string
Multiple --with flags may be specified. Each override applies to the in-memory replay
evidence only — no run record is modified.
realm run diff <run-a> <run-b>
Section titled “realm run diff <run-a> <run-b>”Compares the evidence chains of two runs side by side. Shows which steps produced different results and which fields changed.
realm run diff abc123 def456realm run cleanup
Section titled “realm run cleanup”Marks idle non-terminal runs as abandoned.
realm run cleanup --older-than 30d # abandon runs idle for 30+ daysrealm run cleanup --dry-run # preview without making changes--older-than accepts: Nd (days), Nh (hours), Nm (minutes). Example: 7d, 6h, 30m.
realm run purge [<run-id>]
Section titled “realm run purge [<run-id>]”Permanently deletes a terminal run and every co-located on-disk artifact — the run file, its
idempotency-key pointer, the <id>.attempts.jsonl sidecar, and any orphaned trace-buffer-*.jsonl WAL
files. Irreversible. Unlike cleanup/abandon, this does not just mark a run terminal — it
removes it from disk entirely. See
Operating & recovering runs for the full
abandon-vs-purge distinction.
realm run purge abc123 # dry-run: report what WOULD be deletedrealm run purge abc123 --force # actually delete itrealm run purge --older-than 30d # dry-run over a batch of eligible runsrealm run purge --older-than 30d --force # actually delete the batchrealm run purge --older-than 30d --workflow my-workflow --force # restrict the batch to one workflowDry-run by default — even naming a single <run-id> only reports what would happen until you add
--force (naming a run is selection, not consent). --older-than accepts the same duration format as
cleanup: Nd (days), Nh (hours), Nm (minutes). --workflow <id> restricts a batch to one
workflow (only valid alongside --older-than).
The age a batch measures is last activity, not last progress — and recording a drive failure
counts as activity (issue #401). A run failing
continuously therefore never ages into a batch sweep — and, being non-terminal, it is not
purge-eligible at all. Stop the flapper, make the run terminal (abandon it, or realm run cleanup), then purge.
Safety posture:
- Terminal-only — never touches a non-terminal or
gate_waitingrun. - A run with a future-deadline (
healthy) in-progress claim is never purged, in either mode — a runner is provably still working it, and there is no override. - A run with an indeterminate-age (
claim_unknown_age, no deadline recorded) claim is skipped with a warning in batch mode (a cron sweep cannot prove the runner is dead), but is purgeable via an explicit single-runrealm run purge <id> --force— the operator naming the exact run is a deliberate judgment call a batch sweep won’t make automatically.
Batch mode reports continue-on-error as purged / already-purged (a concurrent purge beat you to it —
not a failure) / failed, plus how many of the selected runs were resumable via realm run resume —
purging one destroys that path permanently.
realm run migrate --stamp-seals
Section titled “realm run migrate --stamp-seals”Materialises the recorded seal arm on every legacy terminal run (issue #367). Records written before
the seal substrate carry no sealed_by; the engine recovers their arm on every read wherever one
is recoverable — and where it is not, the run’s phase still derives correctly from the legacy
ladder — so they are
already correct — this command writes that arm down, which is what makes it visible to external
readers and to the phase stored on disk.
Dry-run by default; --force writes. Each run lands in exactly one bucket:
| bucket | meaning |
|---|---|
| stamped | given its arm |
| already stamped | had an arm. The report splits these three ways: checked against the record, ruled by an operator (a human looked at it — the ruling stands), and unverifiable because the record holds nothing to check the arm against |
| unclassifiable | no arm and nothing to infer one from — printed, never written. Exit 0 (2 with --detailed-exitcode) |
| incoherent | has an arm that DISAGREES with its own prose or markers — printed, never auto-rewritten. Exit 0 (2 with --detailed-exitcode). Adjudicate these yourself; a ruled record stops appearing here |
| skipped | a concurrent writer moved the record; their own next write path owns it |
| failed | an infrastructure error on one record; the sweep continues, exit 1 |
updated_at is preserved on every record it touches, and that is the whole point. Stamping is
not activity. realm run gc --heal materialises the same phases but resets the retention clock on
everything it rewrites, so run this command BEFORE --heal after upgrading across #367 if
those clocks matter to you.
version is bumped on stamped records, so a writer holding a pre-stamp snapshot loses its
compare-and-swap rather than silently erasing the stamp.
Parked incoherent records: the format for adjudicating them is already published. A recorded
seal arm is immutable while the run is terminal — every rewrite is refused — with exactly one lawful
exception: an adjudication write carrying truthful provenance (by, at, previous_arm, and
optionally reason). A ruling that misnames the arm it overwrote is refused exactly as hard as one
carrying no provenance at all, which is what keeps the chain walkable a step at a time. A truthful
SAME-arm ruling is legal too — that is how you close out a parked record you have examined and found
correctly stamped.
A ruling supersedes the record’s own prose. The coherence check that normally guards a stamp is skipped once a seal carries a ruling — permanently, not just on the write that records it — because that check exists to catch SILENT drift and a ruling is the opposite of silent. The prose is never rewritten to match: it is the historical evidence of what the engine said at the time, and the ruling resolves the disagreement without falsifying it. A ruled record also stops appearing in this command’s incoherent bucket, which is what actually closes the loop.
Three consequences worth knowing before building on it. There is ONE provenance slot, so a second
ruling overwrites the first: the record tells you what the previous arm was, not the whole history
of rulings, and that loss is deliberate. Changing an arm requires a FRESH ruling — re-using a
previous one’s provenance is a rewrite, not a ruling. And by is a recorded CLAIM of identity, not
a verified one: there is no auth model behind it, so treat it as attribution-by-assertion.
An adjudication write is ordinary activity — unlike stamping, it advances updated_at and
version.
Where a ruling shows up. realm run inspect renders it as a Ruled: line (above), and
get_run_state carries it as sealed_by_adjudicated. export carries the whole record verbatim,
as it always has. realm run list renders none of it — that surface shows one line per run and
rides a later increment.
The operator verb itself ships when it has a customer. The record format will not change when it does.
Residue is the count the report ends on: terminal runs that will still have no recorded arm when this run finishes. That is the unclassifiable ones plus any whose write failed — and, in a dry run, the ones that would have been stamped, since a dry run writes nothing. Skipped records are reported separately and never counted: their arm state is genuinely unknown until their own writer settles.
Exit codes. 0 means the command did its job — including when it found residue and said so.
1 means the command FAILED at its job: a write errored, or the batch was refused. Residue is
chronic by nature (an unclassifiable record stays unclassifiable), so a nonzero exit on residue
would make every scheduled run fail forever, and a chronic alarm is one that gets silenced.
Automation that wants to gate on residue opts in with --detailed-exitcode, which makes it a
three-way: 0 clean, 1 the command failed, 2 succeeded but residue remains — the same contract
grep, git diff --exit-code and Terraform’s -detailed-exitcode use, and opt-in for the same
reason they made theirs explicit.
One compatibility cost, owner-accepted and disclosed here once. An export bundle taken BEFORE a record was stamped, then re-imported AFTER the sweep, will fail with
STATE_RUN_DIVERGED: the bundle carries the pre-stamp version and the store has moved on. Re-export after migrating if you keep bundles for round-tripping.
realm run gc
Section titled “realm run gc”Sweeps orphaned atomic-write .tmp files — crash residue from a process that died between writing a
temp and renaming it over the target (a run record or an idempotency-key pointer). Unlike purge,
which acts by runId, gc cleans up files that belong to no specific run at all. See
Garbage collection for the full
picture, including what it deliberately does NOT reap.
realm run gc --older-than 1h # dry-run: report what WOULD be reapedrealm run gc --older-than 1h --force # actually delete it--older-than is required (no default — omitting it is an error, not “reap everything”) and has a
1-hour floor: a smaller value is rejected outright, even with --force. Accepts the same duration
format as cleanup/purge: Nd (days), Nh (hours), Nm (minutes). Dry-run by default; --force to
actually delete.
Reaps top-level *.tmp files in runsDir and one level into runsDir/keys/*.tmp. Does not reap
orphaned .lock directories (deferred — issue #164) or run-less trace-buffer-*.jsonl WAL files
(issue #163) — the report always names both so you don’t mistake their presence for a bug. A no-op on
Windows (no temp files are ever produced there).
--heal (issue #293, opt-in) — a one-shot pass that rewrites every run record whose persisted
run_phase disagrees with the phase re-derived from the record today, curing the residue left by a
pre-#282 binary (which could mis-derive gate_waiting for a run that had already reached a real
terminal outcome). This is pure population hygiene, not a correctness fix — every live read path
(get_run_state, realm run list --status, the engine’s own eligibility checks) already derives the
phase fresh on every read, so a stale on-disk value is cosmetic residue, invisible unless you read the
raw JSON file directly. The heal writes each mismatched record back unmodified; the store’s own
versioned write tail corrects run_phase (plus version/updated_at) as its ordinary side effect —
gc itself never constructs or edits a single field.
Ordering trap after upgrading across #367 — read before running
--heal. The seal substrate changed whatderiveRunPhasecomputes for some legacy records: a startup death that used to deriveabandonednow derivesfailed.--healrewrites exactly the records whose persisted phase disagrees with the derivation, so on the first run after the upgrade it will rewrite that whole population — and every rewrite resets the record’supdated_at, which is the clock--older-than,cleanupand your retention policy all read. Runrealm run migrate --stamp-sealsFIRST if retention clocks matter to you — it materialises the same phases and leavesupdated_atuntouched, which is exactly what--healcannot do. Composable with--older-than:--healalone runs without it (healing is safe at any age, no 1-hour floor applies),--older-thanalone runs the temp/ artifact sweeps as always, and both together run all three passes in one invocation.
realm run gc --heal # dry-run: list records that WOULD be healedrealm run gc --heal --force # actually rewrite themrealm run gc --heal --older-than 6h --force # heal, plus the temp/artifact sweeps, togetherA single unparseable run file makes list() throw (the store’s deliberate fail-closed read — issue
#132/#183) — --heal aborts with a non-zero exit and heals nothing, rather than silently healing a
partial, possibly-wrong population. This mirrors the orphan-artifact sweep’s own “couldn’t look, abort
loudly” convention above. Postgres and other external stores are out of scope for --heal — their own
records heal automatically the next time anything writes to them through the store’s normal path.
realm run export <run-id>
Section titled “realm run export <run-id>”Archives a run’s evidence — its record, its failed-attempt sidecar, and any orphaned/in-flight WAL
traces — into one self-contained, human-readable JSON file. The read-only, evidence-preserving
companion to purge: keep this before (or instead of) permanently deleting the rest. See
Exporting a run’s evidence for the full bundle shape.
realm run export abc123 # writes ./abc123.realm.jsonrealm run export abc123 --out ~/evidence/ # writes ~/evidence/abc123.realm.jsonrealm run export abc123 --out bug-1234.json # writes exactly that fileWorks on any run, not just terminal ones — exporting a non-terminal run prints a best-effort
snapshot warning (its artifacts are read at slightly different instants and may be mid-flight) but
still produces the bundle; a terminal run needs no warning. Read-only and lock-free: never writes
into runsDir, never deletes anything, and refuses to overwrite an existing file at the resolved
--out target (error + non-zero exit, naming the path — pick a different --out). Excludes the
idempotency-key pointer file by design — it’s a rebuildable index, not evidence, and the key itself
is already on the run record.
realm run attempts <run-id>
Section titled “realm run attempts <run-id>”Shows failed agent-step validation attempts recorded for a run (the durable <id>.attempts.jsonl
sidecar — see Operating & recovering runs).
These are pre-claim validation rejections that never reach the run record, so this is the only
post-mortem trail for them.
realm run attempts abc123 # table: ts, step, error_code, key count, validation summaryrealm run attempts abc123 --json # raw metadata-only records + capped flagRecords are metadata-only (no raw model output). A run with no recorded failures prints a friendly message; if the sidecar reached its size ceiling, the output notes that later attempts were dropped.
MCP server commands
Section titled “MCP server commands”realm mcp
Section titled “realm mcp”Starts the Realm MCP server over stdio. All workflows registered via realm workflow register
are immediately available. Use this command in AI client configs (Claude Desktop, Cursor, VS Code
MCP) that can spawn a local subprocess.
realm mcpClaude Desktop — claude_desktop_config.json:
{ "mcpServers": { "realm": { "command": "realm", "args": ["mcp"] } }}Cursor — ~/.cursor/mcp.json:
{ "mcpServers": { "realm": { "command": "realm", "args": ["mcp"] } }}VS Code — .vscode/mcp.json:
{ "servers": { "realm": { "type": "stdio", "command": "realm", "args": ["mcp"] } }}realm serve
Section titled “realm serve”Starts the Realm MCP server over HTTP with Bearer token authentication. Designed for hosted agent platforms (OpenClaw, Claude.ai, LangChain cloud, custom backends) that cannot spawn a local subprocess via stdio.
REALM_SERVE_TOKEN=<secret> realm serveREALM_SERVE_TOKEN=<secret> realm serve --port 8080 --host 0.0.0.0realm serve --dev # disable auth for local development only| Option | Default | Description |
|---|---|---|
--port <number> |
3001 |
Port to listen on |
--host <address> |
127.0.0.1 |
Bind address |
--dev |
off | Disable auth (local development only — do not expose to a network) |
Environment variables:
| Variable | Description |
|---|---|
REALM_SERVE_TOKEN |
Bearer token clients must send in the Authorization: Bearer <token> header |
REALM_DEV |
Set to 1 to disable auth — equivalent to the --dev flag |
The server refuses to start if neither REALM_SERVE_TOKEN nor --dev / REALM_DEV=1 is set.
Connecting from an HTTP MCP client (e.g. n8n, VS Code remote):
{ "servers": { "realm": { "type": "http", "url": "http://127.0.0.1:3001", "headers": { "Authorization": "Bearer <your-token>" } } }}realm listen
Section titled “realm listen”Starts an HTTP server that routes inbound webhooks to workflows by their trigger: block. For each
request it verifies authenticity per the workflow’s trigger.auth, optionally filters and dedups,
creates a run, and spawns realm agent --run-id (detached) for it.
realm listen [workflows...] [--port <n>] [--host <addr>] [--body-timeout-ms <n>] [--max-body-bytes <n>] [--max-concurrent <n>] [--dedup-store file|memory] [--log-level debug|info|warn|error] [--sweep-expired-gates <seconds>] [--llm-timeout <seconds>]Arguments:
[workflows...]— workflow files or directories to mount (default:./workflow.yaml). Each must declare atrigger:block; workflows without one are skipped (logged).
Flags:
--port <n>— port to listen on (default3000)--host <addr>— interface to bind (default127.0.0.1; a non-loopback host emits a TLS warning)--body-timeout-ms <n>— max time to read a request body before408(default5000)--max-body-bytes <n>— max request body size before413(default1048576)--max-concurrent <n>— in-flight request ceiling before503(default20)--dedup-store file|memory— durable file dedup (default) or in-memory best-effort--log-level <level>—debug|info|warn|error(defaultinfo)--llm-timeout <seconds>— per-attempt fallback ceiling for model requests on the drives this process spawns (issue #409). A step’s ownllm_timeout_secondsWINS; this fills in for steps that author none. No default here, deliberately: absent means each spawned drive uses its own 600-second fallback, and its run record then carries nodeclared_per_attempt_msat all — a fallback nobody chose is never recorded as a declaration. The flag governs the drives THIS process spawns, and nothing beyond them: a run later re-attached withrealm agent --run-idtakes that invocation’s own--llm-timeout, or the default when it carries none — never this one. The clock is per-drive, resolved when the drive starts. Worth knowing because the stuck-run remedy loop sends operators to exactly that re-attach.--sweep-expired-gates <seconds>— opt-in (issue #291), default OFF: runs a coarse, store-wide sweep every<seconds>that enacts every expired, enactable gate this store holds — not just gates on workflows thislistenprocess has mounted. The user-chosen always-on enactor: realm itself still ships no daemon, but runninglistenwith this flag makes that process one. Never drains finalizers itself (no extension registry for a workflow it hasn’t mounted) — a terminalizing enactment logs an advisory pointing atrealm run drain --expiredinstead. Safe to run alongside any other enactment point (submit/execute_step/drain/an attending process’s own timer, or a secondlistensweeper) — races resolve via the same idempotent arm matrix every enactment point shares.
Verification (trigger.auth.mode): shared_secret (header token — e.g. Gorgias Authorization: Bearer …), github / stripe / hmac (body signature), or none (explicit, discouraged escape
hatch). Verification runs before filtering/dedup; failures receive 403 Forbidden (an unknown path
also returns 403, never 404).
Hardening: binds loopback by default; caps and times out the request body before any verification
work; enforces a --max-concurrent 503 floor. TLS termination, rate limiting, and autoscaling are
reverse-proxy concerns — front realm listen with nginx/Caddy/Traefik for public endpoints.
Startup registration & project extensions: each routed workflow is registered once at
startup (the old per-webhook re-register was removed — it silently reverted fresher
registrations; restart realm listen after re-registering a workflow). Workflows declaring
extensions: have their modules loaded at startup, fail-fast — note the modules are imported in
the listen parent process, so top-level side effects run there. Spawned children re-resolve
extensions when they attach; a child that cannot load them exits nonzero and marks its run
terminal_reason: 'extensions_load_failed' (recoverable — see
Project extensions across commands).
realm webhook(GitHub-only) has been removed. Express the GitHub PR flow as atrigger:block withauth: { mode: github }and aparams_mapof the PR payload dot-paths — byte-parity-equivalent to the old hardcoded mapping. Seeexamples/09-webhook-pr-review/.