> ## Documentation Index
> Fetch the complete documentation index at: https://docs.akter.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Runbooks

> Respond to runner, database, migration, and durable-work failures.

# Operator runbooks

**Responsibility:** provide repeatable recovery actions.\
**Authority:** operational.\
**Owner role:** operations/reliability.
**Change policy:** a change requires operator review when a procedure or limit changes.

Every incident procedure records symptoms, deployment and tenant scope, command or execution ids, relevant traces, safe actions, forbidden actions, and recovery proof.

These are design requirements for runtime operations; a procedure is runnable only where its command or signal exists. `akter defects list`, `akter inspect`, `akter export`, `akter receipts show`, `akter dead-letters retry|discard`, `akter subscriptions list --lagging`, `akter subscriptions skip`, `akter dev`, and `akter workflows check` exist ([CLI reference](../api/06-cli.md), or `durable <command> --help`); the operator commands need an operator token whose grant covers the action and resource, and every use is recorded in `durable.operator_audit`; the other `akter` commands named here are not implemented. Metric names are in [observability](03-observability.md).

| Incident | Safe first actions | Recovery proof |
| - | - | - |
| Mailbox backlog | signal `akter_mailbox_age_ms` and `akter_pool_wait_ms`; reduce ingress, inspect oldest envelopes and runner health, add verified shard capacity | age and depth fall without duplicate application |
| Drain deadline expired | read the `Drain deadline expired` warning and its interrupted turn and job counts; let the process exit so its shards move; inspect interrupted jobs, which stay ambiguous until taken over after their lease | callers' retries commit once or replay their receipt; each interrupted job settles once or dead-letters with `ambiguous: true` |
| Fence or lock failures | inspect generation owner and database locks; drain a stale runner | new turns commit under one generation |
| Deterministic defect | signal `akter_turns{outcome="defect"}`; read the cause with `akter defects list --url <each runner> --actor <T> --since 1h`, follow its trace id, check state migration and input; deploy a code/schema fix | same command succeeds or returns a declared error |
| Relay backlog or stuck rows | signal `akter_relay_lag_ms` and `akter_relay_stuck_rows`; check the rows' `last_error` and warnings, runner health, and for jobs that some runner registers the executor | lag returns to 0 and stuck rows to 0 without duplicate receipts |
| Subscription lag | signal `akter_subscription_lag_ms`, `_lag_events`, and `akter_relay_stuck_rows{kind="subscription"}`; list the failing rows with `akter subscriptions list --lagging --url <runner> --tenant <t>` (each row's lag and `last_error`), fix the handler or its input; watch `akter_subscription_pinned_events` against the hold; a row that keeps failing on one poison event: `akter subscriptions skip --source <T>/<id> --subscriber <S>/<id> --subscription <name> --through <cursor> --reason …`, which the subscriber sees as a gap | lag falls to 0; a pruned id-routed row shows in `akter_subscription_undeliverable_gaps` and is reconciled |
| Retryable redelivery loop | inspect dependency, timeout, and activation restarts | one receipt is committed and replayed |
| Job dead letter | inspect with `akter inspect <T>/<id>`, check the provider outcome and idempotency key; then `akter dead-letters retry <jobId> --actor <T>/<id> --reason …` (add `--provider-checked` for an ambiguous one only after the provider confirms it did not apply the call) or `akter dead-letters discard …` | durable success or explicit discard with audit reason |
| Operator request refused, or its audit row unwritable | signal: `401` (no valid operator token; the runner logs a warning with no credential), `403 access_denied` (the grant lacks the action or covering scope; the refusal is audited as `"denied"` in `durable.operator_audit`), or `500` with a trace id when the audit insert fails (a read is refused when its audit row cannot be written, and a repair rolls back whole). Read `durable.operator_audit` for the operator and the capability that was missing; grant the narrowest capability that covers the action and resource, or fix the database fault behind the failed audit write, then repeat the command. An application credential never works on an operator route ([ADR 0050](../decisions/0050-operator-authority-and-audited-repair.md)) | the command succeeds and its audit row names the operator; a denied attempt changed nothing |
| Neki relay backlog | keep ingress bounded, restore relay health, inspect stranded outbox rows | every committed obligation reaches one logical destination intent; receiver receipts deduplicate redelivery |
| Singleton missing or duplicated | verify sharding reachability and ownership, drain conflicting runner | one `run` owner and one cron tick across two runners |
| Parked-socket storm | rate-limit reconnects and inspect edge/runners | activation wake restores parked state; transport loss reconnects and replays durable events without pinning all activations awake |
| Failed migration | a boot migration rolls back whole, so fix forward and restart; a `BadState` refusal names ids that landed out of order: restore from before the higher id ([migrations](02-migrations.md)) | the runtime starts, conformance and state decode pass |
| RLS refuses startup | rerun the RLS script in [deployment](01-deployment.md#row-level-security) for the table or view the error names; a view must belong to the `durable_views` role, never to the tenant role or a superuser | the runner starts; schema-only reads see one tenant |
| Tenant isolation incident | stop affected ingress, preserve evidence, rotate credentials, audit RLS and predicates | cross-tenant probes fail and affected rows are reconciled |
| Postgres primary failover | signal: turns and mints fail `ActorUnavailable` with `SqlError` `ConnectionError` causes across every runner at once. Promote the synchronous standby only, never a lagging asynchronous replica, and move the database address to it; keep runners running, since they reconnect. Callers retry with the same command ids | new turns commit on the promoted primary, retried command ids replay their receipts, and the outbox drains |
| Restore | stop every runner process and ingress, restore one whole-database snapshot, run the clock and job checks, then start runners before ingress ([backup and restore](04-backup-restore.md)) | expired ids fail `CommandExpired`, snapshot intents are delivered once, in-flight jobs retry with their original job id |
| Clock behind restored data | keep runners stopped; correct the database host's clock until the step 3 checks in [backup and restore](04-backup-restore.md) pass: past the backup's snapshot time and in agreement with NTP | `now()` is past the recorded snapshot time and agrees with NTP, and an expired id fails `CommandExpired` |
| Runner survived a restore | stop it at once; restore again from the same snapshot with every runner stopped | no runner started before the restore is running, and each actor's generation rises past the snapshot's on its first turn |
| Replica lag or failure (queries on the primary) | compare `pg_stat_replication.replay_lag` on the primary with query load there; fix or replace the replica named in `Database.postgres({ replica })`; queries keep reading the primary meanwhile ([ADR 0052](../decisions/0052-read-your-writes-commit-versions.md)) | `pg_last_wal_replay_lsn()` on the replica keeps pace with the primary and query load moves back to it |
| Runners refuse edge assertions | read the refusal code: `ActorUnavailable` means the runner cannot reread its key set, so restore the key-set URL; `invalid_credentials` means check the runner's `issuer`, `audience`, and `region` and that the edge's signing `kid` is published | a request through the edge is admitted and the key set is reread within `refreshEvery` |
| Payload startup refusal | run `akter payloads check`; restore the dropped chain step or wait for `akter payloads clear` ([ADR 0032](../decisions/0032-event-and-effect-payload-evolution.md)) | the check passes and every retained version decodes |
| Content sweep lag or grant-key rotation | inspect `durable.contents` and sweep lag; rotate grant keys with one grant lifetime of overlap ([ADR 0034](../decisions/0034-tenant-scoped-content-addressed-blobs.md)) | unreferenced content older than grant plus grace is gone; no referenced content is |
| `DataDirLocked` or `DataDirVersion` | find the process holding the `dataDir` lock, or dump and reload with the previous core version ([ADR 0035](../decisions/0035-pglite-embedded-production-backend.md)) | one process opens the `dataDir`; the next start after a SIGKILL recovers the last commit |

## Drain a runner

A drain does not hand the runner's shards to the others; its process exit does. Follow this sequence for a rolling deploy, a scale-down, or a runner you suspect:

1. Call `RuntimeControl.drain({ deadline })` with a deadline you choose; there is no default. `GET /ready` answers `503 draining` at once, so the load balancer stops routing to the runner.
2. Wait for the report. On `deadline-expired`, record `interruptedTurns` and `interruptedJobs` from the report or the `Drain deadline expired` warning (see the table above).
3. Exit the process gracefully right away, by closing its layer. The exit ends the runner's activations and releases its shard locks, so the other runners take its actors at once. Until then the drained runner keeps its shards and refuses their commands, which callers retry.
4. Confirm that `GET /ready` answers `200` on the remaining runners and that commands for the drained runner's actors succeed.

Do not leave a drained runner running, and do not restart it because it answers `503`. If the process is killed instead of exiting, the other runners take its shards only after its locks expire. A drain of one runner is not the deployment-wide quiescence that [restore](04-backup-restore.md) needs.

Never delete receipts, generation rows, messages, outbox rows, or dead letters as a first-line fix. Escalate with trace ids, command ids, SQL evidence, deployed versions, and the exact recovery test.

## Fleet views

* **A view is `stale` with `last_error`.** A recompute failed deterministically, for example a `sum` past the `bigint` range. Its other groups stopped advancing and every other view continues. Fix the data or the definition, then run `akter fleet rebuild <View> --database-url <url>`: the view goes back to `building` and the maintainer rebuilds it from its source while writes continue.
* **Every view is `stale` without an error.** The slot was dropped or invalidated (`wal_status = 'lost'` in `pg_replication_slots`, usually after `max_slot_wal_keep_size`). Run `akter fleet setup` again: it recreates the slot, and the maintainer rebuilds every view.
* **Retained WAL grows.** No runner holds `akter/fleet`, or the maintainer cannot keep up. Check `pg_replication_slots` for `durable_fleet` and `pg_locks` for the advisory lock, and alert on `pg_current_wal_lsn() - confirmed_flush_lsn` of the slot. Dropping the slot releases the WAL and marks every view stale; `akter fleet setup` then rebuilds them.
* **A definition changed.** Startup marks the view `stale` and the maintainer rebuilds it; a newly registered view starts `building` and builds from the existing rows.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.