> ## Documentation Index
> Fetch the complete documentation index at: https://docs.akter.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Observability

> Inspect runner metrics, traces, durable work, and defects.

# Observability

**Responsibility:** make behavior explainable in production.\
**Authority:** operational.\
**Owner role:** operations/reliability.
**Change policy:** a change requires operator review when a procedure or limit changes.

[ADR 0049](../decisions/0049-observability-names-metrics-and-defect-spans.md) fixes the span and metric names; they are public, and a rename is a breaking change (invariant O1).

**Spans.** A command's admission is `akter.admission`, its turn on the owner `akter.<Actor>/<Command>`, and the turn's commit group `akter.commit`. Commands that were already waiting run as one turn batch under `akter.<Actor>/batch`, which links each command's request span and carries `batch.size`; a lone command keeps its own turn span, and every command counts once in `akter.turns`. The relay's deliveries are `akter.relay.intent` and `akter.relay.subscription`, and each executor attempt is `akter.job/<Actor>/<Job>`. Admission, turn, and commit share one trace; a durable hop (outbox row, subscription row, workflow resume) starts a new trace on the relay, correlated by the command id or job id. Turn spans carry tenant, actor type and id, command and command id, caller kind, trigger, generation, outcome, and replay status; logs carry the same identities. No span carries a principal's subject, a credential, a payload, or state. A deterministic defect fails its turn span with the cause. Install any Effect tracer exporter, such as `OtlpTracer.layer`, around `Actors.layer` to ship them. The hosted API and edge install `OtlpTracer` and `OtlpLogger` from `OTEL_*` variables and strip URL paths, queries, headers and client addresses from their spans. The command id is also the HTTP `x-request-id`, linking client errors to receipts and server logs. Redact bearer tokens, API keys, database URLs, private payloads, and provider credentials.

**Metrics.** Each runner counts turns by outcome, turn duration, mailbox age, turn-pool waits, resident and started activations, receipts written, replayed, and pruned, events appended and pruned, outbox rows staged, relay deliveries and retries by kind, dead letters, and subscription gaps without a recipient. The runner that maintains fleet views reports, by `view`, `akter.fleet.lag_bytes` (WAL between the server's flush position and what it applied) and `akter.fleet.lag_ms` (time since it last drained the change feed), and counts `akter.fleet.groups_recomputed` by `view` and `akter.fleet.batches`; alert on lag, on a `stale` view in `actor_fleet_views`, and on WAL the `durable_fleet` slot retains. One runner of the deployment samples the database every `observability.sampleEvery` (15 seconds): outbox rows and relay lag by kind, rows claimed at least 8 times (`akter.relay.stuck_rows`), subscription lag in events and milliseconds, and events pinned by subscriptions and open workflows. Attributes name actor types, commands, jobs, subscriptions, kinds, and outcomes, never a tenant or id. `Telemetry.serve()` from `@rikalabs/akter/runtime` serves `GET /metrics` in Prometheus text on a private listener; an application that installs `OtlpMetrics` exports the same series over OTLP. Sum counters across runners; take the `max` of sampled gauges, which appear only on the runner holding the sampler.

**Defects.** Each runner keeps its last `observability.defects` (1,000) defect turn spans in memory. `Operators.serve` serves them at `GET /operator/defects` under the `defects.read` capability ([ADR 0050](../decisions/0050-operator-authority-and-audited-repair.md)). `akter defects list --url <runner> … [--tenant T] [--actor T] [--since 1h]` reads them from every runner named, with the operator token from `AKTER_OPERATOR_TOKEN`, and prints each with its trace id; older history is in the exporter's backend.

Monitor command latency, mailbox depth and age, receipt replay rate, generation-fence failures, lock timeouts, redelivery, transaction retries, actor restarts, singleton ownership, cron lateness, workflow age, job retries and dead letters, parked connections, database saturation, and restore progress. On hosted Neki also monitor relay lag, duplicate suppression, and stranded `actor_outbox` rows.

Bound actor-id and tenant cardinality in metrics; use traces and logs for individual identities. Alert on degraded recovery paths, not only failed requests: relay backlog, old dead letters, repeated deterministic defects, and a singleton without an owner are operational failures.

Committed runtime state is readable from SQL through the read-only `durable` [inspection views](inspection-views.md) (migration `0013_inspection_views`, [ADR 0028](../decisions/0028-sql-inspection-views.md)): actors, state, receipts, events, outbox, timers, jobs, and dead letters. Dashboards and tools read those views under a role granted only the `durable` schema, never the `actor_*` tables. `akter dev` serves a local inspector over the same views, one tenant at a time ([the local inspector](inspection-views.md#the-local-inspector)). The planned `akter dead-letters` command will expose inspection and repair under the job policy and [runbooks](runbooks.md); the command is not implemented. Never infer success solely from an executor attempt; use the durable job outcome. Parked connections, restore progress, database saturation, singleton ownership, cron lateness, and workflow age are not reported yet.

**Delivery timeouts.** When a command waits out its actor type's `deliveryTimeout`, the admitting runner logs `Delivery timed out` at warning level. The log carries the command's identity and this runner's view at that moment: the target activation's handlers, whether Cluster is rebuilding one, its worker phase (`idle`, `turn`, `workflow kick`, `restarting`, or `none` when the activation is not on this runner), and the commands in its mailbox and current batch. It also carries the turn pool's leased and waiting sessions and every other activation on the runner that is being rebuilt or restarted in place, so one line shows whether a stall sat in the target's worker, the pool, or a neighbour's restart.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.