Observability
Responsibility: make behavior explainable in production.Authority: operational.
Owner role: operations/reliability. Change policy: a change requires operator review when a procedure or limit changes. ADR 0049 fixes the span and metric names; they are public, and a rename is a breaking change (invariant O1). Spans. A command’s admission is
akter.admission, its turn on the owner akter.<Actor>/<Command>, and the turn’s commit group akter.commit. Commands that were already waiting run as one turn batch under akter.<Actor>/batch, which links each command’s request span and carries batch.size; a lone command keeps its own turn span, and every command counts once in akter.turns. The relay’s deliveries are akter.relay.intent and akter.relay.subscription, and each executor attempt is akter.job/<Actor>/<Job>. Admission, turn, and commit share one trace; a durable hop (outbox row, subscription row, workflow resume) starts a new trace on the relay, correlated by the command id or job id. Turn spans carry tenant, actor type and id, command and command id, caller kind, trigger, generation, outcome, and replay status; logs carry the same identities. No span carries a principal’s subject, a credential, a payload, or state. A deterministic defect fails its turn span with the cause. Install any Effect tracer exporter, such as OtlpTracer.layer, around Actors.layer to ship them. The hosted API and edge install OtlpTracer and OtlpLogger from OTEL_* variables and strip URL paths, queries, headers and client addresses from their spans. The command id is also the HTTP x-request-id, linking client errors to receipts and server logs. Redact bearer tokens, API keys, database URLs, private payloads, and provider credentials.
Metrics. Each runner counts turns by outcome, turn duration, mailbox age, turn-pool waits, resident and started activations, receipts written, replayed, and pruned, events appended and pruned, outbox rows staged, relay deliveries and retries by kind, dead letters, and subscription gaps without a recipient. The runner that maintains fleet views reports, by view, akter.fleet.lag_bytes (WAL between the server’s flush position and what it applied) and akter.fleet.lag_ms (time since it last drained the change feed), and counts akter.fleet.groups_recomputed by view and akter.fleet.batches; alert on lag, on a stale view in actor_fleet_views, and on WAL the durable_fleet slot retains. One runner of the deployment samples the database every observability.sampleEvery (15 seconds): outbox rows and relay lag by kind, rows claimed at least 8 times (akter.relay.stuck_rows), subscription lag in events and milliseconds, and events pinned by subscriptions and open workflows. Attributes name actor types, commands, jobs, subscriptions, kinds, and outcomes, never a tenant or id. Telemetry.serve() from @rikalabs/akter/runtime serves GET /metrics in Prometheus text on a private listener; an application that installs OtlpMetrics exports the same series over OTLP. Sum counters across runners; take the max of sampled gauges, which appear only on the runner holding the sampler.
Defects. Each runner keeps its last observability.defects (1,000) defect turn spans in memory. Operators.serve serves them at GET /operator/defects under the defects.read capability (ADR 0050). akter defects list --url <runner> … [--tenant T] [--actor T] [--since 1h] reads them from every runner named, with the operator token from AKTER_OPERATOR_TOKEN, and prints each with its trace id; older history is in the exporter’s backend.
Monitor command latency, mailbox depth and age, receipt replay rate, generation-fence failures, lock timeouts, redelivery, transaction retries, actor restarts, singleton ownership, cron lateness, workflow age, job retries and dead letters, parked connections, database saturation, and restore progress. On hosted Neki also monitor relay lag, duplicate suppression, and stranded actor_outbox rows.
Bound actor-id and tenant cardinality in metrics; use traces and logs for individual identities. Alert on degraded recovery paths, not only failed requests: relay backlog, old dead letters, repeated deterministic defects, and a singleton without an owner are operational failures.
Committed runtime state is readable from SQL through the read-only durable inspection views (migration 0013_inspection_views, ADR 0028): actors, state, receipts, events, outbox, timers, jobs, and dead letters. Dashboards and tools read those views under a role granted only the durable schema, never the actor_* tables. akter dev serves a local inspector over the same views, one tenant at a time (the local inspector). The planned akter dead-letters command will expose inspection and repair under the job policy and runbooks; the command is not implemented. Never infer success solely from an executor attempt; use the durable job outcome. Parked connections, restore progress, database saturation, singleton ownership, cron lateness, and workflow age are not reported yet.
Delivery timeouts. When a command waits out its actor type’s deliveryTimeout, the admitting runner logs Delivery timed out at warning level. The log carries the command’s identity and this runner’s view at that moment: the target activation’s handlers, whether Cluster is rebuilding one, its worker phase (idle, turn, workflow kick, restarting, or none when the activation is not on this runner), and the commands in its mailbox and current batch. It also carries the turn pool’s leased and waiting sessions and every other activation on the runner that is being rebuilt or restarted in place, so one line shows whether a stall sat in the target’s worker, the pool, or a neighbour’s restart.