Deploy
Responsibility: take an app from the quickstart to a production process on Postgres, and state what is and is not supported today.Authority: operational.
Owner role: operations/platform.
Change policy: change with the support matrix and deployment when a supported shape or limit changes; those documents win on any conflict.
What is supported today
Akter is alpha. Before you deploy, know the limits the support matrix records:- Postgres only. PGlite is for development and tests, one process per data directory. Production PGlite is not supported.
- Single runner or a TCP runner cluster.
Runner.socketis the public Postgres configuration for separate processes. Runners authenticate each other with mutual TLS (Runner.mtls). Three-process command, relay, singleton, schedule, SIGKILL, and rolling-drain drills run on one host; separate-host networks and hosting providers still require their own evidence. - Embedded or served.
Actors.serveserves commands, reducers, and queries over HTTP, connections over WebSocket, and feeds, streams, and watches over SSE. These transports are verified behind Bun’s HTTP server on loopback; no proxy, load balancer, or hosting provider has been verified. - No managed hosting. Hosted runners are planned.
- Not on npm yet. Install
@rikalabs/akterfrom a locally packed tarball, as in the quickstart, until the first alpha release.
The process
A deployed app is one Bun process that builds the runtime over Postgres and, if clients call it over HTTP, serves it. This is thechat example’s src/main.ts:
routes is Actors.serve({ actors: [Room], auth, openapi: { path: "/openapi.json" } }) from @rikalabs/akter/runtime. An embedded app, like the quickstart’s counter, skips Actors.serve and calls its actors as Effects inside the same process.
The database
- Framework tables are created and migrated when the runtime starts; see migrations for how startup refuses a database it cannot migrate safely.
- Owned tables are yours. Create and migrate them with drizzle-kit before the runtime starts; startup fails if a declared table is missing or claimed by another actor type.
- State upcasts through the migrations declared on
Actor.state, one actor at a time, inside turns. No batch migration is needed.
Redacted.
Several runners
Every process builds the same actor layers and shares the database, but advertises its own directly reachable private address. Runners can all start at once against an empty database: startup serializes creation of the migration bookkeeping and cluster tables. ProvideRunner.socket to Actors.layer with Runner.mtls as its transport:
BunCrypto.layer as in the single-process example. The listener is separate from Actors.serve and uses NDJSON RPC over TLS 1.3, not HTTP. Each runner’s certificate has exactly one subject alternative name, the URI Runner.identity("shop-production") (spiffe://akter/deployment/shop-production), and must chain to an authority in ca. A certificate with any other name besides it is refused. Session resumption is off, so every connection runs a full handshake. Both sides refuse a peer without a certificate, a plaintext peer, an untrusted or expired certificate, and a runner of another deployment before any runner message is exchanged. Any authority that can issue the URI SAN works (cert-manager, SPIRE, step-ca, AWS Private CA); RunnerAuthority.make() and authority.issue({ deployment }) create development credentials.
credentials runs at startup and again every refreshEvery (default one minute). Replace the files to rotate without a restart: new connections use the new certificate, and established ones continue. To rotate the authority, first add the new authority to every runner’s ca and wait a refresh, then replace each certificate and key with ones from the new authority, then remove the old authority. Credentials that fail validation on a refresh, or take longer than refreshTimeout (default 10 seconds) to load, are logged and the previous ones stay in use; at startup they stop the runner. Readiness reports peering once the current certificate has expired or loads have failed for unhealthyAfter (default 5 minutes), so a runner that can no longer peer leaves ingress. A peer’s certificate is checked at the handshake only, so an expired or distrusted peer keeps an established session until either side reconnects.
The platform plaintext layers (Layer.merge(layerSocketServer, layerClientProtocol) from @effect/platform-bun/BunClusterSocket) remain available for a network isolated to one deployment’s runners. They authenticate no one: any process that reaches the port can deliver runner messages.
address advertises a unique private DNS name or IP reachable directly by every peer, never a wildcard or a shared load-balancer address. listenAddress binds an interface and can differ; bind the private interface, and never expose the port to public clients, even over mutual TLS. Never run two incarnations at one address; keep it reserved until the old process exits and releases its locks. Prove your network’s reachability and failure behavior before deploying across hosts.
Defaults are 256 shards per group, one-second assignment refresh, and expiring table locks with the singleton lease check. Survivors wait for the dead owner’s expiration, and stale generations still fail the database fence. Advisory multi-runner locks are unsupported: the pinned storage assigns colliding lock IDs to distinct private holder groups (ADR 0068). Expiration defaults to 35 seconds and must be at least 3 seconds; Cluster caps lock refresh at a third of it. Polling, activation, and retries add time beyond expiration. entityTerminationTimeout defaults to 15 seconds and bounds activation shutdown during handoff.
Migration 0029_runner_configuration records the first public runner’s shard count and expiration. Later runners must match; an embedded runtime cannot join without Runner.socket. Stop the old embedded process before first enabling this configuration. To change a recorded layout, stop every runner, clear stale cluster registrations and locks, update the recorded configuration in maintenance, then restart all runners with matching values; never change layout during rolling deployment. Actor identities and receipts do not change.
Gate traffic on RuntimeControl.readiness or served GET /ready, not on a listening port. A configured runner reports routing until registered and holding every currently assigned shard, including its private holder group. This uses a fresh registration snapshot and local acquired shards, not actor activation. Assignments can change after a probe, so commands still retry normally.
Sizing
Each runner owns its pools. Keepprocesses × (maxConnections + offTurnConnections + queryConnections), plus the coordination pool’s connections when configured, plus operator headroom below the database’s max_connections.
- Connections. Each command holds one Postgres connection until its transaction ends; a pipelined chain keeps its session until the chain ends.
Database.postgres({ maxConnections })defaults to 50,offTurnConnectionsandqueryConnectionsto 10 each; count every pool targeting the server, plus migrations, backups, and operator sessions, under itsmax_connections. A pooler in front of Postgres is unverified. - Memory. A resident actor holds about 20 KiB of JavaScript heap on the measured runner, so the default
maxResidentActorsof 10,000 is about 200 MiB. A command that needs a new activation past the limit failsRunnerAtCapacity, and its handle retries until an idle actor hibernates.
Serving over HTTP
- TLS. Serve
Actors.servebehind TLS. Credentials andIdempotency-Keytravel in headers, and the framework cannot tell whether a proxy terminates TLS, so it does not refuse plain HTTP. - Authentication.
Actors.serverequires an auth provider. UseAuth.jwt({ issuer, audience, jwks, tenant })from@rikalabs/akter/runtimefor tokens from an identity provider,Auth.makefor your own, andAuth.noneonly for deliberately public actors. Authorization is still yourauthorizecallback. - Browsers. List allowed browser origins in
origins. Behind a proxy, list the public origin: forwarding headers are not trusted. - Clients. The OpenAPI document at
openapi.pathgenerates clients in any language; see generating clients. Every client must keep oneIdempotency-Keyacross its retries of a command. - Retry window.
Actors.serverefuses a runtime whose retry window is below 60 seconds. The default is one day.
Releasing a new version
With one process, stop the old process and start the new one. WithRunner.socket, start a compatible replacement, wait for readiness, remove the old runner from external traffic, call its RuntimeControl.drain({ deadline: "30 seconds" }), then close its runtime layer and stop it. Repeat one runner at a time. Drain makes it unready and stops claims; closing the layer releases shards. Retry interrupted commands under their original ids. The drills prove same-code rolling restart, not arbitrary mixed-version compatibility. Never remove decoders, workflow steps, or schemas still needed by durable work or surviving runners.
A command the old process had not committed is not applied. Retrying the same command id gives exactly one result: handles retry until deliveryTimeout, and HTTP clients resend the same Idempotency-Key. Pending work stays in the outbox for survivors or replacements. Zero loss through primary failover requires a synchronously replicated standby; asynchronous promotion can lose acknowledged commits regardless of runner configuration.
Before you release:
- Apply owned-table migrations that the old and new code both accept (expand before contract).
- If you changed a workflow, check that no open execution needs a removed or renamed step. Startup refuses such a deploy.
akter workflows check --entry <module> --database-url <url>, fromapps/cliin the repository (not yet published), runs the same check first. - Keep every event, job, and workflow payload decodable until the records that use it have passed their retention.
Checklist
- Postgres, a database for this app alone, and
DATABASE_URLin a secret. - Owned-table migrations applied before start.
authorizeallows only the callers and tenants you expect, andActors.servehas a real auth provider.- TLS in front of the HTTP server, and
originsset for browser clients. maxConnections × processeswithin the server’smax_connections.- One process per runner address, matching shard count and expiration across the database.
- A private peer network, direct advertisement, readiness gating, and drain before closing old runners.
- Logs collected: deterministic defects, dead-lettered jobs, and outbox retries are logged as warnings and errors. See observability.