The core API exposes three operational endpoints outside the /v1 router. They are mounted in backend/apps/core-api/src/app.ts from @perform/observability, before the rate limiters, so they are never rate limited and need no authentication. Requests to /healthz, /readyz and /metrics are excluded from the HTTP access log.

/healthz

Returns a constant body and touches no dependency:
Use it for process supervision and for the deploy health check. deploy/dev/health-check.sh polls it after each pm2 release, and the workspace scripts tools/wait-for-api.sh and tools/ensure-core-api.sh poll it before starting the web app. A passing /healthz only proves the HTTP server is listening. It does not prove the database is reachable.

/readyz

Runs every registered check in parallel. The checks are built by buildReadinessChecks in apps/core-api/src/health-checks.ts. A check that throws counts as failed. The response depends on which checks fail:
Status 200.
degraded with redis is not harmless. Without Redis, BullMQ queues and schedulers stop, and the replay nonce store behind the signed lanes cannot work. Treat a degraded response as an incident even though the status code is 200. Alert on the body, not only on the status.

Adding a check

Add an entry to the array in health-checks.ts:
Keep checks cheap. They run on every probe. Mark a dependency optional only when the API can serve most requests without it.

/metrics

@perform/observability creates a prom-client Registry and calls collectDefaultMetrics on it. The endpoint returns the registry in Prometheus text format with the registry’s content type. What you get today is the default Node.js process metrics from prom-client: CPU time, resident memory, heap sizes, event loop lag, garbage collection durations, active handles and so on. There are no HTTP request counters, queue depth gauges or business metrics registered at the time of writing. The registry is exported as metricsRegistry. To add a metric, register it against that registry:
prom-client is a dependency of @perform/observability, not of the API. Either define metrics inside that package and export them, or add prom-client to the workspace that defines the metric.
/metrics has no authentication. Do not expose it publicly. Block the path at the reverse proxy and scrape it from the host or a private network.
initTelemetry(serviceName) is called first thing in server.ts but is currently an empty function. No tracing is set up.

Request ids

requestIdMiddleware runs early in the stack. It takes the incoming x-request-id header when present, or generates a ULID, stores it on the request and returns it in the X-Request-Id response header. The same id appears as requestId in the response envelope and in the access log line. Quote it when you report or investigate a failed request. The web app’s server-side client does not send an x-request-id, so each API call gets a fresh id. Read it from the X-Request-Id response header or the envelope.

Web app health

The web app has no dedicated health route. Its deploy health check requests / on port 3043 and accepts any successful or redirect status. To check the whole path from the web server to the API, request /api/health on the web origin. The /api proxy forwards it to the Hono app in the backend, which answers OK. When the API is down, the proxy answers 503 with:

Probing from a shell

Process-level safety nets

  • server.ts logs unhandled rejection and keeps running. On uncaught exception it logs and exits with code 1, and the supervisor restarts it.
  • Under pm2, max_memory_restart restarts the API when its memory passes 1024 MB.
  • A config validation failure at boot exits with code 78 and prints each invalid key. A process that restarts in a loop right after a deploy is often this. Check the first lines of the error log.