Skip to content

Health Endpoints

Both the issuer and the verifier expose two probes, named after the Kubernetes health check endpoints. They answer two different questions, and conflating them is what causes a healthy service to be restarted in the middle of a database blip.

Path Question Answers
/livez Is the process running? 204 whenever the service is serving
/readyz Can it answer requests? 204 when the database answers, 503 when it does not

Both are served on the same port as the rest of the service's API. Neither returns a body, and neither takes an API key — a probe has no credentials of yours. The administrative endpoints they sit alongside do; see API keys.

Liveness

Whatever polls /livez is expected to restart the service when it fails, so it answers for the process and nothing else. It stays 204 while the database is unreachable, while a callback endpoint is down, and while anything else the service depends on is broken.

Readiness

A 204 from /readyz means the database is reachable and its schema readable, not merely that a connection could be opened. Anything else is a 503, with the reason written to the service's log rather than to the response.

Use it to decide whether to send traffic: take the instance out of rotation while it fails, and put it back when it recovers. It requires no restart.

Readiness reports on one connection

The probe borrows a single connection from the pool, and connections are not validated when they are handed out. While a service recovers from a database restart the answer can therefore flip between 204 and 503 until the stale connections have been discarded. Treat a single failure as a signal to stop sending traffic, not as proof that every connection is dead.

Wiring the probes

In Compose, give each backend a health check of its own rather than relying on start order:

verifier-backend:
    healthcheck:
        test:
            [
                'CMD-SHELL',
                'curl -sf -o /dev/null http://localhost:<port>/readyz',
            ]
        interval: 10s
        timeout: 5s
        retries: 6
        start_period: 30s

curl -sf treats 204 as success and fails on 503. With this in place docker compose up --wait blocks until both backends can actually serve, instead of returning as soon as their containers start.

In Kubernetes, point the two probe types at the two endpoints:

livenessProbe:
    httpGet: { path: /livez, port: <port> }
    periodSeconds: 10
readinessProbe:
    httpGet: { path: /readyz, port: <port> }
    periodSeconds: 10

What's Next

  • Database: What readiness is checking, and how the schema is managed.
  • Troubleshooting: Startup and connection failures.