Skip to content

Health

Liveness, readiness, and startup probes — three routes mounted by importing one module, extended with indicators on any provider.

Health probes answer the questions an orchestrator asks: should you keep this pod running? should you route traffic to it? has it finished booting? The probes ship as routes on the HTTP transport. Importing HealthModule mounts three — GET /health/live, GET /health/ready, GET /health/startup — each returning 200 and a JSON body when every indicator passes, 503 with the same shape and error populated when any indicator fails. Routing rides on poem (the same engine HttpModule mounts), and discovery uses inventory gated by the access graph.

Terminal window
cargo add nest-rs --features health
cargo add anyhow

Every #[readiness] / #[liveness] / #[startup] method returns anyhow::Result<()>, so your own indicator names the crate. Reach it through nest_rs::core::anyhow instead and the second line goes too.

apps/api/src/module.rs (from the demo)
use nest_rs::core::module;
use nest_rs::health::HealthModule;
#[module(imports = [HealthModule])]
pub struct ApiModule;

Three routes appear at boot:

Terminal window
$ curl -s http://localhost:3000/health/live
{"status":"up","info":{},"error":{},"details":{}}
$ curl -s http://localhost:3000/health/ready
{"status":"up","info":{},"error":{},"details":{}}
$ curl -s http://localhost:3000/health/startup
{"status":"up","info":{},"error":{},"details":{}}

With no indicators registered every probe reports up with empty buckets — the right answer for a first boot, since there is nothing yet to be unhealthy. Register indicators to make the probes speak for your dependencies: see Indicators for the framework-shipped ones and how to write your own.

ProbeQuestion the orchestrator asksDefault with no indicators
liveIs the process still alive — should I restart it?up
readyShould I send traffic to it right now?up
startupHas it finished its long boot?up

The split mirrors the orchestrator’s: a long-booting app fails startup until ready (so Kubernetes does not kill it for slow liveness during boot); a temporarily overloaded app fails ready (so traffic drains) without flagging live (the process is fine — do not restart it).

Every probe returns the same body. info holds the up indicators, error holds the down ones, details carries the union so an operator can grep one bucket without iterating a list.

Terminal window
$ curl -s http://localhost:3000/health/ready | jq
{
"status": "up",
"info": {
"db": { "name": "db", "status": "up" },
"upstream": { "name": "upstream", "status": "up" }
},
"error": {},
"details": {
"db": { "name": "db", "status": "up" },
"upstream": { "name": "upstream", "status": "up" }
}
}

When one indicator drops, the response moves to 503 and the indicator migrates from info to error with a fixed reason — "check failed", "timed out", or "probe deadline exceeded". Never the indicator’s own error: the probe is routinely unauthenticated and a connection error carries a DSN or an internal hostname. The real cause goes to the log, warn on nest_rs::health (see Indicators).

Terminal window
$ curl -s -o /dev/null -w "%{http_code}\n" http://localhost:3000/health/ready
503
$ curl -s http://localhost:3000/health/ready | jq
{
"status": "down",
"info": { "db": { "name": "db", "status": "up" } },
"error": {
"upstream": {
"name": "upstream",
"status": "down",
"error": "check failed"
}
},
"details": { "db": { /* ... */ }, "upstream": { /* ... */ } }
}

The overall status is down if any indicator is down. The HTTP status maps from that one bit: 200 for up, 503 for down.

Two ceilings bound a probe response, read from NESTRS_HEALTH__* or pinned through HealthModule::for_root — dual-path, like every other module’s config. The kubelet’s timeoutSeconds defaults to 1, and a probe that has not answered by then is scored a failure; on a liveness probe, that is a restart.

KeyMeaning
INDICATOR_TIMEOUT_MSceiling on one indicator; expiry reports it down and names it in a warn — default 750
PROBE_DEADLINE_MSceiling on the whole response, whatever the indicator count; what answered is reported, the rest are down — default 900

Milliseconds, not seconds: the only interval that matters here is one inside the orchestrator’s second. 0 is refused at boot on both — an unbounded probe is what these fields exist to prevent, and a zero one fails every poll.

Indicators run concurrently, so a probe costs the slowest check rather than their sum, and adding a fifth indicator does not move the response’s worst case.

apps/api/src/module.rs (from the demo)
use std::time::Duration;
use nest_rs::core::module;
use nest_rs::health::{HealthConfig, HealthModule};
#[module(imports = [
HealthModule::for_root(
HealthConfig::default().with_probe_deadline(Duration::from_millis(400)),
),
])]
pub struct ApiModule;
  • Indicators — where they live, the framework-shipped ones, writing a custom one with #[indicators].
  • Discovery — how module-gating decides which indicators run in a given binary.
  • Database / Health — SeaOrmHealthModule, the first framework-shipped indicator.
  • HTTP — the transport HealthModule mounts on; same #[controller] shape under the hood.