Skip to content

Retries and failure

What an Err from a #[process] method means — the retry shape, panic handling, and the dead-letter story.

The contract is small: a #[process] method returns anyhow::Result<()>, Err(_) marks the attempt failed, and Ok(()) acknowledges it — once the attempt’s transaction has settled, which is the one way a body that returned Ok can still fail (see below). After that the backend decides — and Redis gives you a fixed retry budget per method, a panic safety net, and a dead-jobs list. Anything richer than that is bolted on top.

crates/features/src/audio/queue/processor.rs
#[process(queue = AudioQueue, retries = 3)]
async fn transcode(&self, job: TranscodeCommand) -> Result<()> {
self.svc.transcode(&job.file).await
}
  • Ok(()) ⇒ the attempt’s transaction commits and the job is removed from Redis; that’s the end of its life. If the commit cannot be honoured, nothing it wrote landed and the attempt fails instead — the one case where Ok is not the end.
  • Err(e) ⇒ the method runs again, at once, with the same job_id and the next attempt. After retries re-runs have failed, the job moves to the dead list (see Where failed jobs go).
  • A panic inside the method is caught, logged as job dead-lettered: handler panicked at error (inside the job’s own span, so it carries queue, processor, job_id and attempt), and the job dead-letters immediately. The worker keeps running — one bad job doesn’t kill the queue’s consumer.

The panic is caught one level in, by the queue port itself (nest_rs_queue::consume::attempt) — that is what keeps the structured event inside the per-job span instead of losing it to an unwind. The job runtime catches an unwind as well, as the backstop for a panic outside the handler, so a single unwrap() in user code never tears the worker down.

Rust’s default panic hook still prints its own thread … panicked at … line to stderr first: catching an unwind does not replace the hook. That line carries no target, so RUST_LOG cannot filter it — the framework’s own record is the job dead-lettered: handler panicked event beneath it. Install std::panic::set_hook in main if you need the raw line gone.

retries = N is a fixed retry count, no backoff: a failed attempt runs again at once, up to N times, while the queue’s one job slot stays held. The number of total attempts a job sees is N + 1 (the first try plus N retries). Default if you omit retries: 0 — the job is tried once.

If you need a different retry curve, today the path is to wrap the business call in the service:

crates/features/src/audio/service.rs
impl AudioService {
pub async fn transcode(&self, file: &str) -> Result<()> {
let mut attempt = 0;
loop {
match self.do_transcode(file).await {
Ok(()) => return Ok(()),
Err(e) if attempt < 5 => {
let delay = Duration::from_millis(100 * 2u64.pow(attempt));
tracing::warn!(target: "features::audio", %file, attempt, error = %e, "transcode failed; retrying");
tokio::time::sleep(delay).await;
attempt += 1;
}
Err(e) => return Err(e),
}
}
}
}

The #[process] budget then acts as the outer safety net — a couple of immediate retries after the service-level backoff already gave up, useful for transient framework-level issues (the container resolution failing under load, a panic in deserialization).

A job whose budget is exhausted, or that dead-lettered, moves to one Redis list shared by every queue, nestrs:queue:dead. Each entry is the job’s JSON record — its queue, its payload, and the error under meta.error — and the newest 1,000 are kept:

Terminal window
$ redis-cli LRANGE nestrs:queue:dead 0 9

The framework does not ship a dead-letter routing layer of its own — the dead list is the storage’s, and inspecting / replaying it is an operational task you run with redis-cli or your own admin handler.

If you need richer dead-lettering — alerting, replay, archival to a separate queue — write a wrapper service that catches the Err from the work and republishes the payload to a dedicated queue (e.g. audio.dead) before returning the error. A #[processor] on audio.dead then runs your DLQ logic. That’s the same pattern any non-Redis backend would expose.

crates/features/src/audio/queue/processor.rs
#[process(queue = AudioQueue, retries = 3)]
async fn transcode(&self, job: TranscodeCommand) -> Result<()> {
match self.svc.transcode(&job.file).await {
Ok(()) => Ok(()),
Err(e) => {
self.svc.report_dead(&job, &e).await?;
Err(e)
}
}
}

Panics, deserialization errors, container misses

Section titled “Panics, deserialization errors, container misses”

Four failure modes the framework converts into the same Err path, so you don’t have to think about them separately:

  • Panics in the method body — caught, with a job dead-lettered: handler panicked event carrying the panic message. The job dead-letters immediately (regardless of remaining retries), the worker stays up.
  • Deserialization errors — the macro-emitted JobHandler does the serde_json::from_value. A payload that doesn’t match the method’s job type returns Err before your method even runs. Retrying won’t help: the payload shape is wrong, so it is classified non-retryable and dead-letters on the first attempt, logging job dead-lettered: non-retryable failure at error. The retries budget is not consumed.
  • Container resolution failures — Container::get of the #[processor] host can fail if a dep wasn’t seeded. The handler returns Err with the resolution error. This usually means a configuration mistake at boot, not a transient — investigate, don’t rely on retries.
  • A transaction that could not be settled — the body returned Ok, and the attempt’s transaction could not be committed, so nothing it wrote landed. The attempt fails rather than reporting writes that were never made, and the database decides whether it is retried — with the answer depending on where it failed. A statement inside the attempt failed on a serialization conflict, a deadlock, a pool that timed out, or a connection the server closed: retryable, because nothing is durable before COMMIT and the framework therefore knows the attempt wrote nothing. A constraint checked at COMMIT, or a commit whose outcome is unknown because the connection dropped mid-flight, is not — it dead-letters at once, because that one may have landed and replaying it would write twice. Retrying replays the whole body, side effects included, so spending the budget on a failure that repeats identically costs several times over and dead-letters anyway. See Repo and executor.

A processor reads only payloads it understands. The wire envelope ({ "v": 1, "payload": ... }) carries a version; the macro-emitted handler validates it. A future release that bumps WIRE_FORMAT_VERSION to 2 makes a v: 1 worker refuse to decode v: 2 jobs — fails closed with a clear error rather than misinterpreting bytes. During a rolling deploy this means jobs queue up; afterward the new worker drains them.

Unversioned legacy payloads (a bare job JSON, no v key) are decoded directly with a tracing::warn — so jobs left in Redis from a release before envelopes existed don’t get stuck. New producers always wrap.

  • No backoff. retries = N is a fixed count, retried back-to-back — see the caution above.
  • will retry is logged on the terminal attempt too. Every failed attempt including the last logs job failed; will retry within the budget; after attempt N + 1 there is no budget left and the job dead-letters with no further line. Read the attempt field on the span, not the message, to tell the last try from the others.
  • The dead list is not routed anywhere. Nothing alerts, replays or archives on your behalf; wire the wrapper above if you need that.
  • A replica that dies mid-job has the job run again. Delivery is exclusive — two replicas never get the same job from the queue — and starting a replica leaves in-flight work alone. But a replica killed mid-job leaves the job in its own processing list, and once its heartbeat has been silent for five seconds a live replica puts the job back on the queue, where it runs again from attempt=1. Treat a #[process] handler as at-least-once and make it idempotent — the practical rule for any durable queue, and load-bearing here because replicas are how throughput scales. The scale-up half is measured in crates/nest-rs-redis/tests/e2e/replicas.rs.
  • A container restarted in place does not re-run its job. A replica is recognised by its hostname and pid, and a container restarted where it stood keeps both, so nothing sees its crashed predecessor as dead: the job that was running stays in the processing list the restarted process now owns, and does not run again. Replicas that differ by hostname are not affected.
  • Observability — what shows up in spans and logs when a job retries or fails.
  • OpenTelemetry — wiring the queue spans into a trace backend.
  • crates/nest-rs-redis/src/worker/consumer.rs — settle: how an attempt’s outcome spends, or does not spend, the retry budget.
  • crates/nest-rs-queue/src/inventory.rs — WIRE_FORMAT_VERSION, the envelope spec.