Retries and failure
What an Err from a #[process] method means — the retry shape, panic handling, and the dead-letter story.
The contract is small: a #[process] method returns anyhow::Result<()>,
Err(_) marks the attempt failed, and Ok(()) acknowledges it — once the
attempt’s transaction has settled, which is the one way a body that returned
Ok can still fail (see
below). After that the
backend decides — and Redis gives you a fixed retry budget per method, a
panic safety net, and a dead-jobs list. Anything richer than that is bolted
on top.
The basic shape
Section titled “The basic shape”#[process(queue = AudioQueue, retries = 3)]async fn transcode(&self, job: TranscodeCommand) -> Result<()> { self.svc.transcode(&job.file).await}Ok(())⇒ the attempt’s transaction commits and the job is removed from Redis; that’s the end of its life. If the commit cannot be honoured, nothing it wrote landed and the attempt fails instead — the one case whereOkis not the end.Err(e)⇒ the method runs again, at once, with the samejob_idand the nextattempt. Afterretriesre-runs have failed, the job moves to the dead list (see Where failed jobs go).- A panic inside the method is caught, logged as
job dead-lettered: handler panickedaterror(inside the job’s own span, so it carriesqueue,processor,job_idandattempt), and the job dead-letters immediately. The worker keeps running — one bad job doesn’t kill the queue’s consumer.
The panic is caught one level in, by the queue port itself
(nest_rs_queue::consume::attempt) — that is what keeps the structured
event inside the per-job span instead of losing it to an unwind. The job
runtime catches an unwind as well, as the backstop for a panic outside the
handler, so a single unwrap() in user code never tears the worker down.
Rust’s default panic hook still prints its own thread … panicked at … line
to stderr first: catching an unwind does not replace the hook. That line
carries no target, so RUST_LOG cannot filter it — the framework’s own record
is the job dead-lettered: handler panicked event beneath it. Install
std::panic::set_hook in main if you need the raw line gone.
The retry shape
Section titled “The retry shape”retries = N is a fixed retry count, no backoff: a failed attempt runs
again at once, up to N times, while the queue’s one job slot stays held.
The number of total attempts a job sees is N + 1 (the first try plus N
retries). Default if you omit retries: 0 — the job is tried once.
If you need a different retry curve, today the path is to wrap the business call in the service:
impl AudioService { pub async fn transcode(&self, file: &str) -> Result<()> { let mut attempt = 0; loop { match self.do_transcode(file).await { Ok(()) => return Ok(()), Err(e) if attempt < 5 => { let delay = Duration::from_millis(100 * 2u64.pow(attempt)); tracing::warn!(target: "features::audio", %file, attempt, error = %e, "transcode failed; retrying"); tokio::time::sleep(delay).await; attempt += 1; } Err(e) => return Err(e), } } }}The #[process] budget then acts as the outer safety net — a couple
of immediate retries after the service-level backoff already gave up,
useful for transient framework-level issues (the container resolution
failing under load, a panic in deserialization).
Where failed jobs go
Section titled “Where failed jobs go”A job whose budget is exhausted, or that dead-lettered, moves to one Redis
list shared by every queue, nestrs:queue:dead. Each entry is the job’s
JSON record — its queue, its payload, and the error under meta.error —
and the newest 1,000 are kept:
$ redis-cli LRANGE nestrs:queue:dead 0 9The framework does not ship a dead-letter routing layer of its own —
the dead list is the storage’s, and inspecting / replaying it is an
operational task you run with redis-cli or your own admin handler.
If you need richer dead-lettering — alerting, replay, archival to a
separate queue — write a wrapper service that catches the Err from
the work and republishes the payload to a dedicated queue (e.g.
audio.dead) before returning the error. A #[processor] on
audio.dead then runs your DLQ logic. That’s the same pattern any
non-Redis backend would expose.
#[process(queue = AudioQueue, retries = 3)]async fn transcode(&self, job: TranscodeCommand) -> Result<()> { match self.svc.transcode(&job.file).await { Ok(()) => Ok(()), Err(e) => { self.svc.report_dead(&job, &e).await?; Err(e) } }}Panics, deserialization errors, container misses
Section titled “Panics, deserialization errors, container misses”Four failure modes the framework converts into the same Err path,
so you don’t have to think about them separately:
- Panics in the method body — caught, with a
job dead-lettered: handler panickedevent carrying the panic message. The job dead-letters immediately (regardless of remaining retries), the worker stays up. - Deserialization errors — the macro-emitted
JobHandlerdoes theserde_json::from_value. A payload that doesn’t match the method’s job type returnsErrbefore your method even runs. Retrying won’t help: the payload shape is wrong, so it is classified non-retryable and dead-letters on the first attempt, loggingjob dead-lettered: non-retryable failureaterror. Theretriesbudget is not consumed. - Container resolution failures —
Container::getof the#[processor]host can fail if a dep wasn’t seeded. The handler returnsErrwith the resolution error. This usually means a configuration mistake at boot, not a transient — investigate, don’t rely on retries. - A transaction that could not be settled — the body returned
Ok, and the attempt’s transaction could not be committed, so nothing it wrote landed. The attempt fails rather than reporting writes that were never made, and the database decides whether it is retried — with the answer depending on where it failed. A statement inside the attempt failed on a serialization conflict, a deadlock, a pool that timed out, or a connection the server closed: retryable, because nothing is durable beforeCOMMITand the framework therefore knows the attempt wrote nothing. A constraint checked atCOMMIT, or a commit whose outcome is unknown because the connection dropped mid-flight, is not — it dead-letters at once, because that one may have landed and replaying it would write twice. Retrying replays the whole body, side effects included, so spending the budget on a failure that repeats identically costs several times over and dead-letters anyway. See Repo and executor.
Wire-format mismatch
Section titled “Wire-format mismatch”A processor reads only payloads it understands. The wire envelope
({ "v": 1, "payload": ... }) carries a version; the macro-emitted
handler validates it. A future release that bumps
WIRE_FORMAT_VERSION to 2 makes a v: 1 worker refuse to decode
v: 2 jobs — fails closed with a clear error rather than
misinterpreting bytes. During a rolling deploy this means jobs queue
up; afterward the new worker drains them.
Unversioned legacy payloads (a bare job JSON, no v key) are
decoded directly with a tracing::warn — so jobs left in Redis from a
release before envelopes existed don’t get stuck. New producers
always wrap.
Limits
Section titled “Limits”- No backoff.
retries = Nis a fixed count, retried back-to-back — see the caution above. will retryis logged on the terminal attempt too. Every failed attempt including the last logsjob failed; will retry within the budget; after attemptN + 1there is no budget left and the job dead-letters with no further line. Read theattemptfield on the span, not the message, to tell the last try from the others.- The dead list is not routed anywhere. Nothing alerts, replays or archives on your behalf; wire the wrapper above if you need that.
- A replica that dies mid-job has the job run again. Delivery is
exclusive — two replicas never get the same job from the queue — and
starting a replica leaves in-flight work alone. But a replica killed
mid-job leaves the job in its own processing list, and once its heartbeat
has been silent for five seconds a live replica puts the job back on the
queue, where it runs again from
attempt=1. Treat a#[process]handler as at-least-once and make it idempotent — the practical rule for any durable queue, and load-bearing here because replicas are how throughput scales. The scale-up half is measured incrates/nest-rs-redis/tests/e2e/replicas.rs. - A container restarted in place does not re-run its job. A replica is recognised by its hostname and pid, and a container restarted where it stood keeps both, so nothing sees its crashed predecessor as dead: the job that was running stays in the processing list the restarted process now owns, and does not run again. Replicas that differ by hostname are not affected.
Going further
Section titled “Going further”- Observability — what shows up in spans and logs when a job retries or fails.
- OpenTelemetry — wiring the queue spans into a trace backend.
crates/nest-rs-redis/src/worker/consumer.rs—settle: how an attempt’s outcome spends, or does not spend, the retry budget.crates/nest-rs-queue/src/inventory.rs—WIRE_FORMAT_VERSION, the envelope spec.