Software Engineering

A Background Worker Recovery Playbook

Follow an item through lease, side effect, commit and acknowledgement, then test the gaps.

On this page
  1. Name the work item before the failure
  2. Draw the acknowledgement boundary
  3. Inspect the actual state before requeueing
  4. Drill one: stop before the effect
  5. Drill two: stop after commit and before acknowledgement
  6. Drill three: make the item permanently invalid
  7. Observe recovery pressure and shutdown
  8. Finish with an auditable outcome

Name the work item before the failure

Recovery starts with the accepted work, not with a restart command. A stable item identifier lets an operator distinguish a missing effect from a repeated delivery. Fig’s worker-design article explains retries and acknowledgement; this companion note describes a bounded recovery review.

Record the item identifier, acceptance time, handler version, attempt number and terminal outcome. Do not log the entire payload merely to obtain correlation. Credentials and unnecessary personal data should stay out of operational events.

This playbook is a test plan, not evidence that a queue or handler has passed recovery testing. Use an isolated queue and a disposable destination. Set a finite number of test items and identify which side effects can be safely repeated or removed.

Draw the acknowledgement boundary

A common path is acceptance, delivery or lease, handler execution, durable effect, and acknowledgement. Failures between these steps create different recovery questions. The queue’s delivery behavior and the handler’s effect behavior must be considered together.

A process can commit an effect and stop before acknowledging the item. The next delivery may repeat the handler. If the effect is a transaction in one database, a stable key and a uniqueness constraint can help make repetition harmless. If the effect is an external payment, message or API operation, the same local transaction cannot cover that remote boundary.

Use the destination’s supported idempotency mechanism when available. When a local state change must cause later remote work, an outbox can record the intent alongside the local transaction. It still needs its own retry and deduplication contract; naming the pattern does not automatically provide exactly-once behavior.

Inspect the actual state before requeueing

For one failed item, collect queue delivery state, the destination record and correlated handler events. A log saying that processing started is not evidence that the effect committed. A missing acknowledgement is not evidence that the effect did not happen.

Classify the state narrowly: pending delivery, currently leased, effect committed but acknowledgement uncertain, or terminal failure. If the available evidence cannot distinguish two states, preserve that uncertainty and choose a recovery action that remains safe in both.

Do not bulk requeue every item with an error event. An error may describe an optional follow-up after the main effect succeeded. Requeueing at that point can duplicate the main effect unless the handler is designed for repetition.

Drill one: stop before the effect

Use a controlled test hook to stop the disposable worker after delivery but before its effect. Observe lease expiry or redelivery according to the queue’s policy. Confirm that another worker eventually sees the item and that the destination receives the intended result.

Record the delay rather than assuming that a restarted process immediately gets the item back. A long visibility timeout, exhausted retry budget or poison-item policy can change the outcome. The queue’s documented rules belong in the setup.

Drill two: stop after commit and before acknowledgement

This is the case most likely to reveal an unsafe duplicate. Arrange a controlled stop after the test effect commits and before acknowledgement. On redelivery, verify both the handler’s response and the destination’s final state.

The expected outcome should be stated before the drill. A deduplicated database write may produce the same state without executing a second effect. An external operation may need an idempotency key and a lookup for a previously completed result. If that destination provides neither, document the remaining duplicate risk instead of calling the handler exactly-once.

A transport success response can also be lost. Include that ambiguous response case when the external boundary is important.

Drill three: make the item permanently invalid

Send one intentionally invalid item to the isolated queue. Verify that validation fails in a controlled way, retrying is bounded, and a terminal outcome is recorded. A permanently invalid payload should not loop forever or block unrelated valid work.

Preserve a minimal failure reason and the identifier needed for investigation. A dead-letter destination can support review, but its retention and access rules still need a budget. Moving an item to another queue does not itself explain or repair it.

Observe recovery pressure and shutdown

Track oldest pending age, retry counts and terminal outcomes alongside queue depth. A stable depth can hide a backlog of items that continually fail while new items complete. Bounded concurrency prevents a retry storm from exhausting a shared destination.

During shutdown, stop accepting new work, give active handlers a bounded completion window, and let unresolved leases follow the queue contract. Do not acknowledge unfinished work simply to make shutdown look clean. Record which part of the handler can be interrupted safely.

Finish with an auditable outcome

For each drill, preserve the item ID, injection point, expected behavior, observed deliveries, destination state, timing and cleanup. The result should establish the behavior of one specified handler and queue configuration, including any ambiguity that remains.

RabbitMQ’s consumer acknowledgement and publisher confirm guide is a useful primary reference for one concrete queue implementation. Other systems have different lease and acknowledgement semantics; translate the playbook into the actual queue contract before using it.