Health Check Failure Drills
Test what a probe causes when a process stalls, a dependency fails or a service warms up.
On this page
Test the decision behind the response
A health endpoint is useful because another component makes a decision from it. A restart controller may act on liveness. A load balancer may act on readiness. An operator may inspect dependency state without either automatic action. The foundational article at Fig separates those roles; this field note turns the distinction into a test plan.
Run the drills in a disposable environment with a small service and a known request path. Record who consumes each check, its interval, timeout and failure threshold. If no component consumes the result, the drill can only establish how the endpoint responds. It cannot establish traffic removal or restart behavior.
This is a procedure for collecting evidence, not a claim that a deployed service has passed the drills. Define the expected action before injecting a failure. Changing the expected action after seeing the result makes the test much less useful.
Record a baseline and a clock
Begin with bounded requests to an ordinary page, a liveness endpoint and a readiness endpoint. Use a timeout shorter than the time you are willing to spend waiting. Record HTTP status, response time and the time of the observation. Keep the response body small enough that a health check is not itself an expensive workload.
curl --max-time 3 -o /dev/null -sS -w '%{http_code} %{time_total}\n' http://127.0.0.1:3000/
curl --max-time 3 -i http://127.0.0.1:3000/api/healthThe example publication exposes one small website-health route; it does not pretend to implement a complete orchestration health model. For a separate test service with distinct readiness and liveness routes, record those routes explicitly. Do not treat an endpoint name as proof that its implementation follows the intended semantics.
Use a consistent clock when comparing injection time with controller action. A timestamp from a container and one from an unrelated host can disagree. Where possible, collect the drill timeline in one place.
Drill one: an optional dependency disappears
Choose a dependency whose absence should degrade a feature while leaving the application process useful. An external article feed is a simple example. Block or replace that dependency in the test environment, then request both the feature and the health route.
The expected outcome is a bounded feature failure and continued availability of unrelated content. A process restart cannot repair a remote feed outage. If the liveness check triggers repeated restarts, the probe has probably coupled the process lifecycle to a dependency it does not control.
Also measure how often the application retries. A timeout bounds one attempt, but many concurrent attempts can still consume connections and memory. A retry delay, shared in-flight request and cached fallback make the behavior more predictable. Verify those properties through request counts, not just the visible error message.
Drill two: a required dependency is unavailable
A request path that cannot safely complete without its data store should stop receiving relevant work when that store is unavailable. In the test environment, make the required dependency fail, observe readiness, and send a bounded request through the actual routing component.
A failed readiness response is only the first observation. Verify that the router removes the instance after its configured threshold and that an already accepted request receives a controlled outcome. Existing connections and in-flight work may behave differently from new connections.
Liveness can remain successful while readiness is unsuccessful. That is not contradictory: the process may be capable of recovering without a restart while it cannot currently accept useful work.
Drill three: the process stops making progress
Use an application-specific test hook or a deliberately stalled disposable process. Avoid freezing a production host or killing a shared service to demonstrate a probe. Observe whether the probe itself remains responsive and whether the controller notices the loss of useful progress.
A response from a separate health thread can stay green while the worker handling normal requests is stuck. Conversely, a check that performs a long dependency request can fail even when the process is healthy. Document what the check actually observes and which failure modes are outside its view.
After recovery, confirm that the instance resumes useful work. A successful restart count is not sufficient evidence that requests now complete.
Drill four: startup takes longer than usual
Introduce a controlled startup delay in the disposable service. Check whether traffic is withheld until initialization is complete and whether the restart policy gives the process enough time to become useful. Initial warm-up and steady-state failure usually need different treatment.
Observe the reverse transition too: readiness should become healthy when the service is ready, not after an arbitrary extra period. Keep startup thresholds grounded in a measured environment, rather than copying a large timeout from another system.
Turn the timeline into a result
For each drill, record the injected fault, the expected controller decision, probe responses, action time, request outcome and recovery. Mark an untested assumption explicitly. A compact result can say that one router removed one instance under a specified timeout policy; it should not claim that every failure is detected.
The Kubernetes probe documentation provides a concrete reference for the separate roles of liveness, readiness and startup probes. The same questions apply even when a much smaller deployment uses a different controller.