An Observability Budget for a Small Host
Estimate the cost of signals, bound log growth and verify that an alert answers a real question.
On this page
Start from questions and constraints
A small host needs a way to explain failure, but it also has a finite memory and storage budget. Installing a large monitoring stack before choosing the questions can turn observability into the dominant workload. The foundational article at Fig develops the relationship between metrics, logs and traces. This note applies that model to a modest environment.
Write down three questions first: is the application completing requests, are resources approaching an unsafe limit, and can one failed request be explained? These questions suggest a small initial signal set. They do not establish that one tool or dashboard is required.
Record the host capacity, the application’s normal resource use and the space already committed to other services. Treat every number in the budget as either an observed value or an assumption. The illustrative calculations below are not benchmark results from a deployed machine.
Give logs a storage budget
A first estimate is event size multiplied by event rate multiplied by retention time. At an illustrative 250 bytes per event and 20 events per second, one day contains approximately 432 million bytes of raw event data, or about 412 MiB. That estimate excludes indexing, replication, framing and compression.
raw bytes per day = bytes per event × events per second × 86,400
example = 250 × 20 × 86,400 = 432,000,000 bytesThe difference between decimal bytes and binary MiB is small compared with an unknown event rate, but write the unit down. Measure representative events and actual traffic before choosing retention. Stack traces, request payloads and repeated errors can make a quiet baseline misleading.
Set a maximum log size and a bounded number of rotated files. A restart policy does not protect the host from unbounded logs. Test rotation in a disposable instance and confirm that older data can be removed without affecting application correctness. Logs should explain work; they should not be the only durable record of accepted work.
Keep the metric label set finite
Start with request count, error count, a useful latency distribution, resource pressure and the age of pending work when a queue exists. Labels should describe bounded categories such as route templates, result classes or a small service set.
A user identifier, raw URL, arbitrary exception message or request ID can create a new time series for each event. Moving those values into logs keeps metric cardinality more predictable. A route template such as /articles/[slug] is different from a label for every possible query string.
The budget should include scrape frequency and retention, not only the collector’s idle memory. A short scrape interval increases samples even when the number of series remains constant. Begin with a rate that can answer the operational question, then revise it when a measured need appears.
Inspect resource pressure at the correct boundary
Use host tools for host memory pressure and runtime-container tools for the application container. A container’s resident memory, the host’s available memory and swap use describe different aspects of the same environment. None is a complete explanation by itself.
free -h
vmstat 1 5
docker stats --no-streamThe container command describes running containers. It should not be assumed to expose every temporary BuildKit build process. During a build failure, inspect the host process list and kernel log too. A compiler can require substantially more memory than the resulting server uses when idle.
For memory-pressure investigations, correlate observation time with process termination and the container’s configured limits. A SIGKILL is a forced termination, but the kernel log or cgroup evidence is needed to establish that a memory limit caused it. Do not publish those logs unredacted: they may contain process names and internal paths.
Choose structured fields that survive a failure
A compact event can include a timestamp, service label, event name, severity, request ID and outcome class. Keep the identifier consistent across the boundary being investigated. It is a correlation value, not a reason to place the full request body in a log.
Avoid credentials, session tokens and unnecessary personal data. Query strings can contain secrets even when the application did not intend them to. Record a route or a minimal event name when that is enough to explain the request. Retention and access control remain part of the design even on a single host.
Test one alert before trusting it
Pick a bounded failure, such as a disposable service returning errors or a test queue accumulating old work. Verify that the signal changes, the threshold is reached, the notification describes the condition, and the alert clears after recovery. Record the complete delay from fault to useful notice.
A dashboard screenshot is evidence that a chart rendered. It is not evidence that the on-call action is correct. If the alert cannot suggest an inspection step, refine the question and the message before adding more signals.
Revisit the budget with observations
After a representative period, compare estimated event volume, actual storage, collector memory and the time needed to explain a failed request. Preserve the assumptions that were wrong; they are useful input to the next version of the budget.
The OpenTelemetry signals overview provides terminology for metrics, logs and traces. The practical goal here remains small: collect enough trustworthy evidence to answer the chosen questions without exhausting the host you are trying to observe.