Bounded Telemetry Buffering at the Edge
Define capacity, durable receipt and reconnect behavior before a link goes away.
On this page
Give the buffer a contract
An edge buffer exists because the source and the destination cannot always exchange data immediately. It can smooth a short interruption, but it cannot preserve unlimited data on finite storage. Fig’s architecture study describes placement patterns; this companion note defines the practical contract for one local buffer.
Specify the accepted record, its identifier, the meaning of acknowledgement and the capacity boundary. Also specify the policy when the boundary is reached. Blocking, rejecting new records, dropping old records and aggregating samples are different choices with different effects on the meaning of the data.
This is a proposed protocol and test plan. Numerical examples are capacity estimates, not measurements from running devices. A completed result would need to record the actual record sizes, arrival rates, storage behavior and recovery observations.
Estimate the outage window in bytes
An illustrative source producing 100 records per second at 200 bytes each generates 20,000 bytes per second. One hour produces 72 million bytes of raw records before storage metadata and indexing. An apparently small event stream can therefore consume a meaningful local budget during a long interruption.
raw bytes = record rate × mean record size × outage duration
example = 100 × 200 × 3,600 = 72,000,000 bytesUse an observed distribution of record sizes rather than only the mean when unusually large records are possible. Reserve space for metadata, checkpoints and the storage engine’s maintenance work. Set both a total storage limit and a maximum accepted record size.
A count limit and a byte limit answer different questions. Ten thousand small records may fit while ten thousand large records do not. Record the expected outage window as a consequence of the budget, not as an independent promise.
Distinguish acceptance from durable receipt
If the buffer acknowledges a record before it survives the required failure boundary, the source may discard the only remaining copy. State whether acknowledgement means entry into memory, completion of a local transaction or confirmation by an upstream destination.
Durable receipt depends on the storage engine, transaction settings, filesystem and device behavior. A successful write call is not a universal statement about sudden power loss. Review the engine’s documented guarantees and test the intended boundary in a controlled environment.
Use a stable record identifier so retries can be recognized. The source may resend after a response is lost even when the local commit succeeded. A uniqueness constraint or equivalent deduplication policy makes that case explicit. Do not infer exactly-once effects from the presence of a queue.
Apply backpressure before the disk is full
Choose a high-water mark below the hard capacity limit. It leaves room for metadata and orderly maintenance. When the high-water mark is crossed, the source needs a documented response: a bounded wait, a rejection it can retry, or a permitted reduction in sampling.
If the policy drops data, record the scope of the drop. A counter can report how many records were discarded, while a separate range or gap marker can explain which time interval is incomplete. An aggregate must preserve enough metadata to distinguish a summary from the original samples.
Keep the retry policy bounded too. A source that immediately retries every rejected record can consume CPU and network capacity without creating more storage. Delay and a finite retry budget are part of the buffer protocol.
Drain with an explicit order and rate
On reconnect, decide whether old records drain before current traffic or share capacity with it. Strict age order preserves a simple sequence but can delay fresh observations. Separate lanes can prioritize current state while carrying older records with their original timestamps.
The drain rate must exceed the arrival rate if the backlog is to shrink. If the source generates 20,000 bytes per second and the effective drain is only 15,000, restored connectivity does not end the capacity problem. Include protocol overhead and destination throttling when measuring effective drain throughput.
Avoid turning reconnect into a burst that overloads the destination. A bounded batch size and concurrency limit make recovery more predictable than starting one request per queued record.
Write the interruption drill before running it
In a disposable setup, establish a baseline, stop the upstream link, continue a bounded source stream and inspect the queue’s age and bytes. Reach the high-water mark and verify the documented source response. Restore the link and measure backlog reduction while new records arrive.
Run a separate restart drill around acknowledgement. Check for records acknowledged before restart, records committed but not acknowledged, and records still pending. Compare identifiers at the source, buffer and destination. These observations reveal duplicates and gaps that a single queue-depth graph can hide.
Report the guarantee and the limits
A useful result says which records survived which failure under which storage settings. It also records the capacity policy and what happened when that policy activated. Avoid a broad claim that buffering makes the system reliable under every outage.
For a possible local transactional store, SQLite’s write-ahead logging documentation explains relevant transaction and synchronization behavior. The tool choice follows the contract; the contract should remain understandable without assuming a particular database.