A useful health view explains delivery, data quality, and freshness without collapsing them into one green light.
Start with the user’s question
A server can answer quickly while returning an old observation. It can also return valid data too slowly for the screen that needs it. Start monitoring with a concrete question: can an authorized user retrieve usable data within the time and freshness requirements of this feature?
Our suggested health view separates request success, response duration, payload validation, and observation age. Keep these measurements distinct so that a fast error response does not improve your apparent delivery performance, and a successful request does not erase a stale-data warning.
Classify response outcomes
RFC 9110 defines status codes as descriptions of HTTP request outcomes. A 200 response indicates success at that level; it does not certify that a market observation meets your application’s freshness threshold. A 503 indicates temporary inability to handle the request, such as overload or maintenance.
Classify transport failures, HTTP errors, invalid payloads, and old observations separately. Record the endpoint pattern and a fixed outcome category. Avoid storing raw response bodies, authentication headers, or full URLs with query parameters just to produce a health chart. [1]
Give every metric a denominator
A success percentage needs a defined population and time window. State whether it includes retries, client cancellations, denied requests, and scheduled probes. Track the number of observations behind it: one successful request tells you less than a sustained record of successful delivery.
We suggest reporting latency for successful requests alongside failure counts, then looking at slower percentiles as well as typical responses. Separate synthetic probes from customer traffic. A probe can establish that one path worked at one moment; it should not be counted as a customer request or proof that every account can connect.
Make alerts explainable
Choose alert thresholds from your feature’s requirements and expected session schedule. An absence of new prices does not, by itself, distinguish a closed market from a collector problem. Combine observation age with known session context and delivery health before assigning a cause.
Test an outage and its recovery using controlled fixtures. Confirm that the incident opens, records repeated failures without creating endless duplicates, and clears only when the recovery condition is met. Keep the last successful check and last observation visible so the next person investigating can see what changed.
References
Primary sources for the technical details in this article. Implementation suggestions are our own.