Web & Mobile · 5 minute read
Observability for Web Apps: Logs, Metrics and Traces
Observability is what lets you diagnose an incident rather than guess. The essentials are structured logs with correlation identifiers, metrics on the paths that matter, traces across service boundaries, and alerting tied to user impact rather than to individual component failures.
Observability is what separates diagnosing an incident from guessing at it, and the difference is almost entirely in what was instrumented beforehand. This guide covers the essentials, drawing on FISTA Solutions' web and mobile work.
What are the three signals for?
Different questions, and none of them substitutes for another.
| Signal | Answers |
|---|---|
| Metrics | Is something wrong, and how much |
| Logs | What happened in this specific case |
| Traces | Where the time went across services |
| Real user monitoring | What the browser actually experienced |
| Error tracking | What is failing and how often |
| Correlation identifiers | Ties all of the above together |
Why do correlation identifiers come first?
Because without them the other signals cannot be joined.
A request identifier generated at the edge and propagated through every service, every log line, and every trace turns an investigation into a query. Without it, you are matching timestamps across systems with different clocks, which is slow and produces wrong conclusions.
Generate it at the earliest point, propagate it in headers, include it in every log line, and return it to the client so support can quote it. That last part turns a customer report into a direct lookup.
What makes logs useful?
Structure and consistency. Logs as structured records with named fields can be queried; logs as formatted text require parsing that breaks whenever someone changes a message.
Include the correlation identifier, the user or tenant where relevant, the operation, and the outcome. Avoid logging entire request bodies, which is expensive and frequently captures data you should not retain.
Agree the fields once and enforce them through a logging library rather than by convention. Conventions drift.
What should metrics cover?
Request rate, error rate, and latency distribution for every important path, plus the business events that indicate the system is doing its job.
Latency should be measured at percentiles rather than as an average, because the tail is what users notice and what averages conceal.
Watch cardinality. A metric labelled with user identifiers produces a series per user, which is expensive and frequently the cause of a surprising telemetry bill.
What do traces add?
The explanation for cross-service latency, which metrics cannot provide.
A slow endpoint metric tells you something is slow. A trace tells you which of the six downstream calls took the time, and whether they ran in sequence when they could have run in parallel.
Sample rather than tracing everything: a small percentage of requests plus all errors and all slow requests gives most of the diagnostic value at a fraction of the cost. See what is distributed tracing for AI.
How should alerting be designed?
On symptoms users experience rather than on component conditions.
A database replica failing is not an alert if traffic is served normally. An error rate rising on checkout is, regardless of which component caused it. Component alerts produce noise, and noise trains people to ignore the pager.
Every alert should have a runbook entry and a plausible action. Alerts that fire and require no response should be removed rather than tolerated. See how to set up AI on-call.
What about the client side?
Real user monitoring is the half that server-side observability cannot see.
Server metrics show what the server did; they do not show script errors, slow rendering on mid-range devices, failed requests from flaky networks, or the user who gave up. Those are the experiences that matter and they are only visible from the browser.
Capture errors, performance timings, and failed requests, segmented by device and connection. See core web vitals guide.
What are the common mistakes?
No correlation identifiers. Text logs. Alerting on components. Tracing nothing or everything. High-cardinality metric labels. And server-side telemetry with no client-side view.
How do you test it?
Verify correlation identifiers propagate through every service, check that an alert fires in a staged failure, and confirm logs contain what an investigation would need.
Run an incident drill using only the telemetry. The gaps surface immediately and they are cheap to close before a real incident finds them.
What does it cost to operate?
Grows faster than traffic, because instrumentation grows too. Log volume, metric cardinality, and trace sampling rate are the three levers.
Set retention tiers: short retention for verbose logs, longer for errors and audit-relevant records. Undifferentiated retention is where telemetry bills come from.
What should you measure?
Mean time to identify the cause during incidents, alert-to-action ratio, telemetry cost per request, and the proportion of incidents diagnosed from telemetry rather than by reproduction.
What changes with AI components?
They need trajectory logging as well as request logging: the prompt, the retrieved context, the tool calls, and the output, tied to the same correlation identifier.
Without that, an incident where the system produced a wrong answer cannot be reconstructed at all. It is the AI equivalent of a stack trace and it must be captured before it is needed. See how to debug an AI agent.
When is this the wrong approach?
Full observability tooling is disproportionate for a small internal application with one service and few users. Error tracking and basic metrics cover it, and the rest is cost without a question it answers.
What should you do first?
Check whether a single identifier ties a user's report to logs across every service. If not, adding that is the change that makes everything else useful.
How FISTA Solutions helps
FISTA Solutions builds and operates production systems through web and mobile, AI enablement, and staff augmentation: correlation identifiers propagated end to end so incidents are queries rather than reconstructions, alerting tied to user impact rather than component state, decisions documented with their reasoning, and handover that leaves your team able to maintain what was delivered. The record is 150+ projects for 50+ companies across 12+ countries.
To scope this work, message FISTA on WhatsApp, or read CI/CD for web apps.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01What should be instrumented first?
A correlation identifier propagated through every request and log line. Without it, investigating an incident means matching timestamps across services, which is slow and frequently wrong.
02Why structured logs?
Because they can be queried. Text logs require parsing with fragile patterns, and the field you need is usually the one nobody thought to format consistently.
03What should alerts fire on?
Symptoms users experience: error rates, latency at the tail, and failed critical journeys. Component-level alerts produce noise, because components fail routinely without users noticing.
04When are traces worth the effort?
As soon as a request crosses more than one service. Traces are what turn a slow endpoint into a specific slow call, and without them the investigation is a sequence of guesses.
05How do you control telemetry cost?
Sampling for traces, retention tiers for logs, and cardinality discipline for metrics. Telemetry volume grows with traffic and with instrumentation, and it becomes a material line item quickly.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. Weâll map the fastest credible path from intent to verified production.