Here's a story most developers will recognise. A customer says checkout is "slow sometimes". The team checks the server's CPU graph, which looks fine, then the logs, which hold several million lines of INFO: request processed. Someone blames the database and someone else blames the payment provider. Two days later the cause turns out to be a retry loop in a third-party address lookup that only fires for certain postcodes.
Nobody in that story was careless. They just couldn't see the system. Monitoring tells you something is wrong. Observability lets you ask questions you didn't think of in advance, like "show me every slow checkout this week and what they have in common".
Why it used to be an enterprise luxury
Until fairly recently, good observability meant picking a commercial vendor, installing its proprietary agent, instrumenting your code with its SDK and paying its bill for as long as you wanted your data. Switching vendors meant instrumenting everything again. For a small team the cost and lock-in were hard to justify, so most made do with logs and a basic uptime check.
OpenTelemetry changed that. It's an open standard, backed by the Cloud Native Computing Foundation, for producing the three core signals of observability: traces, metrics and logs. You instrument your code once against a vendor-neutral API, and where the data goes becomes a config setting you can change later. It might be an open-source backend you run yourself, a managed service or a commercial platform.
Almost every major observability vendor now accepts OpenTelemetry data natively, so instrumentation has become a commodity. That's good news for anyone who'd rather not pay for lock-in.
The three signals
Traces: the story of one request
A trace follows a single request through your whole system, from the web tier to the API, the database queries, the calls to third parties and any background job it kicks off. Each step is a span with a start time, a duration and attributes. Traces would have found that address-lookup bug in minutes. You filter for slow checkouts, open one, and there's the culprit span taking four seconds.
If you only have time for one signal, make it this one. It's the most useful of the three and the one small teams are most often missing.
Metrics: the shape of the system over time
Request rates, error rates, latency percentiles, queue depth, cache hit ratio. Metrics are cheap to store and quick to query, which makes them the right foundation for dashboards and alerts. Pick the few that reflect what users experience, not every number your infrastructure can spit out.
Logs: detail when you need it
Logs don't go away, but they get far more useful when every line carries the trace ID of the request that produced it. Instead of searching millions of lines, you jump from a slow trace straight to that request's logs.
A one-week starting plan
You don't need a platform team for this. Here's the order we use with smaller clients.
Day 1, turn on auto-instrumentation. Most mainstream languages and frameworks have OpenTelemetry auto-instrumentation that captures incoming HTTP requests, database calls and outbound HTTP with little or no code change. Switch it on for one service and look at the traces. That alone is usually an eye-opener.
Day 2, add the collector. The OpenTelemetry Collector sits between your apps and your backend, batching, filtering and routing data. It's where you drop noisy spans, strip sensitive fields and keep volume under control, and it's what makes switching backends painless later.
Day 3, choose a backend. For a small team, a managed service with a generous free tier is usually the pragmatic pick. The instrumentation doesn't care, so you can change your mind later.
Day 4, add business context. Auto-instrumentation knows about HTTP and SQL, but it doesn't know your business. Add a handful of custom attributes like customer tier, order value, feature flag state or tenant ID. That's what turns "some requests are slow" into "requests from customers on the old pricing plan are slow".
Day 5, build one dashboard and three alerts. The dashboard covers the key user journeys, with rate, errors and duration for each. Alert on error-rate spikes, on latency getting worse, and on a heartbeat for critical background jobs. Don't add more until these have earned their keep.
Keeping the bill sane
The quickest way to make observability expensive is to keep everything.
Sample traces sensibly. Keep every trace with an error or unusual slowness and a small percentage of the healthy ones, because you rarely need a million identical successful requests. Set retention on purpose, with detailed traces kept for a couple of weeks and aggregated metrics for longer, since most investigations happen within days of the problem. And drop noise at the collector. Health checks, static asset requests and chatty internal polling can make up a big share of volume while telling you nothing.
The payoff
Faster incident response is the obvious win. The less obvious one is better decisions. Once you can see which endpoints are slow, which features get used and which third-party dependency is dragging everything down, prioritising stops being an argument about opinions.
If your team spends more time guessing than fixing when production misbehaves, we can help you get the basics in place.

