Skip to main content

DevOps

Observability with OpenTelemetry: instrument what actually helps diagnosis

Adding a tracing library does not make a system observable. What matters is attribute choice, context propagation and discipline about stored volumes.

- 10 min read

Instrumenting an application has become easy: current libraries produce traces, metrics and logs with a few lines of configuration. The real work starts afterwards, when a diagnostic question must be answered in under ten minutes.

Start from questions, not components

Every engagement begins with a list of concrete questions asked by operations and development teams. Successful observability answers a few dozen specific questions, not every conceivable one.

  • Which service produced the error the user sees on this journey?
  • How long did the partner system call take, and was it retried?
  • Which application version was deployed when the slowdown started?
  • Which SQL query concentrates processing time on this screen?
  • Was the message processed once, or replayed after an incident?

Attribute choice decides usefulness

A trace without business attributes tells you a request failed, never why. We systematically add low-cardinality attributes such as journey identifier or operation type, plus a limited set of technical attributes such as application version and region.

Personal data stays out of attributes: customer identifiers, addresses and message contents must never appear in a trace, a log or a metric name.

@Service
class PaymentInstrumentation {

    private final ObservationRegistry registry;
    private final PaymentGateway gateway;

    PaymentInstrumentation(ObservationRegistry registry, PaymentGateway gateway) {
        this.registry = registry;
        this.gateway = gateway;
    }

    PaymentResult process(PaymentCommand command) {
        return Observation.createNotStarted("payment.process", registry)
                .lowCardinalityKeyValue("payment.method", command.method().name())
                .lowCardinalityKeyValue("app.version", Version.current())
                .observe(() -> {
                    Span.current().setAttribute("payment.amount.bucket", Buckets.of(command.amount()));
                    return gateway.authorize(command);
                });
    }
}

Grouping amounts into buckets rather than exact values is deliberate: a high-cardinality attribute multiplies storage cost and degrades aggregation readability.

Context propagation and correlated logs

Propagating context across HTTP calls, messages and asynchronous processing is what distinguishes an observable system from a collection of isolated traces. We verify this propagation on the most complex paths: outbound calls, message consumers and scheduled tasks.

SignalQuestion it answersVolume discipline
TracesWhere did time go, in what order did calls occur?Sampling per journey, limited retention
MetricsDoes the service meet its objectives over time?Aggregation, no personal data
LogsWhat happened for this specific request?Controlled levels, retention by type
Business eventsDid the expected process complete?Counting, no detailed content

Alerts tied to objectives

We define a small number of service objectives per critical journey, then one alert per objective. Any alert with no expected action is either removed or turned into a dashboard indicator. That discipline is the only way to durably reduce noise.

A trace without business attributes tells you a request failed, never why.

Controlling storage cost

Observability cost grows in invisible steps: more services, more attributes, more retention. We track three levers: sampling rate per journey type, trace versus log retention, and indexing of search attributes. A thirty to forty percent volume reduction is common without losing diagnostic capability, provided you agree never to sample errors.

What to remember

Useful observability is built around diagnostic questions and chosen attributes, not around a tool. Instrument critical journeys, exclude all personal data, tie every alert to an objective, and monitor cost like any other infrastructure resource.

  • Kubernetes
  • Observability
  • OpenTelemetry
  • Performance

Related articles