Cloud and DevOps
Observability and application performance
I make your applications diagnosable: correlated traces, useful metrics and alerts that signal a real problem. You move from interpreting symptoms to identifying the cause.
What it covers
An observable system is recognised by how quickly a question becomes an answer: why did this request slow down yesterday at three, which service produced an error on this journey, how long did processing actually take. Without correlated traces, those questions turn into hours of discussion.
I start by instrumenting critical journeys rather than every component. Each request carries an identifier propagated across services, calls to databases and external systems are measured, and errors are enriched with the context needed for diagnosis.
Alerts worth waking up for
Alerts are defined from service objectives rather than arbitrary thresholds. They signal degradation perceived by users, with a false-positive rate tracked and reduced over time. Dashboards are organised by business journey.
- Distributed traces with an identifier propagated across services.
- Application and technical metrics tied to user journeys.
- Alerts based on measurable service objectives.
- Dashboards per journey rather than per server.
- Cost analysis for trace and log storage.
Problems addressed
Incidents are diagnosed by elimination, for lack of traces. Too many alerts end up being ignored. Slowdowns are noticed but never precisely attributed. Logs are voluminous but hard to use in an emergency. Trace storage costs grow without control.
Expected benefits
Incident diagnosis in minutes with an identified cause. Fewer useless alerts draining the team. Localised bottlenecks without opinion-based debate. Service indicators tracked and communicated objectively. Diagnostic capability retained during traffic peaks. Observability storage costs controlled and justified.
Method and steps
- 1
Question framing
Interviews with operations and development teams to list the questions observability must answer, ordered by real frequency.
- 2
Instrumentation
Adding traces, metrics and structured logs on critical journeys, with context propagation across services and to external systems.
- 3
Dashboards and alerts
Building views per business journey and defining alerts tied to service objectives, with false-positive rate measurement.
- 4
Noise reduction
Removing useless alerts, adjusting thresholds, introducing retention policies and tracking storage costs.
Deliverables
List of diagnostic questions and instrumented journeys.
Application instrumentation with end-to-end correlation.
Dashboards per business journey.
Alert rules tied to service objectives.
Retention policy and observability cost report.
Diagnostic guide for on-call teams.
Technologies used
- Spring Boot
- PostgreSQL
- Kubernetes
- Prometheus
- Grafana
- OpenTelemetry
- Elastic Stack
Frequently asked questions
Should the whole application be instrumented?
No, and doing so is expensive in volume and complexity. We start with critical journeys and external dependencies, then extend only where a diagnostic need is demonstrated.
How do you avoid alert overload?
By tying every alert to a service objective and an expected action. An alert with no clear action is removed or turned into a dashboard indicator.
What is the performance impact?
Instrumentation has a cost, usually a few percent of processing time. We measure it and tune it, particularly on high-volume processing where sampling is preferable.
Related case studies
Incremental overhaul of a contract management monolith
Modernisation through domain extraction, with no service interruption
Demonstration sample - not a real client reference.
Multi-environment Kubernetes platform for a software vendor
A shared foundation for twelve applications and eighty customer instances
Demonstration sample - not a real client reference.
API and integration foundation across eleven systems
Versioned contracts, idempotency and incident recovery
Demonstration sample - not a real client reference.
Talk about my observability
Describe the situations where your team loses the most time. We will target instrumentation at those specific cases.