DevOps
Industrialising a delivery pipeline and moving to GitOps without breaking everything
A useful pipeline is not the one with the most stages, but the one that answers a simple question fastest: is this change any good? Here is how we rebuild it, then switch deployment to GitOps.
Most pipelines we audit share the same profile: forty minutes of compilation and tests, a high failure rate, and a manual release that concentrates all the risk. Rebuilding always starts with measurement, never with a tool.
Step 1: measure before touching anything
We collect four figures from the last thirty runs: total duration, queue time before start, failure rate, and the share of failures caused by infrastructure rather than code. That last point is often revealing: when a third of failures come from the environment, fixing tests achieves nothing.
- Total duration and queue time before start.
- Overall failure rate, and the infrastructure-related share.
- Average time to fix after a detected failure.
- Number of deployed versions that can be rebuilt identically.
Step 2: split the pipeline by objective
A useful pipeline reads like a sequence of questions. Each stage answers one question, and its failure must be understandable by the person who has to fix it, without reading a thousand-line log.
| Stage | Question it answers | Target duration | Blocking |
|---|---|---|---|
| Compilation | Does the code build? | Under one minute | Yes |
| Unit tests | Do business rules hold? | Under three minutes | Yes |
| Static analysis | Are sensitive points introduced? | Under two minutes | Yes on defined thresholds |
| Integration tests | Does the application work with its real dependencies? | Under eight minutes | Yes |
| End-to-end tests | Do critical journeys work? | Under ten minutes | Yes on critical journeys |
| Image publishing | Is the artefact versioned and signed? | Under three minutes | Yes |
The catch-all stage trap
Many pipelines pile checks into a single stage, which makes failures unreadable and prevents parallelisation. Separating stages lets you run unit tests, static analysis and image building in parallel, then run heavy tests only after the fast checks pass.
Step 3: make the deployed version verifiable
A deployment should never rest on an assumption. We expose the built version and build timestamp on a technical endpoint, queried automatically after every deployment.
@RestController
class VersionController {
private final BuildProperties buildProperties;
VersionController(BuildProperties buildProperties) {
this.buildProperties = buildProperties;
}
@GetMapping("/internal/version")
VersionResponse version() {
String time = buildProperties.getTime() == null
? "unknown"
: buildProperties.getTime().toString();
return new VersionResponse(buildProperties.getVersion(), time);
}
}
This endpoint may look trivial: in reality it is the precondition for fast diagnosis. Without it, finding out which version actually runs in an environment means checking a registry, a dashboard and sometimes a container log.
Step 4: switch deployment to GitOps
GitOps moves the question of deployment: we no longer apply commands, we describe a desired state. The configuration repository becomes the source of truth, and any drift observed on the cluster is an anomaly to be explained.
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: billing-api-staging
namespace: argocd
spec:
project: default
source:
repoURL: https://example.com/git/platform-config
targetRevision: main
path: environments/staging/billing-api
destination:
server: https://kubernetes.default.svc
namespace: billing
syncPolicy:
automated:
prune: true
selfHeal: true
syncOptions:
- CreateNamespace=false
Rollback becomes an ordinary operation: revert to the previous revision of the configuration repository. That simplicity has one condition: database migrations must remain compatible with the previous version of the code for at least one release.
A delivery pipeline is judged by how long it takes to learn that a change is bad, not by how many stages it contains.
Step 5: train the team on a real case
The last step is to deliberately cause a failure in staging: a faulty image, an incompatible migration, an unavailable dependency. The team performs the rollback, then we correct the documentation based on what was actually needed. That session usually reveals two or three missing procedures nobody had identified.
What to remember
Industrialising a pipeline is not a tooling project: it is work on reducing feedback time. Start with measurement, split by question, make the deployed version verifiable, then move deployment to GitOps. The rest follows naturally.
Related articles
Observability with OpenTelemetry: instrument what actually helps diagnosis
Adding a tracing library does not make a system observable. What matters is attribute choice, context propagation and discipline about stored volumes.