Service · Software development

Setting up observability and monitoring.

Software made up of a handful of services can no longer be debugged by reading the logs. We set up metrics, log entries and traces so that you can follow a request from start to finish, and we base alerting on what your users experience rather than on whatever a machine happens to report.

OpenTelemetryTracingAlertingError budget

What we do and don't do here.

Most organisations don't measure too little; they measure too much, and without coordination. Three dashboards that contradict each other, two alert channels one of which is muted, and a log file so large that nobody searches it any more. The question that matters during an incident – where in the chain is this going wrong – still can't be answered without calling in three people.

Good tools already exist for this: Grafana with Prometheus, Loki and Tempo, Datadog, Elastic, New Relic, Honeycomb, Sentry, and the built-in services of Azure, AWS and GCP. If you already run something and it works, we'd advise keeping it. We rarely replace a platform as a standalone solution; adding a second vendor alongside usually fragments the picture further.

What we do provide is the layer in between: instrumentation in your software so there is something to measure, a trace that continues across system boundaries, and alerting tied to a service rather than to a server. We often do this as part of a modernisation project, alongside setting up DevSecOps in the pipeline, or on its own when faults are piling up and nobody can say why.

We do this as a standalone project and as part of a broader custom software development project. If it comes with a move to a managed services model, also look at microservices architecture and optimising your IT operations. You can read how we handle your data in our information security policy.

Three types of engagement we take on here.

Most engagements start with one of these three and grow once the first picture is in place. In the first conversation we'll tell you which option will deliver the most for your landscape, and which you're better off postponing.

First engagement · fixed sprint budget

Setting up the four signals per service

Latency, error rate, traffic and saturation – per service, with the spread shown alongside, not just the average. An average of two hundred milliseconds could mean everyone waits two hundred, or that one in a hundred waits eight seconds; the latter is where your complaints come from. We capture that measurement in the software, bring it together on a dashboard that's about your service delivery, and remove the alerts that concern isolated technical values.

Four golden signalsPercentilesPer-service dashboardAlert clean-up
Mid-sized project · fixed sprint budget

Tracing across system boundaries

A request that passes through four services gets a trace on arrival that follows it everywhere. Through your own code, through the queues in between and through calls to external parties. We build this with OpenTelemetry, so you aren't tied to a single vendor, and make sure log entries and measurements carry the same trace. From that point on, the question "where did this request get stuck" takes a click instead of an afternoon.

OpenTelemetryQueue tracingLinked log entriesExternal calls
Larger project · fixed sprint budget

Error budgets and service agreements

A target value per service – how many requests may fail or take too long before it becomes a problem – and a budget that follows from it. That turns "it's slow" into a number you can have a meaningful conversation about, and it gives your team a fair line between building new features and repairing stability. Including reporting to your client or, where you are the supplier, to your customers.

Target values per serviceError budgetReportingEscalation path

What you have at the end of an engagement.

A working setup that runs on your own platform, with the knowledge in your team to extend it without bringing us back in.

  • Instrumentation in your softwareMeasurements, log entries and traces from the code itself, via OpenTelemetry, so you can switch platforms without rebuilding everything.
  • Dashboards per serviceAn overview per service with the four signals and the spread, plus a home page that shows at a glance whether anything is wrong.
  • Alerting that makes senseAlerts based on what your users notice, with a clear recipient, an escalation path and a short description of what the recipient should do first.
  • A retention policy with its costsFor each type of data, we record how long it is kept and why, including sampling on traces, with the agreement that anything that fails is kept in full.
  • Runbooks for the most common incidentsA short page for each recurring pattern: how you recognise it, where you look first, and what intervention resolved it last time.
  • Handover to your teamTwo working sessions in which your developers instrument a service and set up an alert themselves, so that extending it doesn't remain with us.

When this is worth the effort.

Four situations we keep seeing in organisations that approach us for this. If you recognise one, there is probably something to gain.

Searching takes too long

Every incident is a hunt

Something breaks and three people sit looking at three systems, without knowing whether they're talking about the same request. Recovery time is determined not by the repair but by the finding. That is precisely what a continuous trace removes.

Alert fatigue

Nobody looks at the notifications any more

So many warnings come in that the channel has been muted. The moment something real happens, it no longer stands out amid the noise. The problem is almost never the volume but the nature: alerts fire on machines rather than on the service you deliver.

Improvement cannot be demonstrated

Every discussion comes down to gut feeling

You want to show that an investment in stability has paid off, but there are no measurements from before that point. Without figures, every improvement becomes an opinion, and an opinion won't hold up in a conversation about where the budget goes.

Multiple services

It is no longer an application but a landscape

What started as one application has grown into a handful of services with queues and external integrations between them. The tooling that worked for a single application now gives you a view of each component rather than of the whole.

How such a project runs.

1

Introduction and intake

A conversation about what currently runs, what is already measured and where the last few incidents came from. We ask about the systems in between (queues, external integrations, scheduled jobs), because that is usually where the blind spots are. We also map which integrations become part of the tracing.

2

Baseline measurement and scope

We measure where things stand today: which services already emit signals, which emit none at all, and what the current storage costs. We also review recent outages and ask, for each one, what someone would have wanted to see. At the end, there is a scope that starts with the service where the most time is lost, not with the service that is easiest to instrument.

3

Instrumenting in sprints

We work in sprints and deliver, each sprint, a fully connected service: metrics, log entries with traceability, and traces that run through to what sits behind them. Your own developers take part, because instrumentation added by someone else is rarely maintained.

4

Alerting and runbooks

Only once enough is being measured do we set up alerting. For each alert, we record who receives it, what the first action is, and when it escalates. At the same time we clear out existing alerts, which in practice brings more peace than adding new ones.

5

Handover and aftercare

Two working sessions with your team, a short period in which we observe real incidents alongside you, and then an evaluation in which we compare recovery times against the baseline. If you then wish to place maintenance with us, that is possible, but it is not a requirement.

Frequently asked questions about observability and monitoring.

What development teams, IT managers and those responsible for availability usually ask us before such a project begins.

What is the difference between monitoring and observability?
Monitoring answers questions you thought of in advance: is the service running, is the disk full, how many requests came in. Observability is about the questions you only think of during an incident, and which you must then be able to answer without new code. The difference lies mainly in the richness of what you capture: enough context per event to filter afterwards on something you never built a dashboard for. In practice you need both, and most organisations have the first but not the second.
What exactly do you mean by the four signals?
Latency, error rate, traffic and saturation. Together, these describe whether your service is doing its job and whether it is under strain. The appealing thing about them is that the first three concern your users rather than your infrastructure. A disk that is eighty per cent full is not a problem; a latency that doubles is. Alerting on the first category gives you nights when nothing was actually wrong, and that is precisely how alert fatigue sets in.
Will you replace our current platform?
Usually not, and we rarely recommend it. Grafana, Datadog, Elastic, New Relic and your cloud provider's services do their job well; the problem almost always lies in what your software supplies to them. We set up that side. If you do want to switch, instrumenting with OpenTelemetry is the sensible order: it makes your telemetry independent of the platform, so the later move becomes a configuration matter rather than a second rebuild.
What does storage cost, and how do we keep it under control?
The bill for observability is almost entirely driven by how much you keep and for how long. The approach that holds up is to choose per type of data. Metrics are small and can be retained for a long time, since they show how something evolves over months. Log entries are large and most useful in the days that follow. Traces are the largest, and a sample is usually enough, with the agreement that anything that fails is kept in full. Together these three choices typically cut costs by a multiple, without you losing anything.
Does this also work if our software isn't made up of separate services?
Yes, and it's simpler. With a single application, the gain lies less in tracing and more in bringing together what you already have: metrics per function, log entries with enough context to filter on, and alerting tied to the service. We set that up in the same way; only the step across system boundaries falls away. If you later grow into multiple services, the foundation is already in place.
How does this fit with our security?
The same record that helps you during an outage is what you need to reconstruct what happened after a security incident. That isn't an extra project but the same work with a second benefit, and it ties directly into vulnerability management. We do pay attention to what ends up in log entries: personal data and secrets don't belong there, and that is a choice you make when instrumenting, not something you repair afterwards.
Who on our side should take this on?
Someone who is on duty when something breaks. Observability set up by someone other than the people who work with it is rarely used: the dashboards then answer the questions the builder had, not the ones that come up at three in the morning. We therefore work together with your own developers and operations staff, and the handover is part of the project rather than an appendix at the end.
How quickly will we notice a difference?
The first gains usually come from bringing together what you already have and from tidying up the alerting; you'll notice that within the first sprints. Tracing across system boundaries requires changes to the software and takes longer, service by service. We deliberately start with the service where the most time is lost, so that the first delivery is also the most noticeable.
Can you do this for purchased software too?
For the part we can reach, yes: incoming and outgoing traffic, availability, call latency, and the queues in between. What happens inside such a package depends on what the vendor makes available. In practice that is enough to establish whether a problem lies with you or with them, which is usually exactly the question you want to be able to answer.
What is an error budget, and do we need one?
An error budget is the gap between your target and one hundred per cent. If you agree that ninety-nine point nine per cent of requests must succeed, that last tenth of a per cent is your budget. As long as it isn't used up, new functionality can go out; once it runs out, attention goes first to stability. That is useful once there is debate between building on and repairing, and superfluous as long as that debate doesn't exist. We recommend it for the third option, not the first.
Do you work with OpenTelemetry or with something from the vendor?
Standard with OpenTelemetry. That is the open standard for metrics, log entries and traces, and almost any platform can read it. The advantage is that the instrumentation in your code is not tied to a single vendor: if you switch platforms, you simply point the stream elsewhere. Only where you are already deeply committed to a vendor-specific approach, with no plans to switch, might it be wiser to build on that.

Talk to us about your monitoring.

A free, no-obligation half-hour introductory call. Tell us what last went wrong and how long it took before anyone knew where the problem lay, and we'll tell you where we would start, even if we don't end up working together.

Edit content