What is the difference between monitoring and observability?
Monitoring answers questions you thought of in advance: is the service running, is the disk full, how many requests came in. Observability is about the questions you only think of during an incident, and which you must then be able to answer without new code. The difference lies mainly in the richness of what you capture: enough context per event to filter afterwards on something you never built a dashboard for. In practice you need both, and most organisations have the first but not the second.
What exactly do you mean by the four signals?
Latency, error rate, traffic and saturation. Together, these describe whether your service is doing its job and whether it is under strain. The appealing thing about them is that the first three concern your users rather than your infrastructure. A disk that is eighty per cent full is not a problem; a latency that doubles is. Alerting on the first category gives you nights when nothing was actually wrong, and that is precisely how alert fatigue sets in.
Will you replace our current platform?
Usually not, and we rarely recommend it. Grafana, Datadog, Elastic, New Relic and your cloud provider's services do their job well; the problem almost always lies in what your software supplies to them. We set up that side. If you do want to switch, instrumenting with OpenTelemetry is the sensible order: it makes your telemetry independent of the platform, so the later move becomes a configuration matter rather than a second rebuild.
What does storage cost, and how do we keep it under control?
The bill for observability is almost entirely driven by how much you keep and for how long. The approach that holds up is to choose per type of data. Metrics are small and can be retained for a long time, since they show how something evolves over months. Log entries are large and most useful in the days that follow. Traces are the largest, and a sample is usually enough, with the agreement that anything that fails is kept in full. Together these three choices typically cut costs by a multiple, without you losing anything.
Does this also work if our software isn't made up of separate services?
Yes, and it's simpler. With a single application, the gain lies less in tracing and more in bringing together what you already have: metrics per function, log entries with enough context to filter on, and alerting tied to the service. We set that up in the same way; only the step across system boundaries falls away. If you later grow into multiple services, the foundation is already in place.
How does this fit with our security?
The same record that helps you during an outage is what you need to reconstruct what happened after a security incident. That isn't an extra project but the same work with a second benefit, and it ties directly into
vulnerability management. We do pay attention to what ends up in log entries: personal data and secrets don't belong there, and that is a choice you make when instrumenting, not something you repair afterwards.
Who on our side should take this on?
Someone who is on duty when something breaks. Observability set up by someone other than the people who work with it is rarely used: the dashboards then answer the questions the builder had, not the ones that come up at three in the morning. We therefore work together with your own developers and operations staff, and the handover is part of the project rather than an appendix at the end.
How quickly will we notice a difference?
The first gains usually come from bringing together what you already have and from tidying up the alerting; you'll notice that within the first sprints. Tracing across system boundaries requires changes to the software and takes longer, service by service. We deliberately start with the service where the most time is lost, so that the first delivery is also the most noticeable.
Can you do this for purchased software too?
For the part we can reach, yes: incoming and outgoing traffic, availability, call latency, and the queues in between. What happens inside such a package depends on what the vendor makes available. In practice that is enough to establish whether a problem lies with you or with them, which is usually exactly the question you want to be able to answer.
What is an error budget, and do we need one?
An error budget is the gap between your target and one hundred per cent. If you agree that ninety-nine point nine per cent of requests must succeed, that last tenth of a per cent is your budget. As long as it isn't used up, new functionality can go out; once it runs out, attention goes first to stability. That is useful once there is debate between building on and repairing, and superfluous as long as that debate doesn't exist. We recommend it for the third option, not the first.
Do you work with OpenTelemetry or with something from the vendor?
Standard with OpenTelemetry. That is the open standard for metrics, log entries and traces, and almost any platform can read it. The advantage is that the instrumentation in your code is not tied to a single vendor: if you switch platforms, you simply point the stream elsewhere. Only where you are already deeply committed to a vendor-specific approach, with no plans to switch, might it be wiser to build on that.