For the first decade of my career, observability meant one thing: logs. When something broke, you SSH'd into a box, tailed a file, and grepped for the word "error" until you found the line that explained your bad afternoon. That model worked when we ran a monolith on three servers and a payment failed once a week. It does not work when you operate a distributed payments platform where a single card authorization touches a dozen services, two message brokers, and three external providers before it returns a yes or a no.
I want to talk about what observability becomes once you accept that logs alone cannot answer the questions that actually matter in a regulated fintech environment. Not because logging is wrong, but because it is one signal among three, and treating it as the whole picture is how teams end up staring at a dashboard that says everything is fine while customers cannot move money.
Why Logs Alone Stopped Being Enough
A log line is a discrete statement about a moment in time. It tells you that a thing happened, and if the engineer who wrote it was disciplined, it tells you a little about the context. The trouble is that a payment is not a moment; it is a journey across process boundaries. When a settlement run stalls, the question is rarely "did this one function throw?" It is "where, across forty steps and four services, did the latency accumulate?" Logs are poorly shaped to answer that, because each line lives in isolation and the correlation between them is something you have to reconstruct by hand.
The second problem is volume. At any meaningful transaction rate, you cannot afford to log everything at full fidelity, so you sample or you summarize. The moment you sample logs, you have quietly accepted that the one event you most need during an incident may be the one you threw away. I have lived through a postmortem where the answer sat in a debug log disabled in production for cost reasons.
So the reframing I push on every team I lead is simple: logs are for narrative detail after you already know roughly where to look. They are not the instrument you use to know where to look in the first place. That job belongs to metrics and traces.
The Three Pillars, Without the Marketing
The "three pillars" framing — metrics, traces, logs — gets repeated so often that it has lost meaning. Let me ground it in how I actually use each one. Metrics answer "is something wrong, and how wrong?" They are cheap, aggregable, and they drive alerts. Traces answer "where is it wrong?" They show the path of a single request across services and where the time went. Logs answer "why is it wrong?" They carry the specific values, stack traces, and business context that explain the failure once you have localized it.
The discipline is to use them in that order. An alert fires off a metric. You pivot from the metric to the traces that share the same time window and labels. You find the slow or failing span, and only then do you drill into the logs attached to it. Done well, an investigation that used to take an hour of grepping becomes three clicks.
The metric tells you the building is on fire. The trace tells you which floor. The log tells you which wastebasket. You need all three, but you do not start by inspecting every wastebasket.
Correlation, or It Might As Well Not Exist
None of this works if the three signals cannot be stitched together. The single highest-leverage investment a fintech engineering team can make in observability is propagating a correlation identifier through every layer of the request path and emitting it on every signal. A trace ID that flows from the API gateway through the authorization service, into the ledger, out to the card network adapter, and back is worth more than any amount of clever log formatting.
In .NET the framework now does most of this for you through the Activity API and OpenTelemetry, but you still have to make deliberate choices about what business context to attach. A trace ID is necessary; it is not sufficient. For a payment, I want the masked merchant ID, the transaction type, the currency, and the processor on the span, so that I can filter and group by the dimensions that matter to the business, not just the dimensions that matter to the runtime.
using System.Diagnostics;
public sealed class PaymentInstrumentation
{
private static readonly ActivitySource Source =
new("Payments.Authorization", "1.0.0");
public async Task<AuthResult> AuthorizeAsync(AuthRequest request)
{
using var activity = Source.StartActivity("authorize.card");
// Attach business context, not just runtime context.
activity?.SetTag("payment.type", request.Type);
activity?.SetTag("payment.currency", request.Currency);
activity?.SetTag("payment.processor", request.Processor);
activity?.SetTag("merchant.id", Mask(request.MerchantId));
try
{
var result = await _processor.SendAsync(request);
activity?.SetTag("payment.decision", result.Decision);
if (!result.Approved)
{
activity?.SetStatus(ActivityStatusCode.Error, result.DeclineReason);
}
return result;
}
catch (Exception ex)
{
activity?.SetStatus(ActivityStatusCode.Error, ex.Message);
throw; // The exception is logged downstream, linked by trace id.
}
}
}
The point of the code above is not the syntax; it is the decision to tag spans with domain-meaningful attributes. When you do this consistently, you can answer questions like "what is the p99 latency for cross-border card authorizations through one specific processor in the last hour" without writing an ad hoc query against raw logs.
Cardinality and the Cost Trap
There is a hard engineering constraint hiding inside that last paragraph, and ignoring it is how observability budgets explode. Metrics are cheap precisely because they are low cardinality. The moment you attach a unique value — a transaction ID, a customer ID — to a metric dimension, you create a new time series for every distinct value, and your monitoring bill grows accordingly. I have seen a well-intentioned engineer add a customer identifier to a Prometheus label and triple the storage cost overnight.
The rule I enforce is a clean division of labor by cardinality:
- Metrics carry only bounded, low-cardinality dimensions: service name, endpoint, status class, processor, currency. Things with a known, small set of values.
- Traces carry high-cardinality detail through tags, because traces are sampled and stored differently — a unique transaction ID belongs here, not on a metric.
- Logs carry the unbounded specifics: full request payloads (redacted), exception detail, and the verbose context you only read after you have localized the problem.
Get this division wrong and one of two failures follows. Either you blow the budget by putting high-cardinality data on metrics, or you cripple investigations by leaving it off your traces. The skill is knowing which signal each piece of context belongs to before you write the instrumentation.
SLOs as the Language of Priority
Raw telemetry without objectives is just noise with good production values. The thing that turns observability into an operational discipline is defining service level objectives and measuring against them. For a payments API, my objective is rarely "uptime." It is something like "99.95 percent of authorization requests complete in under 800 milliseconds with a valid decision." That sentence encodes availability, latency, and correctness in one measurable target, and it gives the team a shared definition of "fine."
SLOs also give you an error budget, which is the most useful management tool I know. If the budget says we can afford forty-three minutes of degraded service this month and we have burned thirty of them by the tenth, that is a data-driven signal to slow down and stabilize rather than ship the next feature. It moves the conversation away from gut feeling and toward arithmetic.
Crucially, the SLO is what your alerts should be built on. Do not page a human because CPU hit eighty percent; page them because the error budget is burning fast enough to breach the objective. The first wakes people for nothing; the second for something a customer can feel.
Alerting on Symptoms, Not Causes
The fastest way to destroy an on-call rotation is to alert on causes instead of symptoms. A cause-based alert says "the database connection pool is at ninety percent." Maybe that matters, maybe it does not — the pool might recover in two seconds with zero customer impact. A symptom-based alert says "the rate of failed authorizations has exceeded our error budget burn rate." That always matters, because by definition it is something the customer experiences.
When you alert on symptoms, your traces and metrics become the tools you use to find the cause after the page fires, rather than a wall of pre-emptive warnings that train your engineers to ignore the pager. I would rather have five well-chosen, symptom-based alerts that are always actionable than two hundred cause-based ones that everyone mutes by their third week on call. Alert fatigue is not a personality flaw; it is a design failure in your observability stack.
This is also where the regulated context tightens the screws. In a payments business, a degraded experience is not only an engineering problem; depending on severity and duration it can be a reportable incident. Your alerting needs to distinguish the merely annoying from the genuinely material.
Where Observability Meets Compliance
In an unregulated startup, observability is an engineering convenience. In a regulated fintech, it is partly a control. Auditors and regulators ask questions that telemetry is uniquely positioned to answer: can you demonstrate that this transaction was processed within your stated timeframe, that this customer's data was accessed only by authorized services, that this failed payment was retried according to policy? A well-instrumented system answers those questions from data you already collect.
This raises the stakes on two things. First, retention: the time horizon over which you keep observability data is now a compliance decision, not purely a cost decision, and the two pull in opposite directions. Second, redaction: traces and logs are a notorious leak path for sensitive data. A card number that ends up in a span tag because someone logged the whole request object is a real incident waiting to happen. I treat redaction as a first-class part of the instrumentation pipeline, enforced in shared libraries rather than left to the discipline of individual engineers.
If your observability data cannot survive an audit, it is a liability dressed up as an asset. Build redaction and retention into the pipeline, not into the good intentions of whoever wrote the last log statement.
Making It a Habit, Not a Project
The hardest part of all this is not technical. It is cultural. Observability decays the moment it stops being part of how you write code. A team that bolts telemetry on at the end of a project produces telemetry that looks plausible and answers nothing, because the instrumentation reflects what was easy to add rather than what you will actually need at three in the morning.
The practices that keep it alive are unglamorous. Instrumentation is reviewed in pull requests with the same rigor as the logic it measures. New services inherit a shared library that wires up tracing, metric conventions, and redaction by default, so the easy path is also the correct path. And every incident review asks one standing question before anything else: "could we have seen this sooner, and what signal did we lack?" The answer to that question is the backlog for the next quarter of observability work.
When I interview senior engineers, I ask how they would debug a problem they cannot reproduce. The ones who reach immediately for logs are thinking about yesterday's systems. The ones who talk about tracing the request path, narrowing by metric, and only then reading the logs are thinking about the systems we actually run.

Conclusion
Observability beyond logging is not a tooling purchase; it is a shift in how you reason about a running system. Logs remain essential, but they are the last instrument you reach for, not the first. Metrics tell you that something is wrong and how wrong, traces tell you where, and logs tell you why — and the entire value comes from being able to move between them along a shared correlation identifier. Layer service level objectives on top so that you alert on what customers feel rather than on what your infrastructure happens to be doing, and treat redaction and retention as the compliance controls they genuinely are. Do that, and the three-in-the-morning page stops being an archaeology dig and becomes a guided walk to the one thing that broke. In a business where the product is trust, that difference is not a nicety. It is the job.
