Fraud detection is one of those problems that looks deceptively simple from the outside. Someone makes a payment, you decide whether it is legitimate, and you either let it through or stop it. In practice, the gap between that sentence and a production system that protects real money while not strangling legitimate customers is enormous. I have built and operated these pipelines at several stages of company growth, and the lessons rarely come from clever models. They come from the unglamorous plumbing around them.
This post is a walkthrough of how I think about designing a fraud detection pipeline in a regulated payments environment. It is not a survey of algorithms. It is the architecture, the trade-offs, and the operational discipline that determine whether a fraud system earns its keep or quietly becomes a liability that nobody understands well enough to change.
What the Pipeline Actually Protects
Before writing a line of code, I insist the team agree on what we are defending against and what failure costs. Fraud is not a single phenomenon. Stolen card credentials, account takeover, synthetic identities, friendly fraud chargebacks, and first-party abuse all behave differently and demand different signals. A pipeline tuned for one will be blind to the others, so naming the threat model is the first design decision, not an afterthought.
The economics matter just as much. Every blocked transaction has a false positive cost: a customer who walks away, a support ticket, sometimes a churned account. Every missed fraud has a direct loss plus chargeback fees and, past a threshold, scheme fines that can put your acquiring relationship at risk. I make the team write down rough dollar values for both error types early, because those numbers set the operating point of every model and rule downstream. Without them, you are optimising an abstract metric instead of a business outcome.
Finally, I treat the pipeline as a regulated component from day one. Decisions that decline a customer may need to be explainable, retained, and auditable. That constraint shapes the architecture more than any accuracy target, because a model you cannot explain or reproduce is a model you cannot defend to a regulator or a court.
Ingesting Signals at the Edge
A fraud decision is only as good as the data available at the moment it is made. The hard part is that the richest signals arrive from different systems with wildly different latencies. The authorization request itself is synchronous and gives you amount, merchant, currency, and card metadata. Device fingerprints and behavioural telemetry come from the client. Historical aggregates live in your data warehouse. Bringing these together at decision time without blowing the latency budget is the core engineering challenge.
I separate signals into three tiers by how fresh they need to be. Real-time signals are computed inline during the request. Near-real-time signals are maintained in a streaming layer with a few seconds of lag. Batch signals are precomputed daily and looked up by key. Forcing every feature into the real-time path is the most common way teams destroy their latency budget and their on-call sanity.
- Real-time: transaction amount, velocity counters in the last few minutes, geolocation mismatch against the issuing country.
- Near-real-time: session behaviour, device reputation, rolling spend over the last hour computed from a stream.
- Batch: customer lifetime value, historical chargeback rate, account age, long-window spending patterns.
The Feature Store as the Source of Truth
The single most valuable piece of infrastructure I have built for fraud is a proper feature store, and the reason is consistency rather than convenience. Training a model on features computed one way and then serving with features computed another way is the classic training-serving skew that silently degrades every model in production. A feature store enforces that the same definition produces the same value offline and online.
Concretely, I want every feature to have a single declarative definition, a backfill path from historical data, and an online serving path that returns the value with single-digit millisecond latency. Velocity features are the ones that bite teams most often, because the boundary conditions of the time window are easy to get subtly wrong. Here is the kind of windowed aggregate I expect to see defined once and reused everywhere.
SELECT
card_hash,
COUNT(*) AS txn_count_10m,
SUM(amount_minor) / 100.0 AS total_amount_10m,
COUNT(DISTINCT merchant_id) AS distinct_merchants_10m
FROM transactions
WHERE card_hash = @cardHash
AND created_at > (NOW() - INTERVAL '10 minutes')
AND status IN ('authorized', 'pending')
GROUP BY card_hash;
Note that we never store raw card numbers; we operate on a salted hash so the fraud system stays out of cardholder-data scope wherever possible. That single decision saves enormous compliance pain and is far easier to enforce at the feature layer than to retrofit later.
Rules and Models Are Partners, Not Rivals
There is a persistent myth that machine learning replaces rules. In every system I have run, the two coexist permanently, and trying to eliminate either one is a mistake. Rules are deterministic, explainable, and instantly deployable, which makes them perfect for hard policy constraints and for reacting to a fraud attack you discovered an hour ago. A model cannot be retrained and validated in an hour; a rule can be live in minutes.
Models, on the other hand, capture interactions across dozens of features that no analyst could encode by hand, and they degrade gracefully rather than failing at the edges of their conditions. I architect the pipeline so a gradient-boosted model produces a calibrated risk score, and a rules engine sits on top to apply hard constraints, regulatory blocks, and rapid-response patches. The model informs; the rules layer makes the policy explicit and auditable.
Enjoying this article?
Get more like it in your inbox — practical engineering leadership, fintech, and AI. No spam, unsubscribe anytime.
The model tells you how risky a transaction is. The rules tell you what your company has decided to do about risk. Conflating those two responsibilities is how fraud teams lose control of their own decisions.
Scoring Within the Latency Budget
An authorization decision typically has to return in well under a few hundred milliseconds, and the fraud check is only one part of that budget. This constraint disciplines the entire design. I treat the scoring path as a hot service that does feature lookups in parallel, runs inference on a model loaded in memory, and applies the rules layer, with an explicit timeout on every external call.
The most important pattern here is graceful degradation. If a feature lookup times out, the system must not block the request indefinitely or crash. It returns a default value flagged as missing, the model handles missingness explicitly, and the decision proceeds with slightly less information. A fraud system that fails closed on every dependency hiccup will decline good customers during a partial outage, which is often more damaging than the fraud it prevents.
public async Task<FraudDecision> ScoreAsync(AuthRequest request, CancellationToken ct)
{
using var cts = CancellationTokenSource.CreateLinkedTokenSource(ct);
cts.CancelAfter(TimeSpan.FromMilliseconds(80));
var features = await _featureStore
.GetOnlineFeaturesAsync(request.CardHash, cts.Token)
.ConfigureAwait(false);
var score = _model.Predict(features); // in-memory, sub-millisecond
var decision = _rules.Apply(request, features, score);
_audit.Record(request.Id, features, score, decision);
return decision;
}
Notice that we record the features, the score, and the decision together for every single request. That audit record is not optional logging; it is the foundation of everything in the next two sections.
The Feedback Loop Is the Product
The hardest truth about fraud systems is that your labels arrive late and incomplete. A chargeback can land sixty days after the transaction. A confirmed fraud report comes through a manual review queue. A blocked transaction that was actually legitimate may never be labelled at all, because the customer simply gave up. This delayed and biased labelling is what makes fraud different from most supervised learning problems.
I invest heavily in the labelling pipeline because the model is only as good as the ground truth feeding it. That means wiring chargeback feeds, manual review outcomes, and customer dispute resolutions back into the same store that holds the original decision and its features. Without that join, you cannot measure whether a model change helped or hurt, and you are flying blind.
The selection bias problem deserves explicit attention. If you only ever observe outcomes for transactions you approved, your training data is censored, and the model becomes overconfident about the population it never let through. I mitigate this with a small, deliberately approved holdout of borderline transactions, accepting a known and bounded fraud cost in exchange for unbiased labels. It feels uncomfortable to approve transactions you suspect, but the alternative is a model that slowly poisons itself.
Monitoring for Drift and Attack
Fraud is adversarial in a way that ordinary machine learning is not. The moment your model becomes effective against a pattern, motivated attackers adapt, and the distribution of inputs shifts deliberately to evade you. This means a fraud model is never finished, and monitoring is not a nice-to-have; it is the early warning system that tells you the ground has moved under you.
I monitor three distinct things, and I have learned the hard way that conflating them hides problems. First, the input feature distributions, to catch data pipeline breakage and deliberate evasion. Second, the score distribution and approval rate, to catch sudden shifts in behaviour. Third, the realised outcomes as labels mature, to catch genuine performance decay. A spike in the approval rate with no corresponding feature drift usually means a rule was misconfigured, not that fraud disappeared.
When an alert fires, the team needs to act in minutes, not days. That is exactly why the rules layer exists alongside the model. A new attack pattern gets a targeted rule within the hour to stop the bleeding, while the longer model retraining cycle catches up. The rule is the tourniquet; the retrain is the surgery.
Governance and Explainability
Every automated decision that affects a customer needs to be reproducible and explainable, and in a regulated environment this is a legal requirement rather than a courtesy. When a customer disputes a decline, I need to reconstruct exactly what the system knew at that moment: the feature values, the model version, the score, and which rule, if any, drove the outcome. This is why the audit record captures all of it atomically at decision time.
Explainability also shapes which models I am willing to deploy. I lean toward models whose decisions can be attributed to specific features using stable attribution methods, because I would rather give up a fraction of a point of accuracy than be unable to explain a decline to a regulator or a customer. Versioning is part of this discipline too: every model and every rule set is versioned, and the version is stamped on each decision so we can always answer the question of what logic produced a given outcome.
- Reproducibility: features, model version, and rule set version stored with every decision.
- Attribution: the top contributing features available for any individual decision.
- Retention: decision records held for the regulatory window, with access controls and immutability.

Conclusion
A fraud detection pipeline is not a model with some glue around it. The model is perhaps a fifth of the work and rarely the part that determines success. The pipeline lives or dies on consistent features, honest labelling, graceful degradation under partial failure, a rules layer that lets you respond to attacks in minutes, and an audit trail that makes every decision reproducible and defensible. Build those well, and an ordinary model performs admirably; build them badly, and the best model in the world will quietly mislead you. The discipline that protects the money is the engineering around the edges, and that is exactly where I spend most of my attention.
Get new posts in your inbox
Occasional, practical notes on engineering leadership, fintech, and building with AI. No spam, unsubscribe anytime.
Comments (9)
Leave a Comment
Efua Darko
September 2, 2026
Backend eng in Accra, mostly working on remittance flows. We hit this exact thing with MTN MoMo API rate limits last Jan — tail latency was the presenting symptom, and the "Scoring Within the Latency Budget" section is basically how we untangled it. Ended up adding a shadow queue on Python + FastAPI, cut error rate by 59%. For context: 175k active users.
Kevin Wilson
August 31, 2026
Quick q on "Rules and Models Are Partners, Not Rivals" — how do you handle partial failures when the settlement window drifts out of sync? We're on whatever the seniors set up and our current answer is jitter-and-pray.
Katie Miller
August 25, 2026
Quick q on "Governance and Explainability" — how do you handle partial failures when the partner API drifts out of sync? We're on Envoy and the sidecar reconciler feels overkill for our scale.
James Pemberton
August 16, 2026
Does the "Monitoring for Drift and Attack" still hold on a 7-engineer team? We're at the awkward middle and some of these patterns feel like they need a dedicated ops person to run properly.
Brandon Rodriguez
August 5, 2026
Small pushback: the "What the Pipeline Actually Protects" advice maps cleanly onto high-throughput consumer payments, less so onto regulated custody where operators actively want the slower, more explicit path. In our platform we ended up doing the opposite and it's been the right call.
Adeola Ogunleye
August 3, 2026
This is why I keep coming back to this blog.
Nicole Lee
August 2, 2026
Good topic. Staff SWE at a post-IPO neobank in Chicago here. What we do differently: run a shadow processor comparing against production for a week before cut-over on gRPC. It is not universally better; ops needed six weeks to warm up to it, but the testability is dramatically better and that pays for itself the first time you have to answer a FCA question at 2am.
Musa Aliyu
August 1, 2026
Good topic. Fractional CTO — done 4 fintech engagements in the last 5 years here. What we do differently: run a shadow processor comparing against production for a week before cut-over on Terraform. It is not universally better; ops needed six weeks to warm up to it, but the audit story is dramatically better and that pays for itself the first time you have to answer a CBN question at 4am.
Jennifer Brown
July 29, 2026
Good writeup. One nit on "What the Pipeline Actually Protects": worth mentioning DLQ replay tooling — otherwise the pattern degrades under real load.
