Every distributed payments platform I have built eventually arrives at the same uncomfortable truth: messages get delivered more than once. Not occasionally, not as an edge case to be apologized for in a postmortem, but as a routine, expected, structural fact of the systems we operate. Brokers retry. Networks partition. Consumers crash after doing work but before acknowledging it. The moment you accept at-least-once delivery as the contract you actually have, the question stops being "how do I prevent duplicates" and becomes "how do I make duplicates harmless."
That reframing is the whole game. Idempotency is not a feature you sprinkle on at the end; it is a design discipline that shapes your schema, your message contracts, your transaction boundaries, and your operational runbooks. In regulated fintech, where a doubled ledger entry is not a glitch but a reconciliation incident with audit implications, getting this right is the difference between sleeping through a broker hiccup and explaining a balance discrepancy to a compliance officer. Here is how I think about building consumers that can safely see the same event twice.
Why At-Least-Once Is the Default You Actually Have
Most teams reach for a message broker expecting exactly-once delivery, and most brokers advertise something that sounds like it. The honest reading is that end-to-end exactly-once is a property of the consumer's processing, not of the transport. Kafka, RabbitMQ, SQS, Azure Service Bus, EventBridge: every one of them, under failure, will redeliver. The broker cannot know whether your consumer finished its side effects before the acknowledgment was lost on the wire. So it does the safe thing and sends again.
I have seen the alternative chased at great cost. Teams build elaborate deduplication into the transport layer, lean on broker-side "exactly-once" semantics that only hold within a single partition or session, and then discover the guarantee evaporates the instant a message crosses a system boundary into a database write, an outbound HTTP call, or a second topic. The guarantee you can actually enforce lives in your code and your storage, close to the side effect itself.
So my first principle is to design as though every consumer will be invoked with the same logical event an unbounded number of times. If that assumption holds and the system stays correct, you have an idempotent consumer. If it does not, you have a latent reconciliation bug waiting for a bad afternoon.
Natural Keys Versus Synthetic Idempotency Keys
Idempotency requires identity. To recognize a duplicate, you need a stable key that means "this exact business event" and not merely "a message that arrived." Choosing that key well is the most consequential decision in the whole exercise, and it is where most designs quietly go wrong.
Sometimes the domain hands you a natural key. A card authorization has an issuer-assigned reference. A bank transfer carries an end-to-end identifier that survives the whole payment chain. When such a key exists and is genuinely unique per business intent, use it. When it does not, the producer must mint a synthetic idempotency key at the moment of intent and carry it through every message and retry. The cardinal sin is generating the key at publish time inside a retry loop, because then each retry invents a fresh key and your deduplication never fires.
The idempotency key must be born with the intent, not with the message. If a client retries a request, it must send the same key; if a producer republishes, it must reuse the original. A key created at transmission time deduplicates nothing.
I push this requirement all the way to the edge. Our public API requires an Idempotency-Key header on every mutating request, and we propagate that value into the event envelope so the same identity flows from the customer's first click through the broker and into the consumer's dedup table. One identity, end to end, regardless of how many times any hop retries.
The Deduplication Store and Its Trade-offs
Once you have a key, you need somewhere to remember which keys you have already processed. The naive version is a table of seen keys that you check before doing work. The subtle version recognizes that the check and the work must be atomic with respect to each other, or two concurrent deliveries both read "not seen," both proceed, and you have doubled the side effect anyway.
I strongly prefer to make the dedup record and the business effect share a single transaction in the same database. That way the unique constraint on the key is the concurrency guard, and the database does the hard part for free. The alternatives each have costs worth naming:
- Same-database dedup table: strongest guarantee, because insert and effect commit or roll back together; limited to side effects that live in that database.
- External store such as Redis: fast and shared across services, but now you have two systems that can disagree, and you must reason about what happens when the cache write succeeds and the database write does not.
- Broker-native dedup windows: convenient, but bounded by a time window and a single broker, so they protect against fast redelivery and nothing else.
For anything touching money, I default to the same-database approach and treat external caches as an optimization in front of it, never as the source of truth. The store of record for "did this already happen" should be as durable as the effect it guards.
Pairing Consumers With a Transactional Outbox
Idempotent consumers solve the receiving side, but they only matter if the sending side is honest about what it published. The classic failure is a service that updates its database and then publishes an event, with a crash in between leaving the two out of sync. The transactional outbox pattern closes that gap and pairs naturally with idempotent consumers.
The producer writes its state change and an outbox row in one local transaction. A separate relay reads the outbox and publishes to the broker, marking rows as sent. Because the relay can crash and replay, it will sometimes publish the same outbox row twice, which is exactly why the downstream consumer must be idempotent. The two patterns are complements: the outbox guarantees at-least-once publication of every committed change, and the idempotent consumer guarantees that at-least-once does no harm.
This is the architecture I reach for by default in our settlement and ledger services. It is unglamorous, it adds a table and a relay process, and it has saved us from an entire category of "the database says one thing and the event stream says another" incidents that are miserable to debug after the fact.
A Concrete Consumer Shape
Patterns are easier to trust when you can see the code. Below is the shape I use in our .NET services: a single transaction that inserts the idempotency record and performs the business effect together, relying on a unique constraint to reject duplicates. The catch on the unique-violation turns a duplicate into a no-op rather than an error.
Enjoying this article?
Get more like it in your inbox — practical engineering leadership, fintech, and AI. No spam, unsubscribe anytime.
public async Task HandleAsync(PaymentCapturedEvent evt, CancellationToken ct)
{
await using var tx = await _db.BeginTransactionAsync(ct);
try
{
// The dedup row and the effect commit together.
await _db.ExecuteAsync(
"INSERT INTO processed_events (idempotency_key, event_type, processed_at) " +
"VALUES (@key, @type, @now)",
new { key = evt.IdempotencyKey, type = "PaymentCaptured", now = DateTime.UtcNow });
await _ledger.PostCaptureAsync(evt.AccountId, evt.Amount, evt.Currency, ct);
await tx.CommitAsync(ct);
}
catch (UniqueConstraintViolationException)
{
// We have seen this key before. Roll back and treat as success.
await tx.RollbackAsync(ct);
_log.LogInformation("Duplicate event {Key} ignored", evt.IdempotencyKey);
}
}
The important detail is that the insert and the ledger post are in the same transaction against the same database. If the ledger post fails, the dedup row rolls back too, so a genuine failure is retryable rather than being permanently marked as done. And because the unique index on idempotency_key is enforced at commit, two concurrent workers racing on the same event will see exactly one winner; the loser catches the violation and exits cleanly.
Handling Side Effects You Cannot Roll Back
The transactional model breaks down the moment a side effect leaves your database: sending an email, calling a third-party payment processor, debiting an external account. You cannot enroll a partner bank's API in your local transaction. So you need a discipline for effects that are not undoable.
My approach is to split such work into two committed steps with a durable intent record in between. First, in the protected transaction, I record the intent and the idempotency key and mark it pending. Then I perform the external call, ideally passing my own idempotency key to the downstream provider so that they deduplicate on their side as well. Finally I record the outcome. If we crash mid-flight, replay finds a pending intent and can query the provider to discover whether the call actually went through before deciding to retry.
The practical lesson is to push idempotency keys across system boundaries wherever the counterparty supports them. Most serious payment processors do, precisely because they live with the same redelivery reality we do. When the counterparty does not, you fall back to a status query before retrying, which is slower and more code but keeps you from sending money twice.
Ordering, Stale Events, and Out-of-Order Delivery
Deduplication answers "have I seen this exact event," but it does not answer "should I apply this event now." Redelivery often arrives out of order, and an old event replayed after a newer one can clobber fresh state if you only check for duplicates. Idempotency and ordering are separate problems that people conflate at their peril.
For state that evolves, I add a monotonic version or sequence to the entity and reject any event whose version is not greater than what I have already applied. This makes the consumer not only idempotent but also resistant to stale replays: a reprocessed status-update from five minutes ago simply loses to the version check. For purely additive effects like ledger postings, the idempotency key alone suffices because each posting is its own immutable fact rather than an overwrite.
The distinction I keep front of mind is between events that set state and events that append facts. Append-style consumers need identity. Overwrite-style consumers need identity and order. Conflating the two leads to consumers that correctly ignore duplicates while happily applying a stale update, which is arguably worse than the duplicate would have been.
Retention, Cleanup, and Operational Reality
A dedup table grows forever if you let it, and an unbounded table eventually becomes a performance problem on the exact write path you most care about. So retention is part of the design, not an afterthought. I keep processed-event records for a window comfortably longer than the broker's maximum possible redelivery delay, then archive or purge older rows in a background job.
Choosing that window requires knowing your broker's behavior under its worst case, including dead-letter replays and manual reprocessing that operators might trigger weeks later. If your runbook allows an engineer to replay a day-old dead-letter queue, your dedup retention must outlast that possibility, or the replay will create duplicates the table can no longer recognize. I would rather over-retain and pay for storage than discover a gap during an incident.
Operationally, I also instrument the duplicate path. Every ignored duplicate is logged and counted as a metric, because a sudden spike in duplicates is an early signal that something upstream is misbehaving, even though the consumer is correctly absorbing it. Idempotency that silently swallows duplicates without telling anyone hides the very signal you would want during an incident.
Testing the Guarantee You Claim to Have
An idempotency guarantee you have not tested is a guarantee you do not have. The cheap, dangerous version of testing is to send a message twice in sequence and confirm one effect; that proves the happy path and nothing about concurrency or crash recovery. The failures that hurt are the ones that happen under contention and partial failure.
I test three scenarios deliberately. First, concurrent delivery of the same key from multiple workers, asserting exactly one effect and one winner. Second, a crash injected between the side effect and the acknowledgment, then a replay, asserting no doubled effect. Third, out-of-order delivery, asserting the version check rejects the stale event. These tests are not glamorous, but they exercise the exact conditions under which production breaks.
The mindset I try to instill in the team is adversarial: assume the broker is actively trying to deliver every message twice at the worst possible moment, and write the test that proves you survive it. If the test is hard to write, that difficulty is usually telling you the design is not as idempotent as you hoped.

Conclusion
Idempotent consumers are less about a clever trick and more about accepting reality and designing around it. At-least-once delivery is the contract you have whether you acknowledge it or not, so the durable move is to make duplicates boring: a stable identity born with the intent, a dedup record that commits in the same transaction as the effect, an outbox to keep producers honest, a version check to defend against stale replays, and tested behavior under concurrency and crash. Do those things and a broker hiccup becomes a non-event in your logs rather than a discrepancy on someone's ledger. In regulated fintech, that quiet correctness is worth far more than any guarantee printed on a transport's data sheet.
Get new posts in your inbox
Occasional, practical notes on engineering leadership, fintech, and building with AI. No spam, unsubscribe anytime.
Comments (10)
Leave a Comment
Amanda Anderson
September 2, 2026
23 months into my first eng job at a mid-size fintech, so a lot of this is above me, but the "A Concrete Consumer Shape" bit made a concept click that I had been nodding along to in code review for months. Thanks for writing at a level that doesn't gatekeep newer engineers out.
Musa Lawal
August 30, 2026
Does the "Retention, Cleanup, and Operational Reality" still hold on a 133-service estate? We're at the smaller end of that and some of these patterns feel like they need a dedicated platform team to run properly.
Christopher Hernandez
August 13, 2026
Good topic. Staff SWE at a Series B neobank in SF here. What we do differently: keep an append-only audit log and rebuild state from it on demand on Envoy. It is not universally better; operational complexity is real, but the testability is dramatically better and that pays for itself the first time you have to answer a CBN question at 3am.
Obinna Eze
August 12, 2026
Staff eng on the ledger team at a neobank in Lagos. We hit this exact thing with NIP settlement last May — manual toil was the presenting symptom, and the "A Concrete Consumer Shape" section is basically how we untangled it. Ended up carving off a shadow queue on Argo Workflows, cut p99 latency by 63%. For context: 365 tps.
Amanda Lee
August 10, 2026
This is why I keep coming back to this blog.
Meg Brown
August 9, 2026
Quick q on "Handling Side Effects You Cannot Roll Back" — how do you handle duplicate events when the settlement window sends duplicate callbacks? We're on Aurora Postgres 15 and ops keep asking for manual replay tooling.
Emily Rodriguez
August 7, 2026
Enjoyed this one. One nit on "Pairing Consumers With a Transactional Outbox": worth mentioning FIFO ordering under failover — otherwise the pattern degrades under real load.
Meg Scott
August 3, 2026
If anyone hits this in skip-level 1:1s specifically, we had good luck with a Postgres advisory-lock queue — the operational visibility alone pays for itself.
Megan Hernandez
July 27, 2026
Good topic. Founder-CTO, 5 engineers, 15 months post-launch here. What we do differently: push the matching into the database instead of pulling into app code on Neon Postgres. It is not universally better; ops needed six weeks to warm up to it, but the recovery story is dramatically better and that pays for itself the first time you have to answer a CBN question at 4am.
Maryam Mohammed
July 23, 2026
The framing on "A Concrete Consumer Shape" alone is worth the read.
