Skip to main content
How We Handle Schema Evolution in Events — Anselm Fowel
AI & Technology

How We Handle Schema Evolution in Events

9 min read
670 views
Share:

A few years ago I inherited a payments platform where the events team and the billing team had quietly disagreed about what a field called amount meant. One side stored it in minor units, cents. The other read it as a decimal. Nobody noticed until a reconciliation run flagged a customer who appeared to owe exactly one hundred times more than they did. The schema was technically identical on both sides. The meaning had drifted. That is the whole problem with schema evolution: the shape stays legal while the truth quietly rots.

Events are the part of a fintech system that outlives everything else. Services get rewritten, teams reorganize, but the event log is forever, and every consumer you have ever deployed is an implicit contract you now owe. This is how I think about changing that contract without breaking the people who depend on it.

Events Are Forever, So Treat Them That Way

The first mental shift I ask engineers to make is that an event is not a message, it is a published fact. Once a PaymentAuthorized event has been consumed by a downstream ledger, a fraud model, and a customer notification service, you no longer own it. You cannot recall it, you cannot quietly change what it means, and in a regulated environment you often cannot delete it for seven years. That permanence is exactly why we like event sourcing, and it is also why a careless schema change is so expensive.

Contrast this with a REST API, where a bad response is gone the moment the request completes. If you break an API and fix it ten minutes later, you have inconvenienced whoever called it in those ten minutes. If you break an event schema, you have poisoned a log that gets replayed. New consumers you write next year will read the broken records and have to cope. I have watched a team spend a full sprint writing defensive parsing logic for events that were malformed in 2021, because the log is the source of truth and the log cannot lie about its own history.

Additive by Default, Breaking Almost Never

My default rule, and I will die on this hill, is that every schema change should be additive. You add fields, you never remove them, and you never repurpose one. Adding an optional field is safe because old consumers ignore what they do not recognize and new consumers get the extra data. Removing a field or renaming it breaks anyone still reading the old name, and you almost never have a complete list of who that is.

Here is the concrete taxonomy I hand to new engineers. It fits on an index card, which is the point.

  • Safe: adding a new optional field, adding a new event type, widening a numeric range, adding a new enum value that old consumers can treat as "unknown".
  • Dangerous but sometimes necessary: making an optional field required, adding an enum value that changes routing logic, tightening a validation rule.
  • Never without a versioned migration: renaming a field, changing a type, changing units, changing the meaning of an existing value.

The units example is the one that bites hardest. Changing a field from dollars to cents does not break your deserializer. It compiles, it validates, it passes every schema check you own. It just silently multiplies money by a hundred. This is why I treat semantic changes as more dangerous than structural ones, and why unit-of-measure lives in the field name whenever I can get away with it: amountMinor, not amount.

Picking a Format That Tolerates Change

We standardized on a schema registry with a defined compatibility mode, and honestly the format matters less than the discipline the registry enforces. We use Avro with backward compatibility set at the subject level, which means a new schema must be able to read data written with the old one. Protobuf would have been a perfectly fine choice too; its field-number rules give you similar guarantees almost for free, as long as nobody ever reuses a retired field number.

What I have grown allergic to is raw JSON with no registry at all. JSON is wonderful for humans and treacherous for contracts, because nothing stops a well-meaning engineer from shipping a change that a consumer three teams away cannot parse. If you are going to use JSON on the wire, and there are good reasons to, then wrap it in JSON Schema validation in your CI pipeline and gate merges on a compatibility check. The format is not what saves you. The gate is what saves you.

The registry is not there to describe your data. It is there to stop the merge that breaks production, at nine in the morning, before coffee, when nobody is thinking clearly.

How We Actually Version

When an additive change is not enough and we genuinely need a breaking change, we do not mutate the event. We publish a new event type alongside the old one. PaymentAuthorized becomes PaymentAuthorizedV2, both flow through the same topic for a while, and producers emit both until every consumer has moved. This dual-write period is annoying and it is also the only approach I trust, because it lets consumers migrate on their own schedule instead of on yours.

I resisted this for a long time because carrying two event types feels ugly, and it is ugly. But the alternative is a flag-day cutover where every producer and consumer must deploy in a coordinated window, and coordinated windows across eight teams are where weekends go to die. Versioned types let each team read the migration guide, add a handler for V2, verify it in staging, and cut over when they are ready. The cost is a few weeks of duplication. The benefit is that nobody's pager goes off.

Enjoying this article?

Get more like it in your inbox — practical engineering leadership, fintech, and AI. No spam, unsubscribe anytime.

Here is roughly how a consumer handles the transition in C#. Nothing clever, and that is deliberate.

public async Task Handle(EventEnvelope envelope)
{
    switch (envelope.Type)
    {
        case "PaymentAuthorized.v2":
            var v2 = envelope.Deserialize<PaymentAuthorizedV2>();
            await ProcessAuthorization(v2.PaymentId, v2.AmountMinor, v2.Currency);
            break;

        case "PaymentAuthorized.v1":
            var v1 = envelope.Deserialize<PaymentAuthorizedV1>();
            // v1 has no Currency; upcast to the assumed default.
            var minor = (long)(v1.Amount * 100m);
            await ProcessAuthorization(v1.PaymentId, minor, "EUR");
            break;

        default:
            // Unknown version: log and park, do not throw.
            _log.Warning("Unhandled event version {Type}", envelope.Type);
            await _deadLetter.Park(envelope);
            break;
    }
}

Upcasting on the Read Path

The trick that made versioned events bearable for us is upcasting. Instead of forcing every consumer to understand every historical version, you write a small function that transforms an old event into the current shape at read time. The consumer only ever sees the latest version. All the compatibility mess lives in one place, gets unit tested, and stays out of business logic.

In the snippet above you can see the seed of it: a V1 event has no currency, so the upcaster injects the default that was implicit back when we only operated in one market. That "implicit default" is the archaeology you inevitably do. Every old event carries assumptions that were true at the time and are now written down nowhere. Upcasting is where you make those assumptions explicit and, frankly, where you find out how much your old self knew that your current self forgot.

Make the Compatibility Check a Merge Gate

None of this discipline survives contact with a busy team unless a machine enforces it. We wired the schema registry's compatibility check straight into CI. If a pull request changes an Avro schema in a way that breaks backward compatibility, the build fails with a message that names the offending field. It is not a linter warning that people learn to ignore. It blocks the merge.

This one change did more for our reliability than any amount of documentation. Documentation about schema rules gets read once and forgotten. A red build gets read every time. The first month we turned it on, it caught eleven breaking changes that would otherwise have shipped, and roughly half of them were type changes the author genuinely believed were harmless. People are not careless. They just cannot hold the entire consumer graph in their head, and they should not have to.

Testing the Seams, Not Just the Happy Path

The bugs that hurt in event systems are almost never in the current version talking to the current version. They are in a two-year-old event being replayed into today's consumer. So the tests that earn their keep are the ones that hydrate real historical payloads and assert the upcaster produces something sane.

We keep a fixtures directory of anonymized production events, one per schema version we have ever emitted, captured the day each version went live. Every build deserializes all of them and runs them through the current read path. When someone changes an upcaster, the fixtures tell them immediately whether they broke the past. It is not glamorous work. It is the difference between a replay that rebuilds a customer's balance correctly and one that quietly reconstructs it wrong, which is the single scariest sentence in this whole essay.

What I Would Tell My Younger Self

If I could send one message back to the version of me who designed our first event schema, it would be: put a version number and a units suffix in from day one, even when you are certain you will never need them. The cost of an unused schemaVersion field is one integer. The cost of retrofitting versioning onto a live log with millions of records and no version marker is a quarter of pain and a bespoke migration tool you will never fully trust.

The second thing I would say is that schema evolution is a people problem wearing an engineering costume. The technology, registries, upcasters, versioned types, is all solved and boring. What is hard is getting eight teams to agree on what a field means and to keep agreeing as the business changes underneath them. The most valuable artifact we produce is not the schema. It is the short written note that says what each field means and, crucially, what it does not.

Anselm Fowel, CTO and fintech architect
Anselm Fowel — CTO & fintech architect

Conclusion

Schema evolution rewards paranoia and punishes cleverness. Every time I have tried to be clever, reuse a field, sneak in a semantic change, save myself a versioned type, it has cost me more than the honest, boring, additive path would have. The event log remembers everything, including your shortcuts. So write the extra field, publish the V2, keep the old fixtures, gate the merge, and accept that the ugly, duplicated, slightly-too-verbose approach is the one that lets you sleep. In a system that handles other people's money, a good night's sleep is worth more than an elegant schema.

Enjoyed this article? Share it with others!

Share:

Get new posts in your inbox

Occasional, practical notes on engineering leadership, fintech, and building with AI. No spam, unsubscribe anytime.

Comments (0)

Leave a Comment

Comments are moderated and will appear after review.

No comments yet. Be the first to comment!

About the author

Anselm Fowel

Anselm Fowel

Chief Technology Officer & fintech architect. 16+ years leading engineering across AlliancePay, Mondu, Transalliance, Global Accelerex, and Fidelity Bank — writing here about engineering leadership, fintech architecture, and AI in production.

Read next