Overview
A synchronization service keeping brand records consistent across downstream systems, with robust failure handling and data reconciliation.
Context
Brand records are shared data: many systems across Ceneje’s markets rely on them being right. This service owns getting those records, more than 2 million of them across 9 countries, from their source to every downstream system that needs them, BigQuery among them.
Problem
This was a new requirement, not a replacement. Brand data had to reach several downstream systems and stay consistent with the source, without coupling those systems to the source or to each other, and without silently losing a change when something failed along the way.
Constraints
Scale: more than 2 million brand records across 9 countries.
Correctness: a lost or out-of-order update leaves a brand wrong in a market until someone notices, so every failure had to be visible and recoverable.
Independence: downstream systems consume at their own pace, and none of them may block the source or each other.
D-01Event-driven propagation over RabbitMQ
- Decision
Every change to a brand record is published as an event on RabbitMQ. Each downstream system, BigQuery included, consumes those events and applies them on its own.
- Why
Downstream systems never call the source and don’t know about each other. Adding a consumer means binding a new queue, not changing the publisher, and a slow or unavailable consumer only builds up its own queue instead of blocking anyone.
- Trade-off
Consistency becomes eventual: for a short window a downstream system can lag behind the source. Delivery, duplicates and ordering also become the service’s own problems, which drove the next three decisions.
D-02Transactional outbox for publishing
- Decision
A change and its event are written in the same database transaction, the event into an outbox table. A separate relay publishes the outbox to RabbitMQ.
- Why
Writing to the database and publishing to the broker as two separate steps can leave them disagreeing if the process dies in between. With the outbox, an event exists if and only if the change was committed.
- Trade-off
An extra table and a relay process to run and monitor, plus a small publishing delay. Delivery becomes at-least-once, so consumers have to cope with duplicates.
D-03Idempotent, version-aware consumers
- Decision
Consumers deduplicate messages and compare record versions before applying a change. Replaying a message is harmless, and an older update can never overwrite a newer one.
- Why
At-least-once delivery and retries guarantee duplicates, and messages can arrive out of order. Making every apply safe to repeat is simpler and more robust than chasing exactly-once delivery.
- Trade-off
Every consumer tracks versions per record, so each target needs somewhere to keep and compare them.
D-04Retries with backoff and a dead-letter queue
- Decision
Failed messages are retried with backoff. Messages that keep failing move to a dead-letter queue, where they can be inspected and replayed.
- Why
Transient failures, such as a target that is briefly unavailable, resolve on their own. A poison message gets parked instead of blocking the queue behind it, and because consumers are idempotent, replaying from the dead-letter queue is safe.
- Trade-off
The dead-letter queue has to be watched: a parked message is a record that stays out of sync until it is replayed.
Outcome
Running in production, keeping more than 2 million brand records consistent across downstream systems in 9 countries. New consumers attach to the event stream without touching the source, and failed messages surface in the dead-letter queue instead of disappearing.
What I'd do differently
With hindsight, which decision would you revisit, and how?