mekhi//node
← all work
JPMorgan Chase 2024 โ€” present

Kafka store-and-forward broker

Digital Payment Systems, Lending Innovation

Store-and-forward message broker that took delivery reliability from 98.2% to 99.95% and cut incident-driven outages by 14 per quarter.

kafkagodistributed-systemsdelivery-semantics

A message broker is easy to write and hard to make mean something. The interesting work was not in moving bytes between producers and consumers. It was in defining what happens when either side fails, and then making that definition hold under real traffic.

The problem

The system was losing messages. Delivery reliability sat at 98.2%, which sounds tolerable until you multiply it by the volume flowing through a payment pipeline. Consumers were restarting, producers were retrying, and the failures were not visible as errors โ€” they showed up much later as an incomplete aggregate, which is a far more expensive way to discover a problem.

Worse, the failure mode was not consistent. Sometimes a message vanished. Often a message arrived twice. That second case is the dangerous one: a duplicate that a human notices is a support ticket, but a duplicate inside a financial aggregate is a correctness bug that surfaces as a reconciliation failure days later.

Store and forward

The fix was to stop treating the broker as a pipe and start treating it as a durable buffer. A message is acknowledged only once it is on disk, not once it has been handed to a consumer. Consumers pull, process, and then acknowledge, so a crash between those two steps leaves the message available rather than lost.

The consequences run through the whole design:

  • Acknowledgement is a durability claim, not a delivery event. If the broker says it stored the message, the message survives a restart.
  • Delivery is decoupled from consumption. Slow consumers create lag, which is visible and measurable, rather than data loss, which is invisible.
  • Replays are safe. Because the broker owns a durable copy, reprocessing a range of messages after a consumer bug is a routine operation instead of an incident.

What it bought

| | Before | After | |---|---|---| | Delivery reliability | 98.2% | 99.95% | | Incident-driven outages | โ€” | 14 fewer per quarter |

The outage number is the one I care about most. A quarter is thirteen weeks, so fourteen fewer incidents is roughly one a week that no longer pages anyone at three in the morning. That is the difference between a system people trust and one they work around.

Why this is the work I point at

Store-and-forward is the same problem as a durable workflow engine, one layer down. Both are asking how to keep a computation running correctly when the process executing it keeps dying. Kafka gives you the log; a workflow engine gives you the same guarantee plus a state machine and replay semantics for the business logic on top.

The part that transfers is the reasoning: decide what an acknowledgement promises, make that promise survive failure, and prefer visible lag over invisible loss.