Skip to main content

No Acknowledged Write Left Behind: The Ownership & Recovery Backbone Under 12 Hours of Chaos

· 10 min read
Clustron Team
Distributed Systems Engineering

Zero acknowledged-write loss, proven over 12 hours of chaos

The happy path of a distributed key-value store is easy to get right. A write lands on the primary, streams to the replicas, everyone agrees, you move on. If that were the whole story, durability would be a solved problem and this post wouldn't exist.

The hard problems live in the seams — the brief, awkward moments when the cluster is mid-transition. A node restarts in the middle of a partition migration. A replica falls far behind the active it's copying from, and then that active fails over to someone else. Ownership of a partition moves to a new node while writes are still in flight to the old one. Each of these is a window: a stretch of a few seconds where, if the handoff is not exactly right, an acknowledged write — one the client was told succeeded — can quietly disappear. Not because the storage engine is buggy, but because two parts of the system disagreed, for a moment, about who owned the truth.

This post is about how Zaris closes those windows, and — the part that actually matters — how we convinced ourselves it works, by running the cluster under continuous, deliberate chaos for twelve hours straight.

The claim, stated honestly​

Let's be precise up front, because this is the kind of claim that invites overreach. The guarantee we are making is:

No acknowledged write is lost under sustained chaos — node stops, starts, restarts, failovers, partition migrations, and scale-out churn — demonstrated by a clean 12-hour continuous certification run.

That is not the same as "Zaris can never lose data under any circumstance." No honest engineer signs that sentence. What we validated is a specific, strong property: once the store has acknowledged a write to the client, that write survives the full catalogue of ownership-transition failures we know how to induce, and the cluster converges back to a consistent state afterward. The twelve-hour run is evidence for that claim, not a proof of universal invulnerability.

There is a second piece of honest nuance that matters for anyone reading a chaos log. During chaos, you will see keys that are transiently missing — a read that doesn't find a key it "should." That is expected and is not data loss. It happens because replication has an asynchronous tail (the write committed on the active but hasn't reached the replica you happened to read), or because the node that holds the key is, at that instant, deliberately stopped by the chaos harness. Those keys reappear on convergence. The entire discipline of certifying durability is separating that benign, self-healing class from genuine loss — a key that no surviving node will ever return. The run passes only when genuine loss is zero. Transiently-missing-then-recovered is tolerated; permanently-gone is not.

The three seams, and how each is closed​

Seam 1 — the stranded replica​

Picture a partition with an active copy and a replica. The active fails, the replica is promoted, the old active comes back as a replica — a promotion/demotion flip. In the wrong implementation, the newly-demoted replica can end up frozen far behind the new active: it stopped receiving the normal replication stream during the flip, and nothing ever tells it to catch up. It sits there, stale, pretending to be a valid copy. If the new active then fails, you promote a replica that is missing thousands of writes.

The fix is a replica catch-up mechanism. A replica that is detected to be stranded far behind its active — not just momentarily lagging, but genuinely frozen below the active's state — is identified and re-synced from the active, instead of being left to rot. The key design subtlety is telling "normal async lag" apart from "stuck": a healthy replica is always a little behind and that's fine; a stranded one is behind and not moving. The second condition is the one that triggers a re-pull.

Seam 2 — the cut-over that destroys the only copy​

Migration is where a partition changes owners. The source node hands the partition's data to the target, and at some point there is a cut-over: ownership flips, and the source is now free to drop its copy to reclaim the space. That "free to drop" step is a loaded gun. If the source drops its copy based on a stale cut-over proof — a signal that says "the target has it" when the target has not actually, fully adopted the partition — you have just deleted the only good copy of that data.

The fix is cut-over safety: the source never destroys its only copy on a stale proof. A segment that the new owner has not demonstrably, fully adopted is retained, not dropped — held until it is genuinely safe to let go. If the target still needs data, the source is still holding it, and can re-serve it. The principle is simple to state and easy to get wrong: dropping is the dangerous operation, so make dropping prove its safety, not the other way around.

Seam 3 — recovery that reconciles both directions​

After a churn storm — many failures stacked close together — a replica can end up out of agreement with its active in either direction. The obvious case is being behind (missed writes). The non-obvious case, and the one that bites you, is being ahead: during a chaotic reshuffle a replica can hold a record the current active no longer has, for example an update or delete that reached one copy but not the other before ownership moved.

So recovery has to reconcile both directions back to convergence — pull missing records into a behind replica, and reconcile away the extra records on an ahead replica — until active and replica agree. A reconciliation that only ever fixed "behind" would leave the "ahead" divergence sitting there forever.

One design decision inside this is worth calling out, because it's the kind of thing that separates a mechanism that works in a demo from one that survives twelve hours. The leader-side reconciliation reuses the sync-state counts the cluster is already tracking. It does not stand up a new detection network that fans out probes across nodes, and it stays off the hot map-commit path — the critical path that every ownership change already flows through. Reusing existing signal means reconciliation adds no new failure mode of its own and no new latency to the operations that were already the most delicate. The cheapest reliable detector is the one you already have.

The shape of it​

Put the three seams together and the durability story is one flow: a write is acknowledged, chaos happens, and each mechanism catches its corresponding failure before it can turn into loss.

The dashed path is the honest-nuance branch: the transiently-missing keys that resolve themselves. They flow into convergence too; they just never counted as loss in the first place.

How we actually know: the 12-hour certification​

A design argument is a hypothesis. The test is the evidence. So we run a continuous 12-hour chaos certification against a live cluster, and the bar is clean — the moment it isn't, the run has found a bug and we go fix it rather than ship.

What "chaos" means concretely, sustained for the full duration:

  • Random node churn — nodes stopped, started, and restarted on an ongoing basis, so the cluster is essentially never in a fully-settled state.
  • Failovers — actives killed out from under their partitions, forcing promotions.
  • Partition migrations — ownership moving between nodes while the workload runs.
  • Scale-out churn — the cluster membership itself changing under load.

Against all of that runs a workload that is continuously checking two independent things. One suite checks data-plane convergence: it drives real writes, tracks what was acknowledged, and verifies that every acknowledged write is still retrievable once the cluster converges — with that transient-vs-genuine distinction applied rigorously, so a key missing because its node is currently stopped is not scored as loss, while a key no node will ever return is. The other suite checks API stability: that the control and data APIs keep behaving correctly throughout the churn, not just that bytes survive.

The run passes only when both suites are clean for the entire twelve hours: convergence holds in both directions, no genuine data loss, no replica left failed, no partition left with the wrong owner. That is the result the backbone work was built to produce, and it is the result the endurance run demonstrated — millions of acknowledged operations through sustained chaos, genuine loss of zero, with every transiently-missing-key event explained by the benign async-tail / stopped-node cases and resolved on convergence.

Twelve hours is a deliberate length. An hour of chaos tells you the common transitions work. The rare, ugly interleavings — the restart that lands exactly during a cut-over, the second failover before the first recovery finished — are low-probability per event, so you only accumulate enough of them to trust the system by letting the dice roll for a long time. The seams are, almost by definition, the states the cluster spends the least time in. You have to stay in the storm long enough to keep hitting them.

Why this matters if you run Zaris​

If you operate a data store, you already know that the failures that hurt are not the clean ones. A node that cleanly dies and stays dead is the easy case. The expensive incidents are the messy overlaps — a deploy rolling nodes while a rebalance is in flight, a network blip that triggers a failover that races a migration. Those are exactly the seams this work is about.

The practical takeaway is that Zaris treats the ownership transition — promotion, demotion, migration, recovery — as a first-class part of the durability story, not an afterthought bolted onto a storage engine. An acknowledged write is designed to survive the handoffs, the dropping of a copy is designed to prove its own safety before it happens, and recovery is designed to pull a divergent replica back into agreement no matter which way it drifted. And the claim is backed by continuous chaos endurance testing, with the honest scope stated plainly: no acknowledged write lost under sustained failure, validated empirically — not a magic promise that data can never go wrong.

If you want the mechanics of how partitions, primaries, and replicas fit together underneath all of this, start with Partitioning and replication and the Replica concept in the docs. The backbone described here is what keeps those pieces honest when the cluster is having its worst day.