One Log, Two Jobs: How a Replica and a New Owner Catch Up the Same Way
Two very different-looking events in a distributed store turn out to be the same problem underneath. The first: a replica falls behind its primary — a GC pause, a slow disk, a brief network partition — and now holds a stale copy of a partition it must eventually match. The second: a partition is reassigned to a node that has never held it, and that node must acquire the entire dataset from the current owner. One is replication lag; the other is migration. They have different triggers, different operators, different failure stories.
But the question each one has to answer is identical: how do I converge to the current truth of this partition without freezing writes while I do it? Zaris answers both with a single mechanism — a revision-ordered operation log kept per segment, replayed from a known point. This post is about that log: how it's built on the write path, how it's bounded so it can't grow without limit, how it's replayed, and why the exact same substrate carries both a lagging replica and a brand-new owner to convergence.
The revision is the spine
Everything starts with a single counter. Each segment — Zaris's unit of partitioned storage — owns a monotonic revision in its SegmentStore:
private long _storeRevision = 0;
public Revision CurrentRevision => new(_storeRevision);
Every mutation bumps it. A Put, a Delete, a TTL change — each takes the next revision via Interlocked.Increment(ref _storeRevision) while holding that key's stripe lock, and in the same critical section records the operation:
revision = new Revision(Interlocked.Increment(ref _storeRevision));
// ... apply to the live map ...
AppendOperationLog(revision, key, ReplicationOperation.Put, record);
So the revision isn't a timestamp or a version number you manage — it's a store-owned, strictly increasing sequence number stamped on every change in commit order. That ordering is the whole trick. If you know the revision a copy last saw, you know exactly what it's missing: everything stamped higher.
The log entry itself is deliberately small and flat:
[MessagePackObject]
public sealed class SegmentOperationLogEntry
{
[Key(0)] public long Revision { get; set; }
[Key(1)] public string Key { get; set; }
[Key(2)] public ReplicationOperation Operation { get; set; } // Put | Delete
[Key(3)] public StoredRecord Record { get; set; }
}
The log is a ConcurrentDictionary<long, SegmentOperationLogEntry> keyed by revision — an append-only, revision-indexed write-ahead log. A delete is a first-class entry, not an absence; this matters, because "key X was removed at revision 1202" is information a stale copy must replay, and you cannot represent a removal by simply not sending something.
Snapshot, then tail
You cannot replay a log from the beginning of time — it would be unbounded, and most of it is long overwritten. So catch-up is always two parts: a snapshot of the live set as of some revision S, followed by the tail of the log, every operation stamped after S.
var (snapshotRevision, snapshotEntries) = _store.CreateSegmentSnapshot(segmentId);
// ... later, the delta ...
var operations = _store.GetSegmentOperationsAfter(segmentId, snapshotRevision.Value);
The snapshot is the bulk of the data; the tail is the handful of writes that landed while the snapshot was being built and shipped. The consumer applies the snapshot first, then replays the tail in revision order, and it has arrived at revision S + n.
There's a subtlety here that separates a correct implementation from a hopeful one. Between capturing the snapshot at S and finishing the stream, new writes keep landing. The serve side handles this with a bounded drain loop — it re-queries the tail after the last revision it already sent, and repeats until a pass comes back empty, so the hand-off reflects the source's true current state rather than a point frozen seconds ago. Any writes that arrive after the final drain pass ride the ordinary live-replication path, and the completion message carries an AckedHighRevision watermark so the source can re-emit anything above it. The tear between "snapshot" and "live" is closed deliberately, not wished away.
The log is bounded — and that changes the contract
A revision-indexed WAL that grew forever would be a memory leak with extra steps. Zaris's op-log is bounded: a soft retention target of 512 entries and a hard cap of 200,000, trimmed opportunistically once the live count crosses twice the soft target (hysteresis, so trimming isn't on every write). Trimming removes the oldest contiguous revision range and publishes a trim floor — the watermark below which operations are gone for good:
public long OpLogTrimFloor => Volatile.Read(ref _opLogTrimWatermark);
Bounding the log creates a new failure mode, and the honest design is in how it's handled rather than hidden. If a consumer asks for the tail after revision S, but S is now below the trim floor, the operations it needs no longer exist — the range (S, floor] was trimmed away. A naive read would silently return a partial tail and the consumer would converge to a corrupt, gap-ridden copy. Zaris refuses to do that. The guarded read returns the tail only if it can serve it gaplessly:
// returns false if `after` < trim floor — the delta would have a hole
if (!_localStore.TryGetSegmentOperationsAfter(seg, request.SnapshotRevision, out var finalOps))
{
response.Aborted = true;
response.AbortReason = "FinalDeltaBelowTrimFloor";
// ... release fence, unpin ...
}
A consumer that gets this abort doesn't limp along — it throws away its stale anchor and restarts from a fresh snapshot, whose revision is current by definition. "Below the trim floor" is a loud, recoverable signal, never a quiet gap. That's the gapless-delta guarantee, and it's what lets the log stay small without ever lying about what it contains.
To keep a catch-up that's in progress from being trimmed out from under itself, the serve side pins the floor at the snapshot revision for the life of the transfer:
_store.PinSegmentTrim(segmentId, snapshot.Revision.Value); // released in finally
A pin holds the trim floor at or below the snapshot point so the tail a consumer still needs is retained. Trimming never crosses the lowest active pin. The bound and the pin together mean the log is small in steady state but guaranteed-complete exactly for the transfers that depend on it.
Replica catch-up: the direct path
When a replica needs to resync, ReplicaSyncRequestHandler runs the snapshot-then-tail model directly. The order of the first few lines is the whole correctness argument:
ArmSegmentOpLog(segmentId)— guarantee the log is recording before the anchor is taken (more on why below).CreateSegmentSnapshot(segmentId)— capture the live set at revisionS.PinSegmentTrim(segmentId, S)— pin the floor so the tail survives the stream.- Stream the snapshot as chunks of
ReplicaSyncChunkKind.Snapshot, then the tail as chunks ofReplicaSyncChunkKind.Replay. - Run the bounded drain loop until a tail pass returns empty.
On the apply side, each chunk is routed to a per-segment processor that applies the snapshot entries, then replays the operations strictly in revision order — a delete calls into the local store's delete path, everything else writes the record:
foreach (var op in chunk.Operations.OrderBy(o => o.Revision))
{
if (op.Operation == ReplicationOperation.Delete)
_localStore.DeleteRecord(op.Key, op.Record);
else
_localStore.SaveRecord(op.Key, op.Record);
}
Order matters because the same key may appear several times in one tail — set, delete, set again. Replaying in revision order reproduces the primary's exact sequence; replaying out of order could resurrect a deleted key or roll back a newer write.
Migration: the same rails, with a fence
A brand-new owner doesn't get a different engine — it gets the same one, plus a cutover gate. Migration is two phases. The bulk phase (SegmentMigrationRequestHandler) arms the log, snapshots, pins the trim floor, and streams the records — identical opening moves to replica sync. Then, once the new owner has applied the bulk set, it asks for the final delta anchored at the snapshot revision it was given:
public class SegmentMigrationFinalDeltaRequest
{
public ushort SegmentId { get; set; }
public long SnapshotRevision { get; set; } // the anchor
// ...
}
The source's response to that request is where migration adds the one thing replication doesn't need. It engages a write fence first, captures the fence revision, and then builds the delta via the same guarded tail read — so every operation in the final delta is at or below the fence, because writes are now held:
EngageMigrationWriteFence(seg);
var fenceRevision = GetSegmentCurrentRevision(seg);
if (!_localStore.TryGetSegmentOperationsAfter(seg, request.SnapshotRevision, out var finalOps))
{
// below trim floor -> abort, owner re-drives from a fresh snapshot
}
response.Operations = finalOps.ToList();
response.SourceDigest = ComputeSegmentLiveDigest(seg);
The response also carries a SourceDigest. The new owner replays the final-delta operations — again OrderBy(o => o.Revision), again routing deletes and puts separately — and completes the cutover only if its own recomputed digest equals the source's. If the contents don't match bit-for-bit, the cutover doesn't happen; the owner re-drives. The fence makes the final delta finite and quiescent; the digest makes the handoff verifiable.
So replication and migration differ in exactly two places, and are otherwise one codebase: migration adds a write fence and a digest-equality cutover gate; replication reads the tail under a pin with a drain loop. The log, the snapshot, the revision ordering, the trim-floor contract — all shared.
Why replay uses the "raw" write path
One detail worth pulling out: catch-up doesn't replay through the normal Put/Delete that a client would use. It uses PutRaw / DeleteRaw, which preserve the source's revision and item version instead of re-stamping them. A normal Put would assign a fresh version as if the write originated locally; that would make the replica's versions diverge from the primary's and break the digest comparison migration relies on. The raw path advances the local revision counter to the incoming value (Interlocked.Exchange(ref _storeRevision, Math.Max(_storeRevision, record.Revision.Value))) rather than inventing a new one. Replay is faithful reproduction, not re-origination — the copy ends up not just with the same data but with the same version lineage, which is what makes "do our digests match?" a meaningful question.
The honest optimization: skipping the log when nobody's listening
Appending to the op-log isn't free — it's an allocation and a second dictionary insert on every write, touching shared cachelines on the hot path. If a segment has no replica and this node is its sole primary, nothing will ever replay that log, so Zaris can skip the append. This ships, and it's gated carefully.
The obligation is a per-segment flag, safe-by-default on:
private long _logObligation = 1; // DEFAULT ON — safe by default
It's derived from topology, not from a config file. On every partition-map apply, RefreshOpLogObligations() recomputes it per segment: obligation = !(noReplicas && selfIsPrimary). The log is skipped only when the partition provably has no replica and this node is its sole primary; a missing or ambiguous entry falls back to on. There is no operator toggle that can turn it off incorrectly — it follows the cluster's actual shape.
The skip is one line at the top of the append:
if (Volatile.Read(ref _logObligation) == 0 && _trimPins.IsEmpty) return;
Note the second clause: even if the flag says "no obligation," a held trim-pin forces logging on. If a transfer is in flight, its pin keeps the log recording regardless of topology.
The part that makes this provably safe is the arm-before-snapshot barrier. Every consumer arms the log immediately before it captures its snapshot anchor:
public void ArmOpLog()
{
Interlocked.Exchange(ref _logObligation, 1);
foreach (var stripe in _keyLocks) lock (stripe) { } // happens-before barrier
}
Setting the flag then sweeping all the key-stripe locks establishes a happens-before edge: every write that assigns a revision after the arm will observe the obligation and log. So any write that skipped the log necessarily committed before the arm, which means its revision is at or below the snapshot revision S — and is therefore carried by the snapshot, not the tail. The skipped writes aren't lost; they're just delivered as part of the bulk set instead of the delta. Turning the obligation off is refused while any pin is held, closing the other side of the window.
It's the kind of optimization that's only worth shipping if you can state exactly why it's safe. Here the argument is a single invariant — arm before you anchor — and both the replica-sync and migration serve paths obey it on the two lines before they snapshot.
One substrate, earned
The appeal of a shared substrate isn't elegance for its own sake. It's that the hard reasoning — revision ordering, the snapshot/tail tear, the gapless-delta contract against a bounded log, faithful version-preserving replay — gets done once and is then trusted by every consumer. A lagging replica and a freshly assigned owner look like different operational events, but they ask the same question and get the same answer from the same code. Migration layers a fence and a digest check on top; replication layers a pin and a drain loop. Neither reinvents the log, and neither can drift from the other's correctness.
The log stays small because it's bounded and pinned only when needed; it stays honest because a delta below the trim floor aborts to a fresh snapshot instead of returning a hole; and it stays cheap because the one write that no one will ever replay is the one write it's allowed to skip — provably, under a barrier, and on by default everywhere else.