Skip to main content

What Crosses the Wire: The Anatomy of a Zaris Record

· 12 min read
Clustron Team
Distributed Systems Engineering

The anatomy of a Zaris stored record

Most of what Zaris does well — replicating a write to a backup, catching up a lagging replica, cutting a partition over to a new owner — comes down to one quiet question: what, exactly, is the unit of data that moves between nodes? If that unit is right, convergence is almost boring. If it drops a field or carries the wrong one, you get a resurrected deleted key, a replica that can't tell stale from fresh, or one node's bookkeeping leaking onto another.

In Zaris that unit is a single type — StoredRecord — and it is worth taking apart field by field, because every decision in it is load-bearing. This post is the anatomy: what's in the record, how it's framed on the wire, why a delete travels as a record instead of as an absence, why there are two different version numbers, and — just as important — what is deliberately left behind on each node and never crosses the wire at all.

One payload, everywhere​

A key/value store has a surprising number of places where "a value" has to be described precisely: the in-memory map, the replication stream, the migration snapshot, the operation log. Zaris uses the same record shape for all of them. The thing sitting in the segment's dictionary is the same thing a replica receives, is the same thing a new owner receives during cut-over. That uniformity is why the operation log can do two jobs at once and why a migration snapshot and a replica backfill don't need separate encodings.

Here is the whole type, stripped to its wire surface:

[MessagePackObject]
public class StoredRecord
{
[Key(0)] public ZarisEntry Entry { get; set; } // key + value + metadata
[Key(1)] public Revision Revision { get; set; } // segment-wide ordering
[Key(2)] public ItemVersion Version { get; set; } // per-key CAS version
[Key(3)] public DateTime TimestampUtc { get; set; }
[Key(4)] public bool IsDeleted { get; set; } = false; // tombstone flag

[IgnoreMember] public long AccountedBytes { get; set; } // node-local
[IgnoreMember] public object? MaterializedValue { get; set; } // node-local
}

Five fields cross the wire. Two more live on the object but are marked [IgnoreMember] and never travel. The rest of this post is why each one is where it is.

The framing: MessagePack with positional keys​

Zaris serialises records with MessagePack, and the [Key(N)] attributes are the reason the wire stays small. MessagePack can serialise an object two ways: as a map of field-name → value (self-describing, but every message repeats every field name as a string), or as a positional array where each field is identified by an integer index. The [Key(0)], [Key(1)], … attributes select the array form. On the wire a record is a short array of values in a fixed order — no field names, no JSON-style punctuation, just the payloads.

That matters because this record is the hot data structure of the whole system. It is serialised on every replicated write and every migrated key; at hundreds of thousands of writes per second the difference between shipping "TimestampUtc" as a string a million times and shipping nothing but the position is not a micro-optimisation, it's the budget.

The positional scheme also sets the rule for evolving the format safely: integer keys are append-only. A new field gets the next unused index; existing indices never change meaning. An older node that doesn't know index 7 simply ignores it; a newer node reading an old record finds index 7 absent and falls back to a default. We'll see one field — labels at map key 3 inside the metadata — that leans on exactly this "absent means default" behaviour to cost nothing when unused.

Inside the entry: key, value, and non-null metadata​

ZarisEntry is the Entry field — the part a user would recognise as "their data":

[MessagePackObject]
public class ZarisEntry
{
[Key(0)] public string Key { get; set; } = string.Empty;
[Key(1)] public object Value { get; set; } = default!;
[Key(2)] public EntryMetadata Metadata { get; } // getter guarantees non-null
}

Two details here are the kind of thing you only learn by getting burned once. First, Value is typed object, not byte[] or string — the value's concrete shape is whatever the writer put there, framed inside the entry, which is how the same record type carries a plain string, a counter, or a whole replicated collection without the record needing to know which.

Second, Metadata is served through a getter that can never return null:

private EntryMetadata? _metadata;
[Key(2)]
public EntryMetadata Metadata
{
get => _metadata ??= new EntryMetadata();
set => _metadata = value;
}

MessagePack bypasses the constructor during deserialization and calls setters directly, so a plain = new() initialiser wouldn't survive a round-trip — a record that arrived carrying no metadata would deserialise with a null one, and every downstream call site would need a null check. The getter makes the guarantee structural instead: metadata is always there, even when the wire carried none. One place to enforce the invariant, zero null guards everywhere else.

Metadata that costs nothing when empty​

EntryMetadata carries the optional attributes of a key — CreatedAt, Ttl, ContentType, an optional lease or lock handle, and labels. The labels field is a small masterclass in not paying for what you don't use. Most keys have no labels, and a naive Dictionary initialised per entry is pure waste — in one real store, 390K label-less entries were each holding an empty 80-byte dictionary, about 31 MB of nothing.

The fix is two-sided. On the heap, label-less entries share a single static empty dictionary, so readers still see a non-null (empty) map but allocate nothing. On the wire, the serialised surface is null when there are no labels:

[Key(3)]
public Dictionary<string, LabelValue>? LabelsSerialized
{
get => _labels != null && _labels.Count > 0 ? _labels : null;
set => _labels = value != null && value.Count > 0 ? value : null;
}

A label-less entry writes nil at index 3 and the reader restores "no labels" from it. This is the append-only key scheme earning its keep: the common case costs one byte, and the feature is there in full the moment a key actually needs a label.

Two numbers, two jobs: Revision and ItemVersion​

The most common point of confusion about Zaris records is that they carry two monotonically increasing numbers. They are not redundant — they answer different questions, and conflating them would break one job or the other.

ItemVersion is per key. It starts at 1 and increments by one each time that key is written. It is the number the client sees and the number optimistic compare-and-swap checks against: "commit only if this key is still at version 7." When the store applies a put, the new version is simply the old key's version plus one:

long newVersion = (existing?.Version.Version ?? 0) + 1;

Revision is per segment. It is a single counter shared by every key in the partition, bumped on every write and every delete:

revision = new Revision(Interlocked.Increment(ref _storeRevision));

That makes Revision a total order over all the mutations in a segment, which is exactly what replication and recovery need. A replica doesn't care that key user:1 is on its 7th version; it cares which segment-wide revisions it has already seen so it can ask for everything after that point. The operation log is keyed by revision, the catch-up protocol requests "everything after revision R," and conflict resolution between an incoming record and a local one compares revisions, not per-key versions:

// An incoming record older than what we already hold is stale — ignore it.
if (existing is not null && existing.Revision.Value >= record.Revision.Value)
return;

So: ItemVersion is the unit of client-visible optimistic concurrency on one key; Revision is the unit of cross-node ordering and convergence across the whole segment. Same record, two clocks, each doing a job the other can't.

A delete is a record, not an absence​

Here is the field that surprises people most: IsDeleted. When you delete a key in Zaris, the store does not simply remove the entry from its dictionary. It writes a StoredRecord with IsDeleted = true — a tombstone — stamps it with the next segment revision, and appends it to the operation log as a delete:

var revision = new Revision(Interlocked.Increment(ref _storeRevision));
_deleteRevisions[key] = revision.Value;
AppendOperationLog(revision.Value, key, ReplicationOperation.Delete, new StoredRecord
{
// ...
Revision = revision,
IsDeleted = true,
});

The reason is the whole point of the record format. If a delete were merely an absence — the key quietly vanishing from a map — there would be nothing to replicate. A replica that missed the moment of deletion would keep serving the old value forever, because "nothing" doesn't travel and can't be ordered against the writes around it. By making the delete a first-class record with its own revision, the delete flows through the exact same replication and catch-up path as a write, and it slots into the segment's total order in the right place.

Tombstones also defend against a subtler failure: a resurrecting stale write. Imagine a delete at revision 50 and an older put for the same key, revision 48, still in flight across the network. Without a memory of the delete, the late put would land and bring the key back from the dead. Zaris records the delete revision and rejects anything at or before it:

private bool HasDeleteRevisionAtOrAfter(string key, long revision)
=> _deleteRevisions.TryGetValue(key, out var deleteRevision)
&& deleteRevision >= revision;

The flip side of tombstones is that they can't accumulate forever, or a delete-heavy workload would leak memory. Zaris trims delete revisions once they're safely below the oldest revision any replica still needs — the same low-water-mark logic that trims the operation log — so the tombstone does its convergence job and is then reclaimed. (The trimming discipline is the subject of the operation-log post.)

What stays home: the [IgnoreMember] fields​

The two fields the record carries but never serialises are as deliberate as the five it does. Both are marked [IgnoreMember], and both are node-local by design — correct only for the node that computed them.

AccountedBytes is the memory this record is charged against the store's budget, stamped once at write time so the store can report byte deltas without re-measuring the previous record. It is explicitly not serialised, because each node computes it from the record's own key and value under its own accounting rules. If it rode the wire, a record replicated from node A would carry A's byte estimate onto node B's books — one node's bookkeeping silently distorting another's. Keeping it [IgnoreMember] means every node independently recomputes it, and a record arriving via replication or migration is accounted by the node that now holds it.

MaterializedValue is a lazily-built, live in-memory form of a native collection (a hash, list, set, or sorted set) cached so per-element operations don't re-deserialise the whole value each time. It is pure local acceleration. The authoritative form is always the framed value bytes inside Entry; a record that arrives from another node materializes this cache lazily on first access. Serialising it would be shipping a node's private scratch state as though it were data.

The discipline here is a clean one worth stating out loud: the wire carries truth, the node carries conveniences. Anything a node can recompute from the record itself — its memory cost, its materialized structure — stays local and is marked [IgnoreMember]. Anything that is part of the record's identity and ordering — key, value, metadata, the two version numbers, the tombstone flag — crosses the wire. The attribute boundary is exactly the truth/convenience boundary, and keeping it sharp is what makes the record safe to copy between nodes without a second thought.

Why this is testable, and tested​

Because the record shape is load-bearing for convergence, its wire behaviour is pinned by tests rather than left to inspection. Two invariants in particular earn their own assertions: the tombstone flag must survive a round-trip (a dropped IsDeleted resurrects a deleted key on a replica), and AccountedBytes must come back as zero after a round-trip (proving it never travelled):

back.IsDeleted.Should().BeTrue(
"a dropped tombstone flag resurrects a deleted key on the replica");

back.AccountedBytes.Should().Be(0,
"AccountedBytes is [IgnoreMember]; each node recomputes it, it must not cross the wire");

These read almost like the thesis of the whole post written as assertions: the things that must cross the wire are proven to survive it, and the things that must not are proven to vanish.

The record is the contract​

It's easy to think of a distributed store's hard problems as living in the fancy parts — the failover protocol, the cut-over handshake, the chaos testing. They do. But those protocols are only as correct as the thing they move, and the thing they move is this one small record. A tombstone instead of an absence is why a delete converges. Two version numbers are why one key's optimistic concurrency and a whole segment's ordering stay untangled. Positional MessagePack keys are why the hot path stays cheap and the format can still grow. And a pair of [IgnoreMember] fields are why replication can copy a record between nodes and trust that nothing node-local came along for the ride.

Get the unit of data right and the big protocols have something solid to stand on. That's the whole case for spending a blog post on five fields and two that stay home.