Skip to main content

How a Zaris Cluster Comes Up: Membership, the Map, and Convergence-Gated Readiness

· 14 min read
Clustron Team
Distributed Systems Engineering

How a Zaris cluster comes up: membership, the map, and convergence-gated readiness

Most writing about distributed stores is about steady state — reads, writes, replication — or about the dramatic failures: a failover when an owner dies, a cut-over when a partition moves. But there's a quieter moment that decides whether any of that works: the moment the cluster comes up. Before a single key can be routed, the nodes have to agree on who is in the cluster, who owns which slice of the keyspace, and — the step most systems get wrong — whether it is safe to start serving at all.

This post follows that boot sequence in the order Zaris actually runs it: membership, the partition map, and the readiness gates that stand between "the process is up" and "the node is taking traffic." It also draws a line, honestly, around what Zaris itself implements versus what it inherits from the underlying cluster runtime it's built on.

Who is in the cluster: membership​

The first honest thing to say is where the boundary lies. The low-level membership transport — peer heartbeats, failure detection, and leader election — is provided by the Clustron.Nodus layer that Zaris is built on. Zaris consumes it through an INodusNode interface: it asks Nodus who the members are, which one holds leadership, and gets notified when a peer's heartbeat is missed or fails. Zaris does not reimplement Raft or a gossip-based discovery protocol; it layers its own partition and ownership control plane on top of that foundation.

What Zaris adds is an explicit membership state machine that translates raw heartbeat events into ownership-relevant decisions. It lives in MembershipStateTracker, and its states are deliberately more nuanced than "up / down":

The states fall into two groups, and the split is the whole point. Alive, Suspected, and Unreachable affect only runtime ownership — they can trigger a replica promotion or mark a partition temporarily unavailable, but they never mutate the partition map. Only Lost, Evicted, and Decommissioned make a node eligible for map regeneration.

That distinction exists because regenerating the partition map is expensive and, done too eagerly, dangerous — it can needlessly reshuffle healthy partitions. So when a node's heartbeat fails, Zaris does not immediately rewrite the map. It moves the node to Unreachable and starts a grace window (30 seconds by default). During that window the cluster does runtime failover only — promote a replica so reads and writes keep flowing — while giving the node a chance to come back. A flap, a GC pause, a brief network blip: none of them should rewrite ownership. Only if the grace window expires does the node become Lost, and only when an entire partition group has been evicted does the single failure-driven map-regeneration path fire. The design rule, stated plainly in the code that wires it up, is: a node leaving never regenerates the map by itself.

How nodes find each other​

Discovery is roster-based, not gossip-discovery. The authoritative list of nodes is configuration, not something the cluster sniffs out at runtime. In a hand-written config it's the nodes array — each entry declaring an id, host, cluster (peer) port, and client port. On Kubernetes, that roster is synthesized from StatefulSet pod ordinals instead of being maintained by hand, with a node's preferred-primary role derived straight from its ordinal.

There is a gossip mechanism in Zaris, but it does a different job — it's how nodes keep their view of cluster state converged, not how they discover that peers exist. We'll come back to it when we talk about readiness.

Who owns what: the partition map​

Once membership is known, the cluster needs a plan for the keyspace. That plan is the partition map — a single, versioned document that every node routes against.

From a key to an owner, in two hops​

A key doesn't map to a partition directly. Zaris uses a two-level mapping that is deliberately byte-for-byte aligned with Redis Cluster so the same routing math works for native clients and the RESP front-end alike:

key --CRC16--> slot (0..16383) --proportional--> segment (0..255) --SegmentToPartition[]--> partition

The first hop hashes the key with CRC16 (CCITT/XMODEM, polynomial 0x1021) modulo 16384 — the exact function Redis Cluster uses, including honoring hash-tag co-location, where a key written as user:{42}:profile hashes only the portion inside the braces so related keys land together. The 16384 slots are then mapped proportionally onto the cluster's segments — 256 of them by default, fixed for the cluster's lifetime — and a SegmentToPartition table assigns each segment to a partition. The partition, finally, is what has owners.

Why the extra segment layer instead of hashing straight to partitions? Because segments are a stable, fine-grained unit you can move between partitions when the cluster grows or rebalances, without rehashing every key. The CRC16-to-slot step never changes; only the segment-to-partition assignment does.

What a partition map contains​

The map itself is a PartitionMap — a MessagePack-serialized record whose shape tells you how carefully ownership is tracked. The fields worth knowing:

  • Version — a monotonic counter that bumps on every map change. This is the number that makes commit safe (more below).
  • TopologyVersion — bumps only on a real structural change (a node or partition joining or leaving, a re-tile), not on a routine migration cut-over. Separating "the shape changed" from "an owner changed" lets the rest of the system react proportionally.
  • TotalSegments — 256 by default, constant for the cluster's life.
  • Partitions — the list of assignments, each carrying a PrimaryNodeId (the placement primary), ReplicaNodeIds, and a stable PartitionIdentity (a GUID minted at birth that survives renumbering) plus ownership-lease fields: ActiveOwnerNodeId, a leader-minted OwnershipGeneration, and the LeaderTerm that granted the lease — which is what lets the cluster fence a stale ex-leader out of authoring ownership.

Note the deliberate split between PrimaryNodeId (where the map prefers a partition to live) and ActiveOwnerNodeId (who is actually serving it right now). Those can differ during a failover or migration, and keeping them as separate fields is what makes those transitions representable without lying about the steady-state plan. The consistency model post explains why reads route to that single active owner.

How the map is planned​

The map isn't assembled greedily. It's produced by a pure, deterministic planner — DesiredMapPlanner — that takes the member node ids, the segment count, a desired partition count, and the replication factor, and returns a plan. Two properties make it worth calling out:

  1. Balanced, contiguous segment splits. Segments are divided as evenly as possible, in contiguous runs, with the first few partitions absorbing the remainder when it doesn't divide cleanly.
  2. Disjoint-group placement. The first N members become the distinct partition primaries; the remaining members fill the replica slots round-robin, so no node ends up in two different partition groups.

The reason the planner is written as a pure function of the inputs is a hard-won one: it makes a whole class of broken maps — a partition with a duplicate primary, a non-contiguous segment range, a degenerate assignment — unrepresentable by construction, rather than something you catch after the fact with a validator. It's worth being honest about the current limitation, too: placement today is positional and disjoint-group; a churn-minimizing, rendezvous-style placement that moves the least possible data on every regeneration is an explicit future refinement, not something shipped. What is shipped is a sticky-placement step that, on a membership-churn regeneration, keeps a live serving owner in place so a healthy partition isn't moved for no reason.

What "commit" means​

The map is authored by one node — the control-plane leader — and nobody else. The leader generates a candidate map, runs it through an ownership-log projection (the leader is the single writer of the ownership log, and the served map is a fold of that log), and then commits it locally before telling anyone. Commit is a single, strict rule enforced in TopologyManager.ApplyMap:

// A map is accepted only if it is strictly newer.
if (_currentMap != null && map.Version <= _currentMap.Version)
return false;

A map with a version less than or equal to the one you already hold is rejected, full stop. Only after the leader commits does it broadcast a MapChangedEvent, push direct map snapshots to peers, and publish the routing change to clients. Followers pull the map (on cold start they request it from already-known peers, retrying until one answers) and apply it through the same monotonic guard. Because every node only ever accepts a strictly-higher version, two nodes can never settle on conflicting maps of the same version — the version number is the serialization point.

One honest cold-start wrinkle: a leaderless node coming up alone can mint a single-node provisional map so it isn't dead in the water, flagged as provisional and later superseded by the authoritative multi-node map once the cluster forms. It's an availability affordance, not the steady-state path.

Is it safe to serve: convergence-gated readiness​

Here's the step that separates a store you can trust from one that merely starts. Having a process up, and even having a map, does not mean a node should take traffic. If a node starts serving before its runtime ownership has caught up with the map it just applied, it can route a request to a partition it doesn't actually hold yet. Zaris refuses to do that, with two distinct gates.

Gate one: cluster-formation stability​

The first gate decides when the leader is allowed to generate the first real map. ClusterReadinessGate waits for membership to be stable — at least a minimum number of nodes present continuously for a stable-duration window (3 seconds by default) — before committing a map. This stops the leader from authoring a map against a half-formed cluster where nodes are still joining one by one, which would mean an immediate regeneration a second later. If stability isn't reached within the timeout (30 seconds by default), the gate force-bootstraps anyway: availability wins over waiting forever. The cluster would rather come up slightly early than wedge.

Gate two: the per-node traffic gate​

The second gate is the one that actually decides whether a node opens its client listener. ClusterTrafficReadinessGate.IsReady() returns true only when all of the following hold:

  • a committed map exists,
  • its segment-to-partition table is fully populated (length == TotalSegments),
  • a runtime ownership map exists,
  • runtime.MapVersion == map.Version — the node's runtime ownership is aligned with the exact version of the map it has committed,
  • and every partition has a runtime entry.

That fourth condition is the heart of it. A node will not advertise itself as ready to serve until its own committed map and its own runtime ownership are on the same version. This is what the hero image at the top is pointing at. There's a 60-second backstop so a node that can't converge doesn't hang startup forever — past that it proceeds with a warning rather than wedging — but in the normal case, the gate holds the door until the two planes agree.

Staying converged: digests and detect-and-heal​

Readiness is a boot-time property; staying converged is a steady-state one, and this is where Zaris's gossip actually earns its keep. Every heartbeat carries a small ClusterStateDigest — the node's epoch, who it thinks the leader is, its map version, its node count, and per-partition ownership generations. No extra RPCs; it rides along on traffic the cluster already sends.

A DigestAccumulator collects peers' digests and, every few seconds, compares them. If a node diverges from the majority for two consecutive evaluation cycles, it raises a drift signal, and a HealingOrchestrator acts on it — the leader sends corrective directives (on a cooldown so it can't thrash), and a node that finds itself the lone outlier self-heals by re-pulling. This is convergence by detect-and-repair, not by a pre-serve quorum vote.

That distinction matters for honesty, so let me be precise about it. In production, map-version agreement across the whole cluster is maintained by the per-node traffic gate plus this digest-driven healing loop — not by a global barrier that blocks all serving until every node has voted on a version. The stronger assertion — every node snapshot agreeing on one leader, one leader epoch, and one map version — is what the 12-hour chaos certification harness checks continuously as a pass/fail invariant. The certification proves the property holds; the runtime achieves it through gating and healing rather than a blocking consensus round on the serve path.

Replication factor, and an honest default​

One last piece of the formation story, because it changes what "ready" means for your durability. The replication factor defaults to 1. With RF=1, each partition has a primary and no replica — a single copy. The planner places RF - 1 replicas per partition, so at the default that's zero. A cluster at RF=1 forms, converges, and serves perfectly well, but a node loss is a data loss for the partitions it owned, by design, not by accident. If you want a partition to survive a node failure you set RF >= 2, and you need enough nodes to actually place those replicas on distinct machines — the Kubernetes durability sizing post works through exactly how many pods that takes. Partition count has a similarly honest default: left unset, the cluster uses one partition per node.

The sequence, end to end​

Put together, a Zaris cluster comes up like this:

  1. Membership resolves from the config roster (or k8s ordinals), and Nodus elects a control-plane leader.
  2. The leader waits behind the formation gate for stable membership, then plans a map with DesiredMapPlanner and commits it locally under the strict monotonic-version guard.
  3. The leader broadcasts the map; followers pull and apply it under the same guard, accepting only strictly-higher versions.
  4. Each node holds behind its traffic gate until runtime.MapVersion == map.Version and routing is fully populated — only then does it open its client listener.
  5. In steady state, heartbeat-borne digests keep every node's view converged, with drift detection and leader-driven healing closing any gap.

None of these steps is dramatic on its own. Together they're the difference between a cluster that is running and one that is ready — and the gap between those two words is where a store quietly loses data if it doesn't take the question seriously.


Zaris is the distributed key-value store from Clustron. If you want to see the control plane from the operator's side, the PowerShell walkthrough creates, starts, and inspects a cluster live; and elastic scaling covers what happens to the map when you add a server to a cluster that's already serving.