Skip to main content

One Hop to the Owner: How the Zaris Client Routes a Key

· 12 min read
Clustron Team
Distributed Systems Engineering

How the Zaris client routes a key to its owner

There are two honest ways to build a distributed cache client. The first is to pick any node, send it the request, and let the cluster forward it internally to whoever owns the key — simple on the client, but every request that lands on the wrong node pays for an extra network hop inside the cluster. The second is to make the client smart enough to know which node owns the key before it sends anything, so the first socket it touches is the one that can actually answer. Zaris takes the second path. A GET leaves your process already addressed to the node that holds the key — no proxy, no forwarding hop, no coordinator in the middle.

This post is the inside view of how that works: the function that turns a key into an owner, the lookup table the client keeps in memory, and — the part that's easy to get wrong — how the client stays correct when ownership moves while your traffic is in flight.

The claim, stated precisely​

When you call client.GetAsync("user:42"), the client does three things before any bytes hit the network:

  1. Hash the key to a segment. A pure function maps the key string to one of the cluster's segments.
  2. Look up the segment's owner. An in-memory array maps that segment directly to the one connection that reaches its active owner node.
  3. Send owner-direct. The request goes over that connection and nowhere else.

No step consults a central router. Step 1 is arithmetic, steps 2 and 3 are an array index and a socket write. The client holds one connection per node and a small routing table, and that's the entire apparatus. The interesting engineering isn't the happy path — it's keeping that table honest.

Step 1 — the key becomes a segment​

Zaris hashes keys the exact way a Redis Cluster client does, on purpose. The placement function lives in SegmentMapper, and its doc comment is blunt about why it looks the way it does:

// Maps a key to its segment. Placement is aligned to the Redis Cluster slot model so a Redis cluster client
// (which routes by CRC16(key) % 16384) and Zaris agree on where a key lives: the key's Redis slot is
// computed the exact Redis way (XMODEM CRC16 over the hash-tag, mod 16384), then the 16384 slots are mapped
// proportionally onto the cluster's segments.
// This is the SINGLE shared placement function used by both the client and the server.

That last line is the one that matters most. It is not "the client computes placement one way and the server validates it another way." It is the same code, compiled into the client assembly and the server assembly, producing the same segment for the same key every time. If the client and server disagreed about where a key lives, every write would be a coin flip between landing on the owner and landing on a node that has to reject it. Sharing the function makes that class of bug unrepresentable.

The computation itself is two moves:

public static ushort ComputeSegmentId(string key, int shardCount, ...)
{
var slot = ComputeSlot(key); // CRC16(hash-tag) % 16384
return SlotToSegment(slot, shardCount); // proportional slot -> segment
}

// The Redis Cluster key slot: XMODEM CRC16 over the hash-tag (or the whole key), mod 16384.
public static int ComputeSlot(string key)
{
var (start, len) = HashTagRange(key);
return Crc16(key, start, len) % RedisSlots; // RedisSlots = 16384
}

First the key is reduced to a slot in the fixed range [0, 16384) using XMODEM CRC16 — the same CRC16-CCITT polynomial (0x1021, init 0x0000) Redis uses. Then the 16384 slots are mapped proportionally onto however many segments the cluster actually has:

public static ushort SlotToSegment(int slot, int shardCount)
=> (ushort)((long)slot * shardCount / RedisSlots);

Two slots can land in the same segment; a segment is a contiguous range of slots. That alignment is what lets Zaris answer CLUSTER SLOTS correctly for a Redis cluster client while still routing by its own segment model internally — the two views are the same view at different resolutions.

Hash tags: deciding what co-locates​

There's a deliberate escape hatch in the key hashing. If a key contains a {...} group with at least one character inside the braces, only the bytes inside the braces are hashed:

// Redis hash tag: if the key contains "{...}" with at least one char between the braces, only that
// substring is hashed (so "{u1}:a" and "{u1}:b" co-locate). Otherwise the whole key is hashed.

This is how you force related keys onto the same segment. {u1}:profile and {u1}:sessions both hash on u1, so they share a segment and therefore an owner — which is exactly what you need for a multi-key transaction or any operation that has to touch several keys atomically on one node. Keys without a tag spread across the whole keyspace. The client and server agree on this too, because it's the same HashTagRange on both sides.

One more detail worth noting for the hot path: the CRC16 loop has an all-ASCII fast path that hashes the characters as bytes with no allocation, and only falls back to a UTF-8 encode for non-ASCII keys. Routing every operation means this runs on every operation, so it was written to not allocate in the common case.

Step 2 — the segment becomes a connection​

Now the client has a segment number. It needs the node that currently owns it. The client keeps an immutable RoutingTable whose fields are the whole story:

// one live connection per node, keyed by node id
public ImmutableDictionary<string, ZarisRemoteClient> Clients { get; init; }

// segment -> the active-owner client it routes to (the precomputed answer)
public ZarisRemoteClient?[] SegmentMap { get; init; }

// segment -> partition, and partition -> owning node id
public ushort[] SegmentToPartition { get; init; }
public ImmutableDictionary<int, string?> RuntimeOwners { get; init; }

The key design choice is SegmentMap. It is an array, indexed by segment, whose entries are already resolved to the connection for that segment's owner. Routing a key is not a graph walk from segment to partition to node to connection at request time — that resolution is done ahead of time when the table is built, so the per-request cost is a single array index:

internal ZarisRemoteClient GetClientForKey(string key)
{
var table = _routing;
return GetClientForSegment(
SegmentMapper.ComputeSegmentId(key, table.TotalSegments),
table);
}

private static ZarisRemoteClient GetClientForSegment(ushort segment, RoutingTable table)
{
if (segment >= table.SegmentMap.Length)
throw new InvalidOperationException($"No client mapped for segment {segment} — map may be mid-update, retry");

if (table.SegmentMap[segment] is not { } client)
{
var partitionId = segment < table.SegmentToPartition.Length
? table.SegmentToPartition[segment] : -1;
throw new PartitionUnavailableException(partitionId, segment);
}
return client;
}

Clients holds exactly one ZarisRemoteClient — one connection — per node. The client does not open a fresh socket per request or per key; it multiplexes all traffic for a node over that node's single connection, and SegmentMap just points many segments at the same connection object (reference-equality dedups them when the client needs the distinct set of active owners, for example to fan out a scan). So the memory cost of partition-awareness is one connection per node plus a couple of arrays sized to the segment count — it does not grow with the number of keys or requests.

The payoff is that there is no node in the cluster whose job is to receive misrouted requests and forward them. Every well-routed request is served by the first node it reaches. That removes a whole tier from the latency budget and, just as importantly, from the failure surface — you can't lose a request in a forwarding hop that doesn't exist. (The companion piece on the latency floor walks what is left in the budget once the proxy hop is gone.)

Step 3 — the map moves under you​

Here's the hard part. The routing table is a cache of a fact — "segment S is owned by node N" — that the cluster can change at any moment. An owner can fail and a replica gets promoted; a partition can be migrated to grow or drain a node. The instant that happens, the client's SegmentMap[S] points at a connection that is no longer the owner. A client that trusted its table blindly would send writes to a node that will correctly refuse them.

Zaris keeps the table honest three ways, in order of preference.

Push — the server tells the client. When ownership changes, the cluster emits a RuntimeOwnershipChangedEvent, and the client has a handler subscribed to it:

public Task HandleAsync(IServerEvent @event)
{
if (@event is RuntimeOwnershipChangedEvent runtimeEvent)
_router.TryUpdateRuntimeOwnership(runtimeEvent);
return Task.CompletedTask;
}

TryUpdateRuntimeOwnership doesn't rebuild the whole table. It clones the SegmentMap array, repoints only the segments whose partition actually changed owners, and swaps the immutable table in with a single reference assignment. Readers never see a half-updated table — they either read the old one or the new one, never a torn mix. This is copy-on-write routing: the common case (nothing changed) costs nothing, and a change touches only the affected segments.

Poll — a slow backstop. A timer refreshes the cluster map every ten seconds regardless, so even if a push event is missed (a reconnect, a dropped notification), the client's view self-heals within one interval:

private readonly TimeSpan _pollInterval = TimeSpan.FromSeconds(10);

React — the request itself corrects the map. The last line of defense is the request that routes wrong. If the client sends to a node that is no longer the owner, the node doesn't silently forward it — it returns a status that tells the client its map is stale, and the retry helper refreshes and tries again:

private static bool RequiresMapRefresh(KvStatus status)
=> status == KvStatus.Moved
|| status == KvStatus.ReplicaWriteRejected
// Ownership is mid-handoff: the client must force-fetch the updated
// ownership and retry against the node that now owns the segment.
|| status == KvStatus.OwnershipChanging;

When a retryable op comes back Moved, OwnershipChanging, or ReplicaWriteRejected, the client calls its refresh hook and retries — a bounded loop of a few attempts with a short backoff, not an open-ended spin. There's a subtle lesson baked into the current code here: an earlier version refreshed the map on every retry attempt, which under heavy node churn turned a single failing op into dozens of refreshes all serialized behind one lock. The fix was to refresh only when the status actually says "your map is wrong," not merely because an attempt failed. The deeper treatment of that retry discipline is its own post — retries under ownership churn — but the one-line version is: a stale map is a self-correcting condition, and the correction is cheap exactly because it's rare and targeted.

The three mechanisms layer: push makes the common case instant, poll makes a missed push eventually consistent, and the reactive retry makes any residual staleness invisible to your call — it just takes one extra round trip the moment ownership moved, and is correct from the next call on.

Why do it on the client at all?​

Partition-aware client routing costs something. The client has to carry a model of the cluster, keep it fresh, and ship the same placement logic the server uses. A proxy-based design pushes all of that into the cluster and lets the client stay dumb. So why pay?

Because the alternative pays on every single request, forever. A proxy or a "send anywhere, we'll forward it" design adds a hop to the hot path of each operation — a hop that is pure overhead, present even when nothing has failed and nothing is moving. Zaris pays the cost of routing once, up front, and only again when ownership actually changes — which, in a healthy cluster, is almost never. The steady state is an array index and a socket write to the right node. That's the whole reason reads scale with your cores instead of bottlenecking on a middle tier: there is no middle tier to bottleneck on.

And it composes with everything else. The same GetClientForKey resolution is what a key-watch, a counter increment, a compare-and-swap, and a single-key data-structure operation all route through, so they all inherit owner-direct delivery and the same self-healing map. The client isn't a dumb pipe with routing bolted on; routing is the thing the client is.

The honest trade-offs​

Owner-direct routing is not free of edges, and it's worth naming them.

  • The client must reach every node. Because there is no proxy, the client opens a connection to each node it routes to. In a flat network that's a non-issue; across a boundary — say a client outside Kubernetes — those nodes have to be individually addressable. That's exactly the problem external client addressing exists to solve, and it's a direct consequence of there being no single front door.
  • A brief map lag is possible, not a loss. In the window between ownership moving and the client learning, a request can route to the old owner. It is rejected and retried, not misapplied — the reactive path above guarantees the retry lands on the real owner. You pay a round trip, not a wrong answer.
  • The placement function is a compatibility contract. Because the client and server share SegmentMapper, you can't change how keys hash to segments without both sides agreeing — which is precisely why the hashing was aligned to the stable Redis slot model rather than an internal scheme that might drift.

Stated plainly: the Zaris client knows where your keys live, keeps that knowledge current without you thinking about it, and spends one network hop — the one that actually does the work — on every operation. The cleverness is not in the fast path, which is deliberately boring. It's in the three quiet mechanisms that keep the fast path correct while the cluster reshapes itself underneath it.