Many Keys, Few Round-Trips: How Zaris Fans Out a Bulk Operation
Ask a distributed store for one key and the cost is dominated by a single number: the round-trip to the node that owns it. The lookup itself is nanoseconds; the network is microseconds. Now ask for a thousand keys. If you loop and call Get a thousand times, you pay that round-trip a thousand times, serially, and the store spends almost all of the wall-clock sitting idle waiting for packets. The keys might be spread across four nodes that could have answered in parallel — but a naive loop never gives them the chance.
Zaris has three bulk APIs — GetManyAsync, PutManyAsync, DeleteManyAsync — and a mixed ExecuteBatchAsync underneath them, and their whole reason for existing is to turn "a thousand keys" into "one request per owning node, issued in parallel." This post is the fan-out path end to end: how the client buckets keys by owner, why all four APIs collapse onto a single engine, how an Index field keeps the answers in order through two layers of concurrency, and why — unlike a multi-key transaction — a bulk operation is per-item and partial success is normal.
The cost model: round-trips, not operations
The thing a bulk API is fighting is not CPU. A single-key GET on an owning node is a dictionary lookup — the latency floor is one microsecond of real work inside ten of plumbing. The expensive part is the plumbing: the request has to cross the network to the owner and the response has to come back. Call that round-trip time R, typically tens to hundreds of microseconds on a datacenter network, orders of magnitude more than the lookup.
A serial loop of N keys costs roughly N × R, because each call waits for the previous one's packet to return before sending the next. That is the trap. The fix is not a faster lookup — it is to stop paying R per key. If the N keys live on M owning nodes, the floor is R paid M times concurrently, i.e. about one R total regardless of how many keys each node holds. For a thousand keys on four nodes that is the difference between a thousand serial round-trips and effectively one. The entire design below exists to collapse N × R down to ~R.
One engine: everything is a batch of items
The first surprising thing in the cluster client is that GetManyAsync, PutManyAsync, and DeleteManyAsync are not three separate implementations. Each is a thin shim that turns its arguments into a list of KvBatchItems and hands them to the one method that does the real work, ExecuteBatchAsync. Here is GetManyAsync doing exactly that:
var list = keys.Select((k, i) => new KvBatchItem
{
Index = i, // caller's position — the correlation key
Op = KvBatchOp.Get,
Key = k,
GetOptions = options?.ToGetOptions()
}).ToList();
var resp = await ExecuteBatchAsync<T>(
new KvBatchRequest
{
Items = list,
DefaultConsistency = options?.Consistency ?? Defaults.DefaultConsistency
}, ct).ConfigureAwait(false);
PutManyAsync and DeleteManyAsync do the same with KvBatchOp.Put and KvBatchOp.Delete. So the distinction the samples draw between Bulk (one operation kind at a time) and Batch (a mix of PUT / GET / DELETE in one call) is a surface convenience: the bulk helpers are just a batch where every item happens to share an op. There is exactly one fan-out engine, and it is op-agnostic.
Notice Index = i. Every item is stamped with its position in the caller's list before anything is grouped, reordered, or sent. That stamp is what survives the two concurrent hops to come and lets the final answer be put back in order. Hold onto it.
The fan-out: bucket by owner, one request each
Here is the heart of the whole post — the cluster client's batch dispatch, essentially in full:
var table = _routing;
var buckets = new Dictionary<ZarisRemoteClient, List<KvBatchItem>>();
foreach (var item in batch.Items)
{
var seg = SegmentMapper.ComputeSegmentId(item.Key, table.TotalSegments);
var client = GetClientForSegment(seg, table);
if (!buckets.TryGetValue(client, out var list))
buckets[client] = list = new List<KvBatchItem>();
list.Add(item);
}
var tasks = buckets.Select(kv =>
kv.Key.ExecuteBatchAsync<T>(new KvBatchRequest
{
Items = kv.Value,
DefaultConsistency = batch.DefaultConsistency
}, ct));
var perOwner = await Task.WhenAll(tasks).ConfigureAwait(false);
return new KvBatchResponse
{
Results = perOwner.SelectMany(r => r.Results)
.OrderBy(r => r.Index)
.ToList()
};
Read it as four moves:
-
Resolve each key to its owner, locally.
SegmentMapper.ComputeSegmentIdis the same placement function the single-key path uses — CRC16 over the hash-tag, mod the segment count — andGetClientForSegmentmaps that segment to the already-open connection for its owning node. No proxy, no coordinator: the client computes ownership itself, exactly as described in one hop to the owner. The difference is that here it does it for every key in the set at once. -
Bucket. Keys owned by the same node pile into that node's list. A thousand keys spread over four owners become four lists. This is where
Noperations becomeMrequests. -
Dispatch concurrently.
buckets.Select(...)builds oneExecuteBatchAsynccall per owner andTask.WhenAlllets them all fly at once. Four nodes answer in parallel; the wall-clock is the slowest single node's response, not the sum. -
Recombine in order. Each owner returns its slice of results;
SelectManyflattens them andOrderBy(r => r.Index)restores the caller's original ordering. The caller gets back a list whose position i is the answer for key i, even though the keys travelled to different nodes and came back in whatever order the network delivered.
That is the full amortization claim made concrete: the number of round-trips equals the number of distinct owning nodes in the key set, not the number of keys.
Inside one owner: concurrent again, ordered again
Each per-owner sub-batch lands on that node's BatchCommandHandler. It does not execute the items one after another either — it runs them all at once and then re-sorts:
var items = batchReq.Request.Items.Cast<KvBatchItemWire>().ToArray();
// Execute every item in the sub-batch concurrently.
var tasks = items.Select(item => ExecuteItemAsync(item)).ToArray();
var results = await Task.WhenAll(tasks).ConfigureAwait(false);
return new ZarisResponse
{
Payload = new BatchResponsePayload
{
Response = new KvBatchResponse
{
Results = results.OrderBy(r => r.Index).ToList()
}
}
};
ExecuteItemAsync is a small switch on the op that dispatches each item to the store — GetRawAsync, PutRawAsync, or DeleteRawAsync — so a mixed batch of reads and writes is handled item by item inside the same request. (For the in-process and single-node routers the same shape appears as a Parallel.ForEachAsync over the keys, bounded by MaxDegreeOfParallelism — a MaxParallelism option or Environment.ProcessorCount by default — which is why reads scale with cores: a fat sub-batch keeps every core busy instead of one.)
So the Index stamp earns its keep twice. Concurrency scrambles order on the node, so the handler sorts by Index before replying. Fan-out scrambles order on the client, so the client sorts by Index after merging. The caller never sees either scramble. One small integer, carried on every item from the moment the request is built, is the entire mechanism that makes "throw it all at the cluster in parallel" safe to expose as an ordered list.
A bulk operation is not a transaction
This is the distinction people most often get wrong, so it is worth stating sharply. Each item in a batch carries its own outcome. The per-item result type records success, status, and — for the operations that have them — the revision, version, and TTL of that one key:
private static KvBatchItemResult MapKvToBatch(int index, KvBatchOp op, KvResult r) => new()
{
Index = index,
Op = op,
Success = r.IsSuccess,
Status = r.Status, // e.g. Success, NotFound, Moved
Revision = r.Revision?.Value,
ItemVersion = r.Version?.Version,
Ttl = r.TimeToLive
};
There is no shared commit point. A GetMany over a hundred keys where thirty don't exist returns seventy Success results and thirty NotFound results, interleaved by index — not an error, not an empty list. A PutMany where one key's segment is mid-migration returns that one item with status Moved and the rest Success. Partial success is the normal, expected shape of the answer, and the caller is meant to inspect each item.
That is the exact opposite of a multi-key transaction, which is all-or-nothing by construction: a client-coordinated two-phase commit over optimistic versions where either every key moves or none does. The two features look similar from the outside — both take many keys in one call — and are built for opposite needs. Reach for a batch when the keys are independent and you want throughput; reach for a transaction when an invariant spans the keys and a half-applied result would corrupt it. Zaris keeps them as separate APIs precisely so the semantics are never ambiguous: a batch never silently rolls back, and a transaction never silently half-commits.
Surviving ownership that moves mid-batch
A bulk operation is a wider target than a single GET for the problem of ownership changing underneath it. Between the moment the client buckets a key to node A and the moment node A handles the sub-batch, a partition cut-over may have moved that segment to node B. Node A will answer those items with status Moved rather than guess.
The cluster client wraps the whole fan-out in a bounded retry for exactly this:
public override async Task<KvBatchResponse> ExecuteBatchAsync<T>(
KvBatchRequest batch, CancellationToken ct = default)
{
return await RetryHelper.RetryAsync(
() => ExecuteBatchClusterAsync<T>(batch, ct),
maxAttempts: 2,
backoffMillis: 100,
logger: _logger).ConfigureAwait(false);
}
On a retry the routing table has usually refreshed, so the re-bucket sends the moved keys to their new owner. The retry is deliberately shallow — two attempts, a short backoff — because the goal is to ride out a routing table that is momentarily stale, not to paper over a node that is actually down. This is the same discipline described in retries under ownership churn, applied to a whole key set at once: the retry re-runs the grouping, not just the send, so each attempt reflects the latest ownership rather than replaying a stale plan.
One real bug, because the wire form has two shapes
It is worth showing one place this path broke, because it says something about bulk operations specifically. A GET result's value can arrive on the wire in two forms — wrapped in a ZarisEntry envelope, or as a raw serialized value — depending on how it was stored (see what crosses the wire). The original bulk-GET decoder assumed only the wrapper form and threw on the raw one.
For a single GET that is a one-key failure. For a GetMany it was catastrophic in a subtle way: every key in the batch failed, because the one unhandled encoding threw before any result was decoded. A certification run flagged it as "bulk get failures," and the fix was to route decoding through a shared dual-encoding helper that handles both forms:
// Decode through the shared helper (handles both the ZarisEntry-wrapper and
// the raw-value forms the server can emit). The old inline code assumed only
// the wrapper and threw, failing every key of a bulk GET.
var obj = KvBatchItemResultExtensions.GetValue<T>(r);
The lesson that generalizes: in a bulk path, a defect in the shared per-item machinery is multiplied by the batch size, and an error that is a nuisance for one key is an outage for a thousand. The per-item result model is what contains the blast radius — once decoding is correct, one malformed value fails one item and leaves the other 999 intact.
What the fan-out buys
Put the pieces together and the shape of the win is clear. A bulk call stamps every item with its index, resolves each key's owner with the same placement math the single-key path uses, buckets the keys so that N operations become one request per owning node, fires those requests concurrently, and reassembles the answers in the caller's order — surviving a routing table that shifts mid-flight and reporting each key's fate independently. The round-trip count drops from the number of keys to the number of nodes, which is the only number that was ever going to matter.
None of it required a new storage path or a new wire record. The bulk APIs are a scheduling layer over the primitives Zaris already has: the partition-aware placement that tells the client who owns what, the per-record contract that each item result reuses, and the core-parallel execution that chews through a sub-batch once it lands. The hard part of going fast on many keys was never the keys — it was refusing to pay for the network one key at a time.