Skip to main content

The Latency Floor: One Microsecond of Work Inside Ten of Plumbing

· 11 min read
Clustron Team
Distributed Systems Engineering

One microsecond of work inside ten microseconds of plumbing

Here is a number that reframes how you think about a key/value store. When we isolate the Zaris in-memory store and feed it operations with nothing else in the way — no socket, no serialization, no routing — a single thread serves reads at over a million per second, and scales cleanly to more than eight million across eight threads. The store, in other words, does the actual work of a GET in under two microseconds, and it is emphatically not the bottleneck.

And yet a full GET over a real TCP connection, on the same machine, plateaus at a few hundred thousand per second. The store can do the work five times faster than the server can deliver it. So the interesting performance question about Zaris is not "how fast is the store?" It's the opposite: where does all the time between the socket and the store go? This post is the answer, traced operation by operation, because the gap between 1.7 microseconds of logic and 10-to-13 microseconds of server CPU per operation is not noise — it's the latency floor, and almost all of it is plumbing.

The measurement that starts the whole investigation​

You cannot optimize a path you haven't separated into layers, so the first thing we built was a set of ablation benchmarks — harnesses that strip the system down to one layer at a time.

  • Store only. Call the partition's in-process store directly, in-process, no network. Result: ~1.15M GET/s and ~0.5M PUT/s per thread, scaling linearly to ~8.2M / ~2.0M at eight threads.
  • Bare socket echo. A plain .NET socket doing frames of the same shape, no Zaris logic: ~7.2M frames/s on the same box.
  • Full native TCP path. The real client talking to the real server: ~300-324K ops/s, with only about three of fourteen logical cores busy on each side.

Read those three together and two usual suspects fall away immediately. The store isn't the limit — it's 3-4x faster than the full path. The operating system and the loopback aren't the limit either — a bare socket moves frames twenty times faster. What's left in the middle, between a socket that can do millions and a store that can do millions, is Zaris's own request plumbing, and it is spending roughly 10-13 microseconds of server CPU per operation to wrap about 1.7 microseconds of real work.

That ratio is the subject. Let's follow one GET through it.

The path of a single GET​

Nothing in that diagram is wrong or wasteful at a glance — it's the honest shape of a networked store. The cost is in the details of each arrow: how many times a thread has to wake up, how many times the bytes are copied, and how much fixed per-request work rides along that has nothing to do with the key being read. Those details are where the eleven microseconds live, and they cluster into three groups.

Bottleneck one: the egress hop chain​

The single largest cost is not reading the value — it's delivering it. Sending one response back to the client, as currently structured, crosses three internal channels, wakes up three threads, and copies the payload four times on its way out:

  1. The response is handed from the shard worker to the server pipeline's send path.
  2. The send path hands it to a per-connection write loop, which encodes it into a pooled buffer.
  3. The write loop enqueues it on the transport connection's send queue.
  4. The transport coalesces queued buffers into a stream and writes to the socket.

Each hand-off is a thread wake-up, and each wake-up is a scheduling event the OS has to service — cheap individually, ruinous at hundreds of thousands per second. And at each stage the bytes get copied again: pooled encode buffer, send queue, coalescing buffer, socket. On the inbound side there's a matching cost: every request that arrives on an IOCP completion thread is handed off to a shard worker before any work happens, which is another wake-up and another queue.

The reason this structure exists is reasonable — decoupling receive, dispatch, encode, and transmit makes each layer simpler to reason about and lets slow consumers backpressure without blocking the socket. But for the common case — an RF1 read whose answer is already sitting in memory on the node that received the request — it is enormously more machinery than the work deserves. The answer could, in principle, be computed and written back on the receive path itself, with one copy into the socket buffer and no hand-offs at all.

That is the first and biggest piece of structural work the measurement points to: collapse egress to one hop and one copy for the hot path, and run the RF1 read/write inline on the receiving thread rather than bouncing it through the worker pool. This is the difference between a design that's easy to draw and a design that's cheap to run, and the two are not the same thing.

Bottleneck two: redundant hashing and a per-request lock​

The second group is fixed per-operation overhead — work that happens on every request regardless of which key it touches, and that mostly shouldn't.

CRC16, three times, bit by bit. Zaris routes keys to partitions with a CRC16 hash — the same hash family Redis Cluster uses, which is how the RESP front-end stays slot-compatible. The problem isn't the choice of hash; it's that the implementation is bit-serial (looping bit by bit rather than using a lookup table) and that a single GET computes it three times as the request passes through routing. At around 260 nanoseconds per computation, that's ~784 ns per GET spent hashing — more than the entire MessagePack decode-plus-encode round trip, which comes in around 533 ns. A table-driven CRC16 computed once and carried along would turn the single largest fixed cost into a rounding error.

A license check that takes a lock, every request. More surprising: the request entry point consults the license service on every single operation, and that consult takes a process-wide lock and reads the wall clock (DateTime.UtcNow) to decide whether the current lease is still valid. A lock on the hot path is exactly the kind of thing that doesn't show up in a single-threaded microbenchmark and then quietly serializes everything under concurrency. The license verdict changes at most a few times over the life of a process — when a lease is granted, renewed, or expires. It has no business being recomputed under a lock on every GET. The fix is to publish the verdict into a single volatile field that a background renewal updates, and let the hot path read it with no lock and no clock call. (The lease machinery itself is the subject of the licensing lifecycle post; the point here is purely that reading the verdict must be free.)

Neither of these is a deep algorithmic problem. They're the kind of cost that accrues when each piece is written correctly in isolation and nobody has yet measured what the sum costs on the path that runs a million times a second.

Bottleneck three: the client spends as much as the server​

It would be convenient to blame the server for everything, but the ablation says otherwise: local benchmarks are client-bound, and pinning the client to performance cores alone took it from 169K to 299K ops/s. The client spends roughly 9 microseconds and 1.8 KB per GET, and the costs rhyme with the server's:

  • Completions bounce through the thread pool. Each pending request is tracked by a TaskCompletionSource configured to run its continuation asynchronously, so when the reply arrives the completion is posted to the thread pool rather than resumed inline — another scheduling hop, on every operation.
  • An environment variable is read on every single operation. The client calls Environment.GetEnvironmentVariable("ZARIS_REQUEST_TIMEOUT") per request to decide its deadline. Reading an environment variable walks the process environment block; doing it a million times a second is the single largest Zaris-authored frame on the client. It should be read once at startup and cached.
  • A semaphore and a channel lock per operation, plus per-request allocations — a client ID stamped on each request, entry metadata requested and returned even when the caller didn't ask for it.

The through-line with the server is identical: the work of a GET is tiny, and it's surrounded by scheduling hops, locks, and allocations that each looked free in isolation. The client's hot path has already started down the right road — the move to ValueTask for the core operations removed an allocation per call on synchronous completion — and the rest of the list is the same idea applied further: cache the env var, complete without the thread-pool bounce, receive on the socket with a ValueTask, and stop allocating things nobody asked for.

So what is the floor, really?​

Add it up and the honest picture is this. On a laptop, the native TCP path tops out around 300K ops/s not because the store is slow — it can do millions — but because each operation drags 10-13 microseconds of server CPU and ~9 microseconds of client CPU through a chain of hops, copies, redundant hashes, and locks. The floor is set by the plumbing, not the work.

The encouraging corollary is that the floor is made of fixable things, and the estimate for what fixing them buys is concrete. The structural list is:

  1. Collapse egress to one hop and one copy.
  2. Run the RF1 hot path inline on the receive thread, or batched per receive.
  3. Table-driven CRC16, hashed once per operation.
  4. Client completions with zero allocation and no thread-pool bounce.
  5. ValueTask-based socket receive on the client.
  6. Cache the request-timeout env var instead of reading it per op.
  7. Publish the license verdict in a volatile field — no lock, no clock, on the hot path.
  8. A general allocation diet: pooled responses, no per-request client ID, no unrequested metadata.

Together those are estimated to take server CPU per operation from ~10 down to ~3-4 microseconds — which is the difference between needing a dozen busy cores for a million ops/sec and needing about four.

The honest caveats​

Two things keep this post truthful rather than triumphant.

First, the absolute numbers here are laptop numbers, and they're likely 2-4x low versus real server hardware with dedicated client machines. What transfers between environments is not the absolute throughput — it's the relative attribution: the store is not the bottleneck, egress is the biggest server cost, and the client is spending as much as the server. Those conclusions are machine-independent; the final ops/sec figure is not, and the real validation is at scale on Kubernetes, not on a workstation.

Second, replication changes the ceiling. Everything above is the RF1 story — a single copy, the fast case. An RF2 PUT has to durably replicate to a backup before it can acknowledge, and that is a genuinely different and more expensive path; a realistic target for it is 400-500K per node, not a million. And when you don't need per-operation round trips at all, the batch path amortizes almost all of this fixed cost across many operations and can reach a million per node today. The million-ops single-node figure for individual operations is a target that the structural work above is meant to reach — not a number we're quoting as shipped.

That distinction is the whole discipline of performance work worth trusting: separate what the measurement proves (the store is fast, the plumbing is the floor, and here is where each microsecond goes) from what it promises (a specific number after a specific list of changes). The first is done and solid. The second is engineering we can now do with a map instead of a guess — which is the entire reason you instrument the path before you touch it.