Skip to main content

7 posts tagged with "Performance"

Throughput, latency, and scaling in distributed systems.

View All Tags

Many Keys, Few Round-Trips: How Zaris Fans Out a Bulk Operation

· 12 min read
Clustron Team
Distributed Systems Engineering

Fanning a bulk operation out to its owning nodes

Ask a distributed store for one key and the cost is dominated by a single number: the round-trip to the node that owns it. The lookup itself is nanoseconds; the network is microseconds. Now ask for a thousand keys. If you loop and call Get a thousand times, you pay that round-trip a thousand times, serially, and the store spends almost all of the wall-clock sitting idle waiting for packets. The keys might be spread across four nodes that could have answered in parallel — but a naive loop never gives them the chance.

Zaris has three bulk APIs — GetManyAsync, PutManyAsync, DeleteManyAsync — and a mixed ExecuteBatchAsync underneath them, and their whole reason for existing is to turn "a thousand keys" into "one request per owning node, issued in parallel." This post is the fan-out path end to end: how the client buckets keys by owner, why all four APIs collapse onto a single engine, how an Index field keeps the answers in order through two layers of concurrency, and why — unlike a multi-key transaction — a bulk operation is per-item and partial success is normal.

One Hop to the Owner: How the Zaris Client Routes a Key

· 12 min read
Clustron Team
Distributed Systems Engineering

How the Zaris client routes a key to its owner

There are two honest ways to build a distributed cache client. The first is to pick any node, send it the request, and let the cluster forward it internally to whoever owns the key — simple on the client, but every request that lands on the wrong node pays for an extra network hop inside the cluster. The second is to make the client smart enough to know which node owns the key before it sends anything, so the first socket it touches is the one that can actually answer. Zaris takes the second path. A GET leaves your process already addressed to the node that holds the key — no proxy, no forwarding hop, no coordinator in the middle.

This post is the inside view of how that works: the function that turns a key into an owner, the lookup table the client keeps in memory, and — the part that's easy to get wrong — how the client stays correct when ownership moves while your traffic is in flight.

The Latency Floor: One Microsecond of Work Inside Ten of Plumbing

· 11 min read
Clustron Team
Distributed Systems Engineering

One microsecond of work inside ten microseconds of plumbing

Here is a number that reframes how you think about a key/value store. When we isolate the Zaris in-memory store and feed it operations with nothing else in the way — no socket, no serialization, no routing — a single thread serves reads at over a million per second, and scales cleanly to more than eight million across eight threads. The store, in other words, does the actual work of a GET in under two microseconds, and it is emphatically not the bottleneck.

And yet a full GET over a real TCP connection, on the same machine, plateaus at a few hundred thousand per second. The store can do the work five times faster than the server can deliver it. So the interesting performance question about Zaris is not "how fast is the store?" It's the opposite: where does all the time between the socket and the store go? This post is the answer, traced operation by operation, because the gap between 1.7 microseconds of logic and 10-to-13 microseconds of server CPU per operation is not noise — it's the latency floor, and almost all of it is plumbing.

The Worker-Pool Ceiling: When a Node Plateaus at 25% CPU

· 8 min read
Clustron Team
Distributed Systems Engineering

A node pinned near 25% CPU with flat throughput because the dispatch worker pool was the ceiling

Here is a benchmark result that looks like a contradiction. A Zaris node serving a read-heavy workload tops out at a little under half a million ops/sec and refuses to climb — add more client connections, more load generators, and the number doesn't move. The obvious diagnosis is "it's saturated." Except the CPU graph says otherwise: the node is sitting at roughly a quarter of its available CPU. Three cores out of four are idle, and throughput is flat.

That is not a core-bound node, and — if you've ruled it out carefully — it's not a client-bound benchmark either. It's a third thing, and it's a classic: the request-dispatch worker pool was narrower than the work wanted to be, so requests queued while cores sat on their hands.

Alloc-Free Hot Paths: How the Zaris .NET Client Uses ValueTask

· 6 min read
Clustron Team
Distributed Systems Engineering

The Zaris .NET client uses ValueTask on its hot-path operations

Most of the time, Task<T> is exactly the right return type for an async method, and you should not think twice about it. But there's a specific place where it quietly costs you: a hot path that completes synchronously, called millions of times.

Here's the mechanism. Task<T> is a reference type. Every call that returns one allocates a Task object on the heap — even when the operation completes synchronously. A cache hit served straight from a local buffer, an awaitable that's already done: none of that requires real asynchrony, but you still pay for a heap object to carry the result. One allocation is nothing. At the call rate of a high-throughput client — a tight loop doing millions of small Gets and Puts — that's a steady stream of small heap allocations, and that stream feeds the garbage collector. More gen-0 collections, more CPU spent on GC instead of your work, more jitter in your tail latencies.

The Zaris .NET client attacks this exactly where it matters and nowhere else. This post is about that decision: which methods changed, why, the rules you have to follow to use them correctly, and — deliberately — which methods we left alone.

Big Values, Small Pauses: Large-Object Chunking in Zaris

· 6 min read
Clustron Team
Distributed Systems Engineering

Large-object chunking keeps big values off the LOH

There's a number every .NET performance engineer eventually memorizes: 85,000 bytes. Any single allocation at or above roughly that size doesn't go on the normal small-object heap — it goes on the Large Object Heap (LOH). The LOH is a different animal. It's only collected during a gen-2 (full) garbage collection, and by default it isn't compacted. That combination is fine for a big buffer you allocate once and keep. It is not fine for a low-latency in-memory store that repeatedly writes and overwrites large values.

Picture the workload: big serialized blobs, large JSON documents, images, protocol payloads — stored, updated, replaced, over and over. Every one of those large writes lands a fresh object on the LOH. Because the LOH isn't compacted, freed slots don't merge back into contiguous space; they fragment. And because reclaiming any of it means a full gen-2 collection, the pauses get longer and land more often exactly as the working set grows. For a store whose entire pitch is low latency, that's the wrong kind of surprise.

Zaris ships an opt-in feature to sidestep the LOH entirely for these values. It's off by default, and this post is about what it does, how it works, and — just as important — when you should and shouldn't turn it on.

Reads That Scale With Your Cores

· 5 min read
Clustron Team
Distributed Systems Engineering

Zaris GET throughput scales with cores

Redis is fast. It is also, by design, single-threaded on the data path: one event loop, one core, one command at a time. That's a genuinely good design — it makes Redis simple to reason about and removes a whole category of concurrency bugs. But it has a ceiling you can't buy your way out of. Give a single-event-loop engine a 4-core box, an 8-core box, a 32-core box, and read throughput lands in the same place. The extra cores sit idle.

Zaris takes the other road. Keys are partitioned and each partition has an owner, so reads for different keys land on different cores and run at the same time. The result is throughput that moves when you add cores — instead of a flat line, a slope.