Skip to main content

The Worker-Pool Ceiling: When a Node Plateaus at 25% CPU

· 8 min read
Clustron Team
Distributed Systems Engineering

A node pinned near 25% CPU with flat throughput because the dispatch worker pool was the ceiling

Here is a benchmark result that looks like a contradiction. A Zaris node serving a read-heavy workload tops out at a little under half a million ops/sec and refuses to climb — add more client connections, more load generators, and the number doesn't move. The obvious diagnosis is "it's saturated." Except the CPU graph says otherwise: the node is sitting at roughly a quarter of its available CPU. Three cores out of four are idle, and throughput is flat.

That is not a core-bound node, and — if you've ruled it out carefully — it's not a client-bound benchmark either. It's a third thing, and it's a classic: the request-dispatch worker pool was narrower than the work wanted to be, so requests queued while cores sat on their hands.

First, rule out the usual suspects​

Two explanations have to go before this one is even on the table, because they're far more common.

Core-bound. If the node is at or near 100% CPU and throughput is flat, you're genuinely out of CPU — the answer is more cores or more nodes, and Zaris scales reads by adding cores. That's not this case: here the CPU is idle.

Client-bound. A single load generator saturates long before the server does, and then you're measuring your client, not Zaris. If throughput is flat as you add server cores, suspect your client first — add generators and connections until the server-side number actually responds. We covered this trap in the reads post, and it's the one to eliminate before going any further.

What's left after both of those is the interesting case: enough offered load to keep the node busy, plenty of CPU headroom, and throughput still flat. When the bottleneck is inside the server but it isn't CPU, the thing to look at is concurrency — how many requests the node will work on at once.

The dispatch pool is the concurrency limit​

Every request that arrives at a Zaris node is handed to a request-dispatch worker pool. A worker picks the request up, routes it to the partition that owns the key, runs the operation against that partition's in-process store, and writes the reply back. For a read, the operation itself is a microsecond-scale in-memory lookup — genuinely tiny.

The width of that pool — how many workers can be mid-request simultaneously — sets a hard ceiling on in-flight requests, and that is a different number from "how fast is one operation." The original default derived the width from a small multiple of Environment.ProcessorCount. On a 4-vCPU node, a couple-times-cores pool is only a handful of workers.

Here's the subtlety that makes the CPU graph look wrong. If every worker spent its whole slot burning CPU on the lookup, a pool sized to the cores would be exactly right and the node would hit 100% CPU. But workers don't spend their whole slot on CPU. A worker occasionally parks — an await continuation, a socket write that doesn't complete synchronously, a brief lock on a shared structure. A parked worker burns no CPU, but it still holds its pool slot. With only a handful of slots, a few parked workers are enough to stall the whole node: new requests queue behind them, and because the busy workers are mostly waiting rather than computing, the cores never light up. Flat throughput, idle CPU. The pool width was the ceiling, not the silicon.

The tell that distinguishes this from a client-bound benchmark: requests are queuing at the server (latency climbs, throughput is flat) while server CPU stays low. Client-bound looks different — the server is idle because nothing is arriving fast enough. Here, work is arriving; the node just won't take enough of it at once.

The fix: size the pool for the work, not the core count​

The repair isn't exotic. Reads are microsecond in-memory lookups, so the pool should be set wide by default — wide enough that a handful of parked workers can't starve the node — and it should be tunable, because the right width depends on the node and the workload, not on a fixed multiple of the core count.

The default pool width is now CPU count × 16, clamped to a sane band of 64–1024, with a floor of max(4, CPU count × 4). Two environment variables override it:

VariableDefaultWhat it does
ZARIS_SERVER_MAX_WORKERSCPU count × 16 (clamped 64–1024)Hard override (1–4096) for the dispatch pool width. Raise it when a node has spare CPU and requests are queuing.
ZARIS_SERVER_MIN_WORKERSmax(4, CPU count × 4)Lower bound (1–4096) on the pool width.
# A node with spare CPU and requests queuing — widen the pool.
ZARIS_SERVER_MAX_WORKERS=512 ZARIS_SERVER_MIN_WORKERS=64 \
dotnet Clustron.Zaris.Server.dll

One property worth calling out: this is a server-side change and nothing else. The wire protocol doesn't change, the client SDK doesn't change, and there's no NuGet bump to chase — you tune a node by restarting it with a different environment, and every existing client keeps talking to it unchanged.

Tune it honestly — widening isn't free​

The reflex after reading this is to crank ZARIS_SERVER_MAX_WORKERS to the ceiling on every node. Don't. A wider pool helps in exactly one situation and does nothing — or mildly hurts — in the others.

  • It helps when requests queue while CPU is idle. That's the case in this whole post: in-flight concurrency is the limit, so admitting more of it fills the cores. This is where you'll see throughput jump.
  • It does nothing when you're already CPU-bound. If the node is at 100% CPU, more workers won't conjure throughput out of nowhere — they just add context-switching overhead to a machine that has no spare cycles. More workers is not more compute.
  • It can hurt if you take it to an absurd width. Thousands of workers thrashing a 2-vCPU node spend real time context-switching instead of serving. That's why there's a clamp, and why the knob tops out at 4096 rather than infinity.

The honest tuning loop is: watch for the signature — latency climbing and throughput flat while CPU sits well below saturation — raise MAX_WORKERS, and measure. If CPU climbs and throughput follows, the pool was your ceiling. If nothing moves, it wasn't, and you should put the knob back and look elsewhere (the network path, the client, or genuinely needing more nodes).

Measure CPU where it's actually true​

This whole investigation hinges on trusting the CPU number, and that number is easy to get wrong.

  • Read CPU from the cgroup, not from a cluster metrics-server estimate. On Kubernetes especially, the metrics-server figure lags and smooths — it will tell you a node is at 40% when the cgroup accounting says it's pinned, or the reverse. Diagnose a "flat throughput, idle CPU" case from the cgroup, or you'll chase a number that isn't real.
  • Local benchmarks are host-bound and will hide the ceiling. On a laptop, the load generator and the server fight over the same cores, so you never present the server with clean, sufficient load — the plateau you measure is your test rig, not Zaris. The real worker-pool ceiling showed up at scale on a managed cluster with the load driven from separate machines, not on a developer box. Validate throughput where the client can't be the bottleneck.

The takeaway​

A flat throughput curve has three very different causes, and the CPU graph tells you which one you're looking at. At 100% CPU, you're core-bound — add nodes. With the server idle because nothing arrives fast enough, you're client-bound — fix the test. But a node that's flat at a quarter of its CPU, with requests visibly queuing, is neither: it's a concurrency limit, and on Zaris that's the dispatch worker pool. The default is now wide enough that most workloads never meet it, and ZARIS_SERVER_MAX_WORKERS is there for the ones that do — to be raised deliberately, when the CPU graph has earned it.

For the complete environment-variable surface and the exact defaults, see the configuration reference. If you're sizing a cluster for throughput, the benchmarking guide walks the settings that actually move the numbers, and Reads That Scale With Your Cores covers the other half of the story — why, once the pool is out of the way, adding cores adds read capacity instead of hitting a single-event-loop wall.