The Memory Ceiling: Why a Zaris Node Never Grows Without Bound
Here is a question that sounds simple and isn't: what happens when an in-memory store runs out of memory?
The naive answer is "it crashes" — the process climbs until the operating system kills it. The slightly-less-naive answer is "it refuses writes." Zaris does neither by default. A Zaris node is an in-memory store, but it is deliberately not an unbounded one: each node is a bounded buffer with a hard ceiling, and when the data it owns approaches that ceiling the node evicts its coldest data to make room. Writes keep succeeding; the limit holds; the process doesn't balloon. The node behaves like a cache, because underneath, it is one.
Every node is a bounded buffer
A Zaris node does not grow without limit. Each node continuously accounts for the size of the data it actively owns and holds that running total under a configured ceiling. The accounting is never optional — a node is always a bounded buffer. The only things you configure are how big the ceiling is and, in principle, what the node does when it reaches it.
The default ceiling is 1 GiB of actively owned data per node. There is, by design, no "unbounded" setting. If you don't set a ceiling, you get the default — you never accidentally get infinity.
One subtlety matters for how you reason about a cluster's footprint: the ceiling covers the data a node actively owns — the primary copies it serves to clients. Replica copies a node holds for other partitions are tracked separately and don't count against the active-owner ceiling. So a node's real resident memory is its owned working set plus whatever replica copies it carries for its peers; the ceiling governs the former, not the sum.
Two ways to hold the line
There are exactly two things a bounded buffer can do when a write would push it past its ceiling: drop something old to make room, or refuse the new write.
| Mode | Behavior at the ceiling | A write that would exceed it |
|---|---|---|
| Eviction enabled (default) | Evicts older data to make room | Admitted — colder data is dropped |
| Eviction disabled | Refuses new data | Rejected with KvStatus.CapacityExceeded |
With eviction on, a node is a cache: under pressure it drops the least-valuable data so writes keep flowing. With eviction off, it's a reject-at-capacity store that protects existing data by turning new writes away.
Here is the honest part. In the current release a node always runs with eviction enabled, using the LRU policy. The no-eviction mode and the alternative policies (LFU, TTL-aware) exist in the engine, but they are not yet selectable through configuration. The practical consequence is that CapacityExceeded — the reject-at-capacity status — is essentially unreachable today: a full node evicts rather than refusing. It's a real, typed status you can write code against, but under the always-on LRU policy you won't normally see it. The one memory knob you actually control right now is the size of the ceiling itself, memory.maxSizeBytes.
The eviction run rides a quiet band
The interesting engineering isn't that a node evicts — it's how, because a naive "evict on every write once you're full" design thrashes miserably. Imagine a node sitting exactly at its ceiling: every single incoming write would trigger an eviction to make room for itself, and you'd pay that cost on the hot path forever.
Zaris avoids that with two moves. First, eviction is driven by a background loop, not by the individual write — the write that tips a node over the edge is admitted, and the loop reclaims space shortly after, off the request path. Second, the loop uses hysteresis: a high trigger and a lower target, so it reclaims in bursts and then goes quiet instead of firing continuously at the edge.
- Trigger: a run starts when active usage rises above 80% of the ceiling.
- Target: once started, the run keeps evicting until roughly 30% of the ceiling is free again — until usage falls back to about 70%.
- Cadence: the loop wakes on a fixed interval (one second by default) and samples a bounded set of candidate items each cycle.
Because the post-eviction target (70%) sits strictly below the trigger (80%), each run drops usage into a quiet band rather than stopping right at the point that would immediately re-fire it. That's the sawtooth in the hero image: usage climbs to 80%, a run clears it back to 70%, and it climbs again — never pressed against the ceiling, never thrashing one eviction per write.
The eviction policy itself is LRU — least recently used — and it's sampled, not strict. Rather than maintaining a globally ordered list of every key (which would cost memory and lock contention on the hot path), Zaris stamps approximate recency signals in place on the read path and, when a run fires, samples a small set of candidates and drops the lowest-scoring ones. It's the same sampled approach Redis uses, and the same tradeoff: you give up perfect "evict the single oldest key" ordering in exchange for eviction that costs almost nothing to maintain and never needs a separate index.
The other ceiling: per-item size
The overall memory ceiling is a limit on the total. There's a second, independent limit on any single item: by default a value may not exceed 4 MiB. This is checked on every write regardless of eviction mode, and it's where a write hits a genuinely reachable rejection.
The distinction matters because the two failures mean opposite things and want opposite responses:
var result = await client.PutAsync("report:2026-07", payload);
if (!result.IsSuccess)
{
switch (result.Status)
{
case KvStatus.ItemTooLarge:
// This one value is too big. No amount of eviction helps —
// split it, or store a reference to it elsewhere.
break;
case KvStatus.CapacityExceeded:
// The node is full and eviction is disabled (not the default).
// Retry later or shed load.
break;
}
}
ItemTooLarge is a live, reachable outcome in every mode — eviction can't make room for a value that's over the per-item cap, because the cap isn't about free space, it's about one pathological value. CapacityExceeded, as covered above, only appears in the no-eviction mode you can't select yet. Treating them identically is a bug waiting to happen: retrying an ItemTooLarge write forever will never succeed.
A related wrinkle lives at the boundary between this per-item cap and the .NET runtime: values of about 85,000 bytes and up land on the Large Object Heap, so raising maxItemSizeBytes to admit big values is a deliberate decision with GC consequences, not a free dial. If you're hot-updating large plain values, chunking them into sub-LOH pieces is the mechanism that keeps the LOH from churning — a separate feature with its own tradeoffs, which we unpacked in Storing Large Objects Without Churning the Heap.
What you configure, and what you don't
The memory ceiling lives in the store configuration under the memory section:
{
"memory": {
"maxSizeBytes": 2147483648
}
}
That's 2 GiB. Set maxSizeBytes to null, or leave it out, and you get the 1 GiB default. In the current release this is the memory setting you control — the LRU policy, the 80%/70% hysteresis thresholds, and the 1-second cadence are fixed defaults, not knobs. It's worth being clear-eyed about that: the blog posts and the engine describe a richer policy space than the configuration surface currently exposes, and this post describes what you can actually set today.
What it means for sizing a cluster
Because each node is a bounded buffer, a store's total capacity is roughly the per-node ceiling multiplied by the number of partitions a node owns, adjusted for the replication factor (replicas hold additional copies that live outside the active-owner ceiling). Two practical consequences fall out of that:
- Size the ceiling to your working set plus eviction headroom, not to your total data. A node runs an always-on LRU cache today, so cold data is dropped under pressure rather than rejected. If your working set fits under ~70% of the ceiling with room for the band to breathe, eviction stays quiet and the hot set stays resident. If you size the ceiling below your working set, you've built a cache that's constantly evicting data it's about to need again — correct, but slow, and the sawtooth turns into a grind.
- Don't forget the replica footprint. The ceiling governs owned data, but a node's actual RAM is owned data plus replica copies. When you provision memory per pod or per machine, budget for both, or a node sized "to its ceiling" will still run hotter than the ceiling implies.
For the capacity math in the context of partitions, replication factor, and pod counts, Sizing a Zaris Cluster for Durability on Kubernetes walks the full calculation.
The takeaway
An in-memory store that grows without bound isn't a feature — it's an outage waiting for a traffic spike. Zaris treats every node as a bounded buffer with a real ceiling (1 GiB by default, no unbounded option), an always-on sampled-LRU cache underneath it, and a hysteresis band that reclaims space in quiet bursts rather than thrashing one eviction per write at the edge. The total ceiling and the per-item cap are two different limits with two different typed statuses, and the one that's actually reachable today — ItemTooLarge — wants a different response than a retry loop. Size the ceiling to your working set, budget for replica copies on top, and the node holds its line without you watching it.
The complete model, including the policies the engine carries but doesn't yet expose, is in Memory and eviction; the exact fields and defaults are in the configuration reference, and every capacity status is catalogued in the status and exception reference.