Skip to main content

42 posts tagged with "Clustron"

Articles related to Clustron distributed key-value store.

View All Tags

Many Keys, Few Round-Trips: How Zaris Fans Out a Bulk Operation

· 12 min read
Clustron Team
Distributed Systems Engineering

Fanning a bulk operation out to its owning nodes

Ask a distributed store for one key and the cost is dominated by a single number: the round-trip to the node that owns it. The lookup itself is nanoseconds; the network is microseconds. Now ask for a thousand keys. If you loop and call Get a thousand times, you pay that round-trip a thousand times, serially, and the store spends almost all of the wall-clock sitting idle waiting for packets. The keys might be spread across four nodes that could have answered in parallel — but a naive loop never gives them the chance.

Zaris has three bulk APIs — GetManyAsync, PutManyAsync, DeleteManyAsync — and a mixed ExecuteBatchAsync underneath them, and their whole reason for existing is to turn "a thousand keys" into "one request per owning node, issued in parallel." This post is the fan-out path end to end: how the client buckets keys by owner, why all four APIs collapse onto a single engine, how an Index field keeps the answers in order through two layers of concurrency, and why — unlike a multi-key transaction — a bulk operation is per-item and partial success is normal.

One Hop to the Owner: How the Zaris Client Routes a Key

· 12 min read
Clustron Team
Distributed Systems Engineering

How the Zaris client routes a key to its owner

There are two honest ways to build a distributed cache client. The first is to pick any node, send it the request, and let the cluster forward it internally to whoever owns the key — simple on the client, but every request that lands on the wrong node pays for an extra network hop inside the cluster. The second is to make the client smart enough to know which node owns the key before it sends anything, so the first socket it touches is the one that can actually answer. Zaris takes the second path. A GET leaves your process already addressed to the node that holds the key — no proxy, no forwarding hop, no coordinator in the middle.

This post is the inside view of how that works: the function that turns a key into an owner, the lookup table the client keeps in memory, and — the part that's easy to get wrong — how the client stays correct when ownership moves while your traffic is in flight.

An Owner That Restarts Has No Right to Its Own Emptiness

· 13 min read
Clustron Team
Distributed Systems Engineering

Rebuilding a restarted owner from its replicas

A failover covers the moment an owner dies: the cluster picks a survivor and keeps serving. A planned migration covers moving a healthy owner's partition on purpose. This post is about the third act, the one that runs after the sirens stop — the original owner process restarts, rejoins the cluster, and wants its partition back. Its on-disk-nothing, in-memory-everything store came back empty. The data it is responsible for is still alive, but it is alive on the replicas that kept serving while it was down.

The whole correctness problem is contained in one sentence: the restarted owner cannot be allowed to trust its own emptiness. An empty store looks exactly like a store someone deliberately cleared. If the node assumes "I'm the owner and I have no data, therefore this partition is empty," it will happily serve NotFound for millions of live keys and, worse, let its empty state flow back onto the replicas and delete them. This post walks the five small, pure policies Zaris uses to make sure that never happens.

Two Clusters, One Change Feed: Replicating Across a WAN Without Stretching Zaris

· 10 min read
Clustron Team
Distributed Systems Engineering

Two independent clusters linked by a resumable change feed across a WAN

Sooner or later someone asks the question: we have a cluster in one region and users in another — can we just add a few nodes over there and let one cluster span both? It is a reasonable instinct. One cluster, one namespace, one connection string, data everywhere. The instinct is also a trap, and it is worth being precise about why.

A Zaris cluster is engineered for a datacenter network. Membership is tracked with heartbeats on short timers. Ownership of a partition moves between nodes through a handshake that assumes the other side answers quickly. Failover fences a suspected-dead owner and promotes a replica on the order of seconds. Every one of those mechanisms is calibrated for a link where the round trip is measured in microseconds and a "slow" node is genuinely sick — not merely eighty milliseconds and an ocean away. Put half the nodes across a WAN and you haven't built a geo-distributed cluster; you've built one cluster that now mistakes normal WAN latency for failure, flaps ownership, and risks split-brain the first time the link between regions hiccups.

So the right shape is not one stretched cluster. It is two independent clusters — each self-contained, each with its own ownership and failover staying entirely on its own LAN — joined by an asynchronous, resumable change feed. This post is about that pattern: why independence is the load-bearing decision, and how Zaris's operation log already hands you the one thing a WAN replicator actually needs — a cursor you can resume from.

How a Zaris Cluster Comes Up: Membership, the Map, and Convergence-Gated Readiness

· 14 min read
Clustron Team
Distributed Systems Engineering

How a Zaris cluster comes up: membership, the map, and convergence-gated readiness

Most writing about distributed stores is about steady state — reads, writes, replication — or about the dramatic failures: a failover when an owner dies, a cut-over when a partition moves. But there's a quieter moment that decides whether any of that works: the moment the cluster comes up. Before a single key can be routed, the nodes have to agree on who is in the cluster, who owns which slice of the keyspace, and — the step most systems get wrong — whether it is safe to start serving at all.

This post follows that boot sequence in the order Zaris actually runs it: membership, the partition map, and the readiness gates that stand between "the process is up" and "the node is taking traffic." It also draws a line, honestly, around what Zaris itself implements versus what it inherits from the underlying cluster runtime it's built on.

The Latency Floor: One Microsecond of Work Inside Ten of Plumbing

· 11 min read
Clustron Team
Distributed Systems Engineering

One microsecond of work inside ten microseconds of plumbing

Here is a number that reframes how you think about a key/value store. When we isolate the Zaris in-memory store and feed it operations with nothing else in the way — no socket, no serialization, no routing — a single thread serves reads at over a million per second, and scales cleanly to more than eight million across eight threads. The store, in other words, does the actual work of a GET in under two microseconds, and it is emphatically not the bottleneck.

And yet a full GET over a real TCP connection, on the same machine, plateaus at a few hundred thousand per second. The store can do the work five times faster than the server can deliver it. So the interesting performance question about Zaris is not "how fast is the store?" It's the opposite: where does all the time between the socket and the store go? This post is the answer, traced operation by operation, because the gap between 1.7 microseconds of logic and 10-to-13 microseconds of server CPU per operation is not noise — it's the latency floor, and almost all of it is plumbing.

Scanning a Keyspace That Won't Hold Still: How SCAN Iterates a Cluster

· 10 min read
Clustron Team
Distributed Systems Engineering

Many sorted segments merged into one ordered, paged stream

A single-key GET is easy to reason about: hash the key, find its owner, ask that one node. SCAN is the opposite kind of operation. It has no key — it asks a question about the whole keyspace: give me every key in this range, a page at a time. And in a partitioned, replicated store that keyspace isn't one thing in one place. It's split across 256 segments spread over however many nodes you run, each node holding only the slices it currently owns, and the whole set is being written to while you iterate. There is no global sorted index to walk. So where does the ordering come from, what exactly does a resume token point at, and what happens to your scan when a key is written — or a partition moves — halfway through it?

This post follows a SCAN down through the three layers that actually serve it: a single segment, a single node, and the client that stitches the nodes together. Each layer solves one piece of the problem, and the honest limits of the operation fall out of how those pieces fit.

The Memory Ceiling: Why a Zaris Node Never Grows Without Bound

· 9 min read
Clustron Team
Distributed Systems Engineering

A Zaris node holding memory usage in a quiet band below its ceiling, evicting at 80% down to 70% so writes never cross the limit

Here is a question that sounds simple and isn't: what happens when an in-memory store runs out of memory?

The naive answer is "it crashes" — the process climbs until the operating system kills it. The slightly-less-naive answer is "it refuses writes." Zaris does neither by default. A Zaris node is an in-memory store, but it is deliberately not an unbounded one: each node is a bounded buffer with a hard ceiling, and when the data it owns approaches that ceiling the node evicts its coldest data to make room. Writes keep succeeding; the limit holds; the process doesn't balloon. The node behaves like a cache, because underneath, it is one.

Migrating a Partition With Zero Data Loss: The Cut-Over Protocol

· 13 min read
Clustron Team
Distributed Systems Engineering

The Zaris partition cut-over protocol

A failover is what happens when an owner dies — the cluster has seconds to pick a survivor and the partition map never moves. Migration is the calmer, more dangerous cousin: the owner is perfectly healthy, and you want to move a partition off it on purpose — to grow the cluster, to rebalance, to drain a node you're about to retire. Nothing has failed, which is exactly why the bar is higher. A failover is allowed a few honestly-lost tail writes in the worst case; a planned migration is not. If moving a partition from node A to node B can lose a single acknowledged write, migration isn't a feature, it's a liability.

This post walks the cut-over protocol Zaris actually runs, in order, and spends most of its time on the one step everyone gets wrong: the moment the source is finally "free" to drop its copy. That step is a loaded gun, and the interesting engineering is the safety on it.

What Crosses the Wire: The Anatomy of a Zaris Record

· 12 min read
Clustron Team
Distributed Systems Engineering

The anatomy of a Zaris stored record

Most of what Zaris does well — replicating a write to a backup, catching up a lagging replica, cutting a partition over to a new owner — comes down to one quiet question: what, exactly, is the unit of data that moves between nodes? If that unit is right, convergence is almost boring. If it drops a field or carries the wrong one, you get a resurrected deleted key, a replica that can't tell stale from fresh, or one node's bookkeeping leaking onto another.

In Zaris that unit is a single type — StoredRecord — and it is worth taking apart field by field, because every decision in it is load-bearing. This post is the anatomy: what's in the record, how it's framed on the wire, why a delete travels as a record instead of as an absence, why there are two different version numbers, and — just as important — what is deliberately left behind on each node and never crosses the wire at all.