Skip to main content

Reading Your Own Writes: The Zaris Consistency Model, Stated Honestly

· 12 min read
Clustron Team
Distributed Systems Engineering

Reading your own writes — one owner per partition

There's a question every developer eventually asks a distributed store, usually right after a confusing bug report: if I write a value and then immediately read it back, am I guaranteed to see what I wrote? It sounds like it should have a one-word answer. It doesn't — not for any partitioned, replicated store, and not for Zaris. The honest answer is "yes, within a partition, because reads go to the same single owner your write did" — and the interesting part is everything packed into that qualifier.

This post is the consistency model for Zaris, assembled in one place. We've written about the pieces separately — optimistic compare-and-swap, transparent retry under ownership churn, the replayable op-log that feeds replicas, multi-key transactions. Here we tie them together from the reader's side and answer the only question that matters at a keyboard: when I read, what am I actually promised? Including, especially, where the promise stops.

The model in one sentence​

Zaris is an AP-leaning store with the consistency boundary drawn around ownership, not around every write. The data plane stays available — a lone node keeps serving the segments it holds — and the single-writer discipline that keeps reads sane is enforced in the ownership plane: there is exactly one active owner per partition at any instant. Writes to a partition are serialized through that one node. There is deliberately no quorum or consensus on the write path, and Zaris does not claim global strict linearizability.

Almost everything a reader observes follows from that one structural fact, so it's worth making the chain explicit:

key → segment → partition → one runtime active owner

A key hashes deterministically to a segment; a segment belongs to a partition; a partition has exactly one active owner that accepts its writes and serves its reads. Ownership is a lease carried on the partition map as (LeaderTerm, OwnershipGeneration) — granted by the cluster leader, durable across restarts, and propagated to every node. That lease is what guarantees the "exactly one" — it replaced an older highest-epoch-wins scheme that could reset on restart and briefly produce two nodes each believing they owned the same partition.

Why owner-routed reads give you read-your-writes​

Here is the part that answers the opening question. In Zaris, reads route to the active owner, the same node your write went to. Clients do not read from replicas by default. So when you PutAsync a key and then GetAsync it, both requests resolve — through the same routing table, built from the partition map plus the runtime-ownership snapshot — to the same owner node, and that node already has your write in its local store.

That's the whole mechanism. Read-your-writes isn't bolted on as a session guarantee or a client-side cache; it falls out of the routing. You read your writes within a partition because a single node is the authority for that partition's reads and writes both, and writes to it are serialized one at a time. The replica in that diagram might lag for a few milliseconds behind the owner — replication has an asynchronous tail — but by default you never read the replica, so that lag is invisible to you.

The qualifier "within a partition" is load-bearing, and we'll come back to what it excludes.

The version you're allowed to see​

Every value the owner serves carries two monotonic numbers, and they're what make "the same value" a precise claim rather than a vibe:

  • a store-owned Revision — a strictly increasing sequence number the owner stamps on every mutation in commit order, and
  • a per-item ItemVersion — the version of that specific key, which is what IfMatchVersion guards in a compare-and-swap write.

Because the owner assigns these in commit order and serves reads from the same committed state, successive reads of a key from the owner move the version forward, never backward. You won't read version 42, then read version 40 a moment later. That monotonicity is exactly what lets the CAS loop be correct: a reader sees a version, computes against it, and writes back conditioned on it still being that version. If the model let reads flap between versions, optimistic concurrency would be built on sand. One owner, commit-ordered revisions, owner-served reads — that's the foundation the entire coordination story in the CAS post stands on.

The window where a read honestly says "not yet"​

Owner-routed reads are clean in steady state. The interesting moments are the transitions — a failover, a migration, a preferred-primary failing back — because during a handoff the node you'd route to may not currently be the authority. Zaris's choice here is the honest one: rather than serve a read from a node that isn't the settled owner, it returns a retryable status. A read (or write) during such a window can come back as:

  • Moved — your routing table is stale; the key lives on a different node now.
  • OwnershipChanging — ownership for that segment is actively in transition right now; it's a handoff, not a settled routing bug.
  • Migrating — the segment is mid-migration to a new owner.
  • ReplicaSyncPending — this node is bringing a segment's copy up to date and won't serve a possibly-incomplete read from it yet.
  • Unavailable — the partition has no reachable owner at this instant.

None of these is an exception you have to parse out of a stack trace — they're machine-readable outcomes, and the .NET client treats them as transient and retries with a reroute, refreshing its cluster map and going to whoever the map now names as owner. The effect is that your GetAsync just takes a few milliseconds longer and then returns the right value. The design decision underneath is worth stating plainly: Zaris would rather make you wait a beat than hand you a read it isn't sure is authoritative. A retryable "not yet" is a better answer than a confident wrong one.

This is also why read-your-writes survives a failover: the write-fence work means an acknowledged write isn't left behind in the handoff window (we'll get to the honest edge of that below), and the promotion only makes a node the owner once it's eligible — so the read that reroutes after a failover reroutes to a node that has your data.

The dial you can turn — and what it actually does today​

Zaris exposes a read-consistency surface. GetOptions carries a ReadConsistency of Leader, Quorum, or Any, and a store has an AllowReplicaReads flag (default off). The intent is the classic trade: let reads be served from replicas to spread read load, in exchange for accepting that a replica may be behind the owner — a bounded staleness you opt into deliberately.

Here is the honest current state, because it's exactly the kind of thing marketing copy would blur: today, reads are owner-only. AllowReplicaReads and ReadConsistency.Any are a configuration and request-shaping surface that exists and flows through the store definition, but there is not yet a server code path that actually serves a read from a replica. In practice that means you get the strong end of the dial right now — every read is an owner read, so read-your-writes holds — and you do not currently get the stale-replica-read behavior, even if you flip the flag.

Why mention a dial that's effectively pinned to one setting? Because it tells you what the consistency contract will let you trade, and because the asymmetry matters when you reason about staleness. When replica-serving reads land, turning AllowReplicaReads on is the moment you'd stop being guaranteed read-your-writes: a read routed to a replica can observe a version behind the owner's, for as long as the asynchronous replication tail is behind. Until then, the guarantee is simpler and stronger than the knob suggests. We'd rather you know that than assume you're already load-balancing reads across replicas when you aren't.

What "acknowledged" means for the next reader​

Read-your-writes is about you reading your own write on the owner. A different and sharper question is: once the owner tells you a write is acknowledged, is that write safe if the owner then dies, so that the reader who reroutes to the new owner still sees it?

The honest ack contract today: a write is acknowledged when it lands in the active owner's local store. Replication to the RF−1 replicas rides an asynchronous tail behind that ack. The consequences are worth planning around rather than hand-waving:

  • At redundancy 1 — no synced replica — a write is genuinely best-effort. It must not be treated as single-fault-durable, because there is exactly one copy.
  • Even with replicas, if the sole active owner is lost in the narrow window after it acked a write but before that write reached a replica, the acked write can be lost.

The stronger property — acknowledge only once the write is on at least one synced replica, so an acked write survives the loss of any one node — is the designed target (the "W1" ack-gate), not yet the shipped ack semantics. It's the direction the durability backbone work is consolidating toward, and that post is careful about the same distinction: the strong claim it validates is scoped to acknowledged writes, when a synced replica exists, under the failure classes the chaos run induces — not an unconditional "nothing can ever be lost." The op-log and migration mechanics are how a replica or a new owner catches up to the owner's committed state in the first place; the ack-gate is about when the owner is allowed to say "done."

So the precise reader-facing statement is: read-your-writes on the owner is solid; cross-fault durability of an acked write is strong-and-getting-stronger, with the current caveat stated above, and a hard floor at RF1 that no store can engineer away — one copy is one copy.

Across partitions, there is no shared clock​

The last place the "within a partition" qualifier bites: there is no global ordering or snapshot across partitions. Each partition is serialized through its own owner; two different partitions are two different owners with two independent revision sequences. If you write key A (partition 1) and then key B (partition 2), a concurrent reader can observe B's new value and A's old one, or vice versa — there is no cluster-wide instant at which "both happened." Bulk and batch operations are per-key underneath; they are convenient, not atomic.

When you genuinely need an invariant to span keys — move a balance from A to B, keep an index in step with the record it points at — that is exactly what multi-key transactions are for. They provide Serializable isolation over a client-coordinated commit across optimistic versions, and they are the only cross-key atomicity surface. Ordinary multi-key reads and writes are not a transaction and shouldn't be mistaken for one.

What is, and isn't, guaranteed​

Put plainly, so you can hold the whole model at once:

PropertyZaris today
Owner per partitionExactly one active owner, held as a durable lease.
Read pathRoutes to the active owner; replicas not read by default.
Read-your-writesYes, within a partition — same owner serves your write and your read.
Monotonic reads of a keyYes — owner serves commit-ordered revisions; versions move forward.
Reads during handoffRetryable status (Moved / OwnershipChanging / Migrating / ReplicaSyncPending), client retries and reroutes.
Replica reads / staleness dialAllowReplicaReads + ReadConsistency.Any exist as a surface; owner-only in practice today (no replica-serving path yet).
Ack durabilityOwner-local today; single-fault-durable-with-synced-replica is the target (W1 ack-gate).
Cross-partition orderNone — no global snapshot; use transactions for cross-key atomicity.
Global linearizabilityNot claimed.

The honest shape of it​

A good consistency model isn't the one with the longest list of guarantees; it's the one whose guarantees you can actually name and whose edges you can actually see. Zaris draws its boundary around ownership and is candid about the trade: you get read-your-writes and monotonic reads within a partition for free, because one owner serves both sides and hands you commit-ordered versions; you get a retryable "not yet" instead of a confident wrong answer during a handoff; and you give up global linearizability and cross-partition order, which you buy back explicitly with a transaction when a particular invariant needs it.

The two caveats worth carrying in your head are the ones this post stated rather than smoothed over: the replica-read dial is, today, effectively pinned to the strong setting, so you're more consistent than the knob implies; and "acknowledged" currently means on the owner, with full single-fault durability a target the durability work is landing, and a genuine floor at redundancy 1. Know those two things and the rest of the model is exactly as simple as it looks: find the owner, read your writes.

For the deeper design, see Consistency and Partitioning and replication in the docs.