Skip to main content

One Primitive, Many Guarantees: Race-Free Coordination on Optimistic CAS

· 12 min read
Clustron Team
Distributed Systems Engineering

Race-free coordination on Zaris optimistic CAS

It's tempting to judge a data store by the length of its command list. Redis has INCR, SETNX, MULTI/EXEC, WATCH, Lua, SCAN, pub/sub — a verb for every occasion. The Zaris .NET client, by contrast, looks almost austere. There is no server-side atomic increment. There is no general multi-key transaction that runs arbitrary logic. There is no prefix scan. Native pub/sub exists only through the RESP front-end, not the typed .NET client.

What the .NET client gives you instead is one small, sharp coordination primitive: optimistic versioned compare-and-swap. This post is about how far that one primitive actually goes — because the honest answer is: surprisingly far. A rate limiter, a wallet that never goes negative under a thousand concurrent debits, a job queue where exactly one worker claims each job, and a feature-flag ruleset that never loses an update under concurrent admin edits. All of them, from the same primitive, with no extra server-side machinery.

The primitive​

Every value in Zaris carries a monotonic version. You don't set it; the store owns it and bumps it on every successful write. Reading gives you both the value and the version it was at:

KvResult<int> read = await client.GetAsync<int>("counter:api");
int value = read.Value; // the current value
long version = read.Version; // the version that value was written at

Writing back lets you attach a precondition on that version:

var opts = new PutOptions { IfMatchVersion = version };
KvStatus status = await client.PutAsync("counter:api", value + 1, opts);
// status == KvStatus.Ok -> your write committed; version is now version+1
// status == KvStatus.Conflict -> someone else wrote first; your value was stale

IfMatchVersion means commit only if the key is still at this exact version. If any other writer committed in the gap between your read and your write, the version has moved, your precondition fails, and you get KvStatus.Conflict instead of silently clobbering their change. There is one more flavour — IfAbsent = true, which means commit only if the key does not exist yet — a create-once operation we'll use for the job queue.

That's the whole primitive. A read that returns a version, and a conditional write that commits only against an unchanged version.

Why this is enough to serialise writers​

The reason this works is that the version is monotonic and store-owned. Two clients can read version 7 at the same instant. Both compute a new value. Both call PutAsync with IfMatchVersion = 7. The store processes writes to a given key one at a time, so one of them arrives first — it matches, commits, and the version becomes 8. When the second write is evaluated, the key is at 8, not 7, so its precondition fails and it gets Conflict. Exactly one writer wins; the other is told, precisely, that it lost. No lock was taken, no writer blocked waiting, and there is no window in which both writes "succeed" and one is silently lost.

The loser doesn't fail the operation — it just didn't win this round. It re-reads the now-current value and version, recomputes against fresh state, and tries again. That loop is the heart of every pattern below:

async Task<T> MutateAsync<T>(string key, Func<T, T> update, int maxAttempts = 50)
{
for (int attempt = 0; attempt < maxAttempts; attempt++)
{
KvResult<T> read = await client.GetAsync<T>(key);
T next = update(read.Value);

var status = await client.PutAsync(
key, next, new PutOptions { IfMatchVersion = read.Version });

if (status == KvStatus.Ok)
return next; // we won this round
// Conflict: someone wrote between our read and write. Loop and retry
// against the value they committed — never against our stale copy.
}
throw new ConcurrencyException($"{key} too contended after {maxAttempts} attempts");
}

Two things make this correct rather than hopeful. First, the retry always re-reads — you never retry with a stale value, so you can't resurrect an overwritten state. Second, update must be a pure function of the value you read. If it computes read.Value + 1, that's safe; if it reads the wall clock or a mutable field captured from outside, the retry won't be a faithful replay. Keep the transformation pure and the loop is a genuine read-modify-write that behaves as if the winning writers ran one after another.

Pattern 1 — a distributed rate limiter​

A rate limiter is an atomic counter with a ceiling. There's no INCR on the wire, but the CAS loop is an atomic increment: read the count, add one if there's headroom, commit against the version.

async Task<bool> TryAcquireAsync(string bucket, int limit)
{
for (int attempt = 0; attempt < 50; attempt++)
{
KvResult<int> read = await client.GetAsync<int>(bucket);
int count = read.Found ? read.Value : 0;
if (count >= limit)
return false; // over the limit, no write at all

var status = read.Found
? await client.PutAsync(bucket, count + 1,
new PutOptions { IfMatchVersion = read.Version })
: await client.PutAsync(bucket, 1,
new PutOptions { IfAbsent = true }); // first caller creates it

if (status == KvStatus.Ok)
return true; // we took a slot
// Conflict -> another caller moved the count; re-read and re-check.
}
return false;
}

Point a thousand concurrent callers at the same bucket with limit = 100 and you get exactly 100 true results and 900 false — never 101, never 99. The losers of each CAS round re-read and discover the ceiling has moved; the moment the count reaches the limit, every subsequent caller short-circuits with no write. The IfAbsent branch handles the cold-start race where several callers find the bucket missing at once — only one create-once wins, the rest fall into the normal increment path on their next attempt.

Pattern 2 — a wallet that never goes negative​

Money is the demanding case, because the invariant is a guard, not just a tally: a debit must never drive the balance below zero, and two debits must never both succeed off the same starting balance. That's exactly the overwrite that CAS forbids.

async Task<bool> TryDebitAsync(string account, decimal amount)
{
for (int attempt = 0; attempt < 50; attempt++)
{
KvResult<Wallet> read = await client.GetAsync<Wallet>(account);
Wallet w = read.Value;
if (w.Balance < amount)
return false; // insufficient funds: reject, don't write

var next = w with { Balance = w.Balance - amount };
var status = await client.PutAsync(
account, next, new PutOptions { IfMatchVersion = read.Version });

if (status == KvStatus.Ok)
return true;
// Conflict -> a concurrent debit landed first. Re-read the real balance
// and re-check the guard against it, never against our stale copy.
}
throw new ConcurrencyException($"{account} too contended");
}

Fire a thousand concurrent debits of one unit each at an account holding 400 and exactly 400 succeed; the balance lands on exactly zero and never dips below it. The guard holds because it is always re-evaluated against the committed balance after a conflict. A debit that was valid against a stale balance of 5 doesn't get to commit against a real balance of 0 — its version no longer matches, so it loops, re-reads 0, and correctly rejects. This is the pattern's sharpest edge: the precondition isn't just protecting the write, it's forcing the business rule to be checked against truth.

Pattern 3 — a job-queue claim​

When N workers race to process the same job, you need exactly one to win the claim. This is a single CAS flip on a status field — or, in its simplest form, a create-once:

async Task<bool> TryClaimAsync(string jobId, string workerId)
{
var claim = new Claim { Owner = workerId, ClaimedAt = DateTimeOffset.UtcNow };
// IfAbsent: the claim key is created by exactly one worker; the rest see Conflict.
var status = await client.PutAsync(
$"job:{jobId}:claim", claim, new PutOptions { IfAbsent = true });
return status == KvStatus.Ok;
}

Every worker that loses gets Conflict and moves on to the next job — no polling, no lock lease to renew in the happy path, no tie to break. If instead you model the job as a record with a Status field, the claim is a version-guarded flip from Pending to Claimed using the MutateAsync loop, which additionally lets you re-claim a job whose previous owner crashed (read it, see it's Claimed but stale, flip it back) — the same primitive, one field deeper.

Pattern 4 — a feature-flag ruleset with no lost updates​

Two admins open the feature-flag console and edit the same ruleset at the same time. Admin A toggles a flag on; Admin B adds a new targeting rule. With a naive "read JSON, edit, write JSON back" flow, whoever saves second wins and silently erases the other's change — the classic lost update. CAS turns that into a detected conflict instead:

async Task SaveRulesetAsync(string flagKey, Func<Ruleset, Ruleset> edit)
{
for (int attempt = 0; attempt < 10; attempt++)
{
KvResult<Ruleset> read = await client.GetAsync<Ruleset>(flagKey);
Ruleset updated = edit(read.Value); // apply this admin's change

var status = await client.PutAsync(
flagKey, updated, new PutOptions { IfMatchVersion = read.Version });

if (status == KvStatus.Ok)
return;
// Conflict -> the other admin saved first. Re-read THEIR version and
// re-apply this admin's edit on top, so neither change is lost.
}
throw new ConcurrencyException($"{flagKey} edited concurrently; please retry");
}

The second admin's save conflicts, re-reads the ruleset including the first admin's change, re-applies its own edit on top, and commits. Both changes survive. The important nuance: this works cleanly when the two edits touch different parts of the ruleset, because re-applying is composable. If both admins edit the same rule, the second save is where you surface a "someone else changed this" prompt to a human rather than blindly merging — the conflict is a feature, giving you the hook to ask.

The idempotency corollary​

Because a conflict means "retry," and a retry means "do it again," you have to make sure again is safe. The wallet debit above is the cautionary case: if your network call ambiguously fails after the store committed but before you got the Ok, a blind retry could debit twice. The CAS-native fix is an own-index / dedupe key: record the operation under a unique id with IfAbsent, and check it before acting.

// Create-once the operation record; if it already exists, this is a retry of
// a debit that already happened — return its recorded outcome, don't repeat it.
var seen = await client.PutAsync(
$"op:{operationId}", new OpMarker(), new PutOptions { IfAbsent = true });
if (seen == KvStatus.Conflict)
return await ReadRecordedOutcomeAsync(operationId); // idempotent replay

The same primitive that serialises your writers also gives you the exactly-once hook. You're not reaching for a second mechanism — IfAbsent is the dedupe.

Where this stops — and what to do about it​

CAS is not free, and pretending otherwise would be the marketing fluff this post is trying to avoid. Three honest limits:

Hot-key contention. Every pattern above funnels concurrent writers through one key. The more writers, the more conflicts, the more retries, and throughput on that single key degrades. CAS wins exactly one writer per round by design — so a key under heavy write contention becomes a serialisation point. The fix is to shard the key: a rate limiter bucket per (client, second) instead of one global bucket; a wallet split into sub-balances reconciled periodically; a job queue partitioned so workers contend on disjoint key ranges. Spread the contention and the retries fall away.

O(n) whole-value rewrites. A CAS write replaces the entire value, not a field. A feature-flag ruleset with ten thousand rules is re-serialised and re-written on every single edit — you pay for the whole object to change one flag. For small values this is a non-issue; for large, frequently-mutated aggregates it's a real cost, and the answer is again to decompose: store independently-mutable parts under separate keys so each edit rewrites only what changed, and reserve the single-key aggregate for things that are genuinely small or rarely written.

No cross-key atomicity. CAS is atomic on one key. If an invariant spans two keys — move money from account A to account B — a single CAS can't cover both. You either model the transfer as a state machine on one key (a "transfer" record that both sides observe), or accept that you need a saga with compensation. Don't reach for CAS to fake a two-key transaction; that's where it genuinely doesn't reach.

One primitive, earned its keep​

The short command list turns out not to be a poverty. A store-owned monotonic version plus a version-guarded conditional write is enough to express an atomic counter, a guarded decrement, a one-winner claim, and a lost-update-free edit — four named distributed-systems problems, one primitive, no locks held and no writer blocked. What you give up is the convenience of a purpose-built verb for each; what you get is a single mechanism whose correctness you can reason about once and then trust everywhere, plus the dedupe hook for idempotency riding along for free.

The discipline it asks of you is small and worth internalising: keep the transform pure, always re-read on conflict, check your invariant against the committed value, and shard before a key gets hot. Do that, and one primitive carries a remarkable amount of coordination.