Skip to main content

Keyspace Notifications That Survive Failover

· 9 min read
Clustron Team
Distributed Systems Engineering

Redis keyspace notifications on Zaris's replicated native watch

Keyspace notifications are one of Redis's most useful features and one of its most quietly fragile. The idea is lovely: every time a key is written, deleted, or expires, the server publishes an event on a well-known pub/sub channel, and anyone who cares can PSUBSCRIBE and react — invalidate a cache, wake a worker, refresh a projection. No polling, no GET in a loop asking "has it changed yet."

The fragility is in the fine print. Redis keyspace notifications are fire-and-forget pub/sub, emitted by whichever node processed the command, to subscribers connected to that same node. In a clustered deployment that means a subscriber only hears about writes that happened to land on the node it's attached to, and if that node fails over, the subscription — and the stream of events it was carrying — goes with it. The feature works beautifully on a single box and gets subtle fast once there's more than one.

Zaris speaks the exact same protocol — CONFIG SET notify-keyspace-events, the __keyspace@0__ and __keyevent@0__ channels, a vanilla PSUBSCRIBE from any Redis client in any language. But the plumbing underneath is not a reimplementation of Redis's notifier. It's a thin bridge onto a feature Zaris already had: a native, cluster-wide, replicated watch. That changes what the notifications are worth.

Turning it on — exactly the Redis way​

There is nothing Zaris-specific in how you enable or consume the feature. Point any Redis client at the RESP front-end and use the commands you already know:

# Enable notifications for keyspace (K) and keyevent (E) channels
redis-cli -p 6380 CONFIG SET notify-keyspace-events KEA

# Subscribe to every key-event, cluster-wide
redis-cli -p 6380 PSUBSCRIBE '__keyevent@0__:*'
1) "pmessage"
2) "__keyevent@0__:*"
3) "__keyevent@0__:set"
4) "user:42"

From a typed client — here StackExchange.Redis — it's the same shape, no Zaris SDK involved:

var mux = await ConnectionMultiplexer.ConnectAsync("localhost:6380");
var sub = mux.GetSubscriber();

await sub.SubscribeAsync(
new RedisChannel("__keyevent@0__:expired", RedisChannel.PatternMode.Literal),
(channel, key) => Console.WriteLine($"key expired: {key}"));

The point worth underlining: no application change. Code written against Redis keyspace notifications runs against Zaris unmodified. That's the whole promise of the RESP front-end — your clients don't need to know what's behind the socket.

The two channel families​

Zaris publishes the two standard Redis channel families, with the standard payload convention — which one you subscribe to depends on whether you're keyed by the key or by the event:

ChannelPublished when flag…PayloadUse it when you care about…
__keyspace@0__:<key>K is setthe event name (set / del / expired)a specific key — "tell me everything that happens to session:abc"
__keyevent@0__:<event>E is setthe key that changeda class of events — "tell me every key that just expired"

The @0 is the database index, and in Zaris it is always 0 — Zaris doesn't carve the keyspace into numbered databases, so there is a single, flat notification namespace rather than sixteen. The three event names map to the three things that can happen to a key:

  • set — the key was written (a SET, an HSET, an LPUSH, any mutation)
  • del — the key was removed
  • expired — the key was evicted because its TTL elapsed

What actually happens underneath​

When you send CONFIG SET notify-keyspace-events KEA, Zaris doesn't light up a bespoke Redis event emitter wired into the command path. It registers one whole-keyspace prefix watch on the node's own routing client — the same Watch primitive the native .NET client exposes for reactive, poll-free reads. From that point on, the only logic that lives in the RESP layer is a naming map: a native watch event arrives, and the bridge turns it into the right Redis channel name and payload.

// The entire RESP-side translation, essentially verbatim:
var eventName = e.EventType switch
{
WatchEventType.Put => "set",
WatchEventType.Delete => "del",
WatchEventType.Evicted => "expired",
_ => null, // snapshots / heartbeats carry no keyspace meaning
};

if (keyspace) PubSub.Publish($"__keyspace@0__:{e.Key}", eventName);
if (keyevent) PubSub.Publish($"__keyevent@0__:{eventName}", e.Key);

That Publish call is the native pub/sub — the same subscribe/deliver machinery, with its own cluster-wide fan-out, that the store already runs. The RESP front-end owns the vocabulary (channel names, payload convention); the cluster owns the delivery. This is a pattern we hold to deliberately across the whole product: every feature is a native capability first, and RESP is a thin skin over it — never a parallel implementation with its own bugs.

Why "the watch is native" is the whole point​

Here is where the two systems diverge, and it's not cosmetic.

A Redis keyspace notification is produced on the node that ran the command and delivered to subscribers connected to that node. Nothing about the subscription is shared with the rest of the cluster, and nothing about it is durable. It is pub/sub in the strict sense — if you weren't listening on the right node at the right instant, the event simply didn't exist for you.

A Zaris keyspace notification rides a native watch, and a native watch is a cluster-wide object:

  • The subscription is replicated across members. WatchSyncService keeps watch subscriptions in sync between nodes — the leader pushes the current set to a joining node, and a node coming up asks the leader for what's active. The subscription isn't a fact known only to one process; it's cluster state.
  • It's re-homed across churn. When partitions move — failover, migration, a node leaving — the watch moves with the data it's watching, the same way key ownership does. The write still produces an event after the dust settles, because the watcher followed the partition.
  • Delivery fans out cluster-wide. Because the publish goes through native pub/sub, a client connected to any node receives events for writes that happened on any node. You are not quietly subscribed to one shard's worth of activity.

The practical upshot: a subscriber watching __keyevent@0__:expired to drive cache invalidation keeps working through a node failure. In Redis Cluster that same subscriber is attached to one node and sees one node's expirations; the honest mitigation there is to open a subscription to every node and de-duplicate. On Zaris it's one subscription, cluster-complete, and it survives the failover that would have silently severed it.

The honest edges​

This is a bridge over best-effort pub/sub, not a magic reliability upgrade. Three things to keep straight:

Delivery is still best-effort, not a guaranteed log. The bridge publishes each event fire-and-forget: a failure delivering to one subscriber must never stall the watch pipeline for everyone else, so a notification to a momentarily-disconnected client is dropped, not buffered. This matches Redis's own contract — keyspace notifications were never an exactly-once event log, and you should not treat Zaris's as one either. If a missed event would corrupt your state, notifications are the wake-up, not the source of truth: on the event, go read the authoritative value.

No backfill. The watch is armed with no initial snapshot, so you get changes from the moment notifications are enabled forward — never a replay of what the keyspace already contained. Enable notifications, then start reasoning about deltas.

Flag granularity is coarse. Redis lets you select event classes with a dense flag string (g, $, l, s, h, z, x, e, …). Zaris reads the two flags that decide where to publish — K for the keyspace channels, E for the keyevent channels (and A as the usual "all classes" shorthand) — and bridges the three native event types whenever either is set. It does not filter down to per-class subscriptions. In practice you subscribe to the channels and patterns you care about and let the pattern match do the narrowing, which is how most keyspace-notification code is written anyway.

If you need notifications scoped to part of the keyspace rather than all of it, set ZARIS_RESP_KEYSPACE_PREFIX on the node and the underlying watch arms on that prefix instead of the whole store — a cheap way to keep a busy keyspace from publishing events no one subscribes to.

When you want a log instead of a doorbell​

Keyspace notifications are a doorbell: they tell you something happened, right now, to whoever's listening. When what you actually need is a replayable, durable record — every event, in order, readable by a consumer that joined late or crashed and came back — that's not a notification, that's a stream. Zaris has native Redis Streams for exactly that: XADD to append, XREADGROUP with consumer groups for at-least-once processing with acknowledgements, blocking reads for low latency. Reach for notifications when a missed event just means "check again next time"; reach for Streams when a missed event is a bug.

Used for what it is — a low-latency, zero-polling wake-up that, on Zaris, happens to be cluster-complete and failover-durable — keyspace notification is one of the nicest things you can turn on with a single CONFIG SET. The difference is only visible on the day a node goes away, which is precisely the day you'd have discovered the single-node version's limits the hard way.