Health and gossip
The hardest question in a distributed system is not "is that node alive?" but "how long do I wait before deciding it is dead?" Wait too little and a momentary hiccup evicts a healthy node, triggering a needless election and a membership flap. Wait too long and the cluster keeps routing work to a node that is already gone.
Nodus answers this with a heartbeat-based failure detector and a suspicion window, surfaced through node.Health and node.Cluster. On top of the same heartbeat traffic, it offers a gossip plane (node.Gossip) so a product can piggyback its own state summary without running a second protocol.
Heartbeats and failure detection
Every member sends a periodic heartbeat to its peers. Receiving one refreshes that peer's liveness; missing enough of them makes the peer suspect, and continued silence eventually evicts it.
The detector runs on three fixed intervals:
| Setting | Value | What it governs |
|---|---|---|
| Heartbeat interval | 5 s | How often a node emits its own heartbeat. |
| Heartbeat timeout | 12 s | How long without a heartbeat before a peer is suspected. |
| Suspicion window | 20 s | How long a peer stays suspected before it is evicted from membership. |
The suspicion window is the key to avoiding flap: a peer that misses a beat is not immediately removed. It is marked suspected, and only if it stays silent through the window is it declared failed and evicted — at which point MemberLeft fires and a re-election may run.
A node that leaves gracefully short-circuits this: on shutdown it stops heartbeating and broadcasts its departure, so peers evict it promptly instead of waiting out the window. This is why StopAsync produces a fast, clean MemberLeft, whereas a killed node takes up to the suspicion window to disappear.
Alive vs. reachable
Nodus tracks two distinct liveness facts about each peer, and the distinction is deliberate.
| Meaning | Query | |
|---|---|---|
| Alive | We have received traffic from the peer — it is actively live. | node.Cluster.IsAlive(nodeId) |
| Reachable | We can dial the peer — an outbound connection succeeded within the reachability window. | node.Cluster.IsReachable(nodeId) |
A peer can be reachable before it is alive: you can open a socket to it (reachable) before it has sent you anything (alive). Reachability is backed by a maintained set, so IsReachable is O(1) and never opens a socket on the calling path. Use ReachableNodeIds to enumerate everything you can currently dial, independent of liveness.
You can also inspect when you last heard from a peer:
DateTime? last = node.Cluster.LastHeartbeatUtc("node-2");
if (last is { } t && DateTime.UtcNow - t > TimeSpan.FromSeconds(15))
Console.WriteLine("node-2 is going quiet");
Health surface
node.Health (IClusterHealth) rolls the above into a small snapshot of this node, its peers, and cluster-level counters:
NodeHealthStatus me = node.Health.Local; // this node
IReadOnlyList<NodeHealthStatus> peers = node.Health.Peers; // each known peer
ClusterMetrics m = node.Health.Metrics; // cluster-level counters
Console.WriteLine($"active peers: {m.ActivePeers}");
foreach (var p in peers)
Console.WriteLine($"{p.NodeId}: {(p.IsAlive ? "alive" : "down")}, last seen {p.LastSeenUtc:HH:mm:ss}");
NodeHealthStatuscarriesNodeId,IsAlive,LastSeenUtc, andRoles.ClusterMetricscarriesActivePeers,TotalMessagesSent, andAvgLatencyMs.
The gossip digest rider
Heartbeats are already flowing between every pair of nodes on a steady cadence. The gossip plane lets a product piggyback its own compact state digest on that existing traffic — so peers can passively detect state drift without any additional RPC or a second protocol.
You supply a digest provider, and you observe peers' digests as they arrive:
// attach your digest to every outbound heartbeat
node.Gossip.SetDigestProvider(new MyDigestProvider());
// observe the digests peers attach to theirs
node.Gossip.DigestReceived += (fromNodeId, digest) =>
{
if (digest.Epoch < node.Leadership.Epoch)
Console.WriteLine($"{fromNodeId} is behind: epoch {digest.Epoch} vs {node.Leadership.Epoch}");
};
A provider implements IClusterStateDigestProvider, returning a ClusterStateDigest snapshot:
public sealed class MyDigestProvider : IClusterStateDigestProvider
{
public ClusterStateDigest GetCurrentDigest() => new()
{
Epoch = /* my current election epoch */,
LeaderId = /* who I think leads */,
MapVersion = /* my view version */,
// ... other compact scalars ...
};
}
ClusterStateDigest is intentionally small — a handful of scalars such as Epoch, LeaderId, MapVersion, and MapNodeCount, plus optional per-partition generation maps. It is a summary that lets a peer notice "your view and mine disagree," not a full state transfer. Keep the field count small; it rides on every heartbeat.
The heartbeat is the one message guaranteed to flow between every pair of nodes, continuously, whether or not there is application traffic. Attaching a digest to it gives you cluster-wide, always-on drift detection with no new connection, no new schedule, and no extra failure mode.
This is how Zaris keeps its partition map convergent across nodes: it rides its map version and per-partition ownership generations on the Nodus heartbeat and reacts when a peer's digest reveals a stale view.
Next steps
- Membership — what the failure detector adds to and removes from.
- Leader election — what a detected leader failure triggers.
- Querying health and membership — the task-focused how-to.