Skip to main content

The Control Plane in Your Shell: Running a Zaris Cluster from PowerShell

· 11 min read
Clustron Team
Distributed Systems Engineering

Provisioning and running a Zaris cluster from PowerShell

Most of what you read about a data store is about the data path — reads, writes, the wire protocol, how fast it goes. But before any of that matters, somebody has to stand the thing up: decide how many partitions it has, how many copies of each, where the ports live, start it, confirm it came up healthy, and watch it while it runs. That's the control plane, and in Zaris it lives in PowerShell.

This is a day-one operator walkthrough. We'll create a workspace, provision a store, start it, and then read its live state back — all from the AdminShell module. No web console required, nothing hand-edited. And because the whole point of an operator story is trust, we'll be explicit about the choices New-ZrStore puts in front of you and what each one actually costs.

Two shells, two planes​

Zaris ships its command-line surface as two distinct PowerShell modules, and the split is not cosmetic — it mirrors the two planes of the system:

  • AdminShell is the control plane. Its cmdlets provision, start, stop, inspect, and scale stores. They talk to the Zaris Management Service (the "manager"), not to the data nodes directly. These are the New-ZrStore, Get-ZrNode, Watch-ZrStoreMetrics verbs.
  • ClientShell is the data plane. Its cmdlets — Connect-ZrStore, Get-ZrItem, Set-ZrItem — open a connection straight to the store's nodes and move key-value data. That's the surface an application uses, and it's the subject of the connection strings post.

An operator standing up a cluster lives almost entirely in AdminShell. The manager it talks to is the component that knows the desired shape of every store and drives the real processes toward it. So the first thing to do is tell your shell which managers to talk to.

Step 0 — a workspace to target​

A workspace is a named set of managers that admin cmdlets target, persisted under ~/.clustron so it survives sessions. You create one once and switch into it whenever you open a shell:

# Create a workspace pointing at the cluster's managers (does not activate it)
New-ZrWorkspace -Name "prod" -Managers "10.0.0.10:7801","10.0.0.11:7801"

# Make it the active target for this and future sessions
Use-ZrWorkspace prod

Each manager endpoint is an observer of one cluster — list all of them and the admin cmdlets aggregate state across the whole cluster rather than whatever one machine happens to know. If your cluster has security enabled, record an admin token on the workspace (-Token) or authenticate the session with Connect-ZrManager -Credential (Get-Credential) -Force, which exchanges a username and password for a session token. (Connect-ZrManager is the older ad-hoc connect and still works, but the workspace cmdlets are the current path — they're what the web console's saved workspaces map to.)

One honest fork in the road here: the walkthrough below is the supervisor model, where each manager forks and owns its local Zaris processes — the native/Windows deployment. If your nodes instead run under an orchestrator like Kubernetes (the attach model), the manager doesn't create processes; you define the store with your orchestrator and adopt it with Register-ZrStore instead of New-ZrStore. New-ZrStore detects an attach-mode manager and tells you so rather than failing deep in the stack.

Step 1 — provision the store​

This is the cmdlet with the real decisions in it. The minimal production form is short:

New-ZrStore -Name "orders" `
-ReplicationFactor 2 `
-BaseClusterPort 7811 `
-BaseClientPort 7861

That creates a store named orders across every machine in the workspace, with two copies of each partition. Four parameters carry the weight, so let's be precise about each.

-ReplicationFactor is the total number of copies of each partition — primary plus replicas. For production it must be at most the number of machines, so each copy lands on a different machine; that's what makes it fault tolerance rather than decoration. On a single dev box the constraint is relaxed with a warning — replicas will share a machine, which is fine for testing and buys you nothing in a real outage. The cmdlet says so out loud rather than letting you believe otherwise.

-PartitionsPerNode defaults to 1, which is the production topology. The total partition count is derived from live cluster membership — PartitionsPerNode × live nodes — and recalculated on every membership change, so it isn't a number you freeze at creation time. You raise it above 1 mainly for local testing, where you want to watch partition-aware behaviour (segment routing, rebalance on scale-out) on a single machine without standing up a rack.

-BaseClusterPort / -BaseClientPort are the first ports in two ranges: the internal gossip/cluster port and the client-facing port, incremented once per process on each machine. The cmdlet validates that the two ranges don't overlap and refuses to create a store that would collide with itself — a small guard that saves a confusing first-boot failure.

There are three more that are worth knowing because they encode a genuine durability trade-off:

New-ZrStore -Name "orders" -ReplicationFactor 3 `
-BaseClusterPort 7811 -BaseClientPort 7861 `
-ReplicationMode Sync `
-WriteQuorumPolicy Majority `
-AllowReplicaReads
  • -ReplicationMode is Async by default: the primary writes locally and returns immediately, replicating in the background. That's the fastest option and the right default for a cache — but be clear-eyed that a write acknowledged under Async is not yet guaranteed to be on a replica. Choose Sync and the primary waits for replica acknowledgement before it returns Ok, trading latency for durability.
  • -WriteQuorumPolicy only bites under Sync: All (every replica must ack), Majority (floor(RF/2)+1 must ack), or BestEffort (failures are logged but don't fail the write).
  • -AllowReplicaReads lets clients read from replica nodes to spread read load — with the honest caveat that a replica read may return slightly stale data, because a replica can trail the primary. Leave it off when you need read-your-writes; turn it on when you'd rather have the throughput and can tolerate the lag.

Notice what the defaults say about Zaris's posture: Async + read-from-primary is a fast, correct-by-default cache. The durability knobs are there when you need them, named plainly, each with a cost you can see.

Step 2 — start it, and confirm it came up​

Creating a store defines its shape; it doesn't run the processes. Start-ZrStore enumerates the store's instances on each manager and starts them, reporting success or failure per instance:

Start-ZrStore -Name "orders"

Now the part that separates an operator from someone who ran a command and hoped. Get-ZrStore reads the store's live state back from the managers — overall status, uptime, and the running state of every node:

Get-ZrStore orders
Store   : orders
Status : Running
Uptime : 2h 14m (since 2026-10-03 09:41:12)

Nodes (4)
───────────────────────────────────────────
● 10.0.0.10-n0 10.0.0.10:7861 Running up 2h 14m
● 10.0.0.10-n1 10.0.0.10:7862 Running up 2h 14m
● 10.0.0.11-n0 10.0.0.11:7861 Running up 2h 14m
● 10.0.0.11-n1 10.0.0.11:7862 Running up 2h 14m
───────────────────────────────────────────
Running: 4 Stopped: 0

The store status is a roll-up: Running when every node is up, Stopped when all are down, and Partial when some are up and some aren't — which is exactly the state you want surfaced loudly, because a Partial store is one that booted unevenly. The uptime is deliberately the cluster formation time reported by the core — the earliest node start that survived churn — not the oldest currently-running process, so a store that's been serving for hours doesn't misreport "a few minutes" just because one node restarted.

For the next layer down, Get-ZrNode shows roles and the partition map:

Get-ZrNode orders
Nodes (4)
─────────────────────────────────────────────────────────────────────────
● 10.0.0.10-n0 10.0.0.10:7861 Running Primary v7
● 10.0.0.10-n1 10.0.0.10:7862 Running Replica v7
● 10.0.0.11-n0 10.0.0.11:7861 Running Primary v7
● 10.0.0.11-n1 10.0.0.11:7862 Running Replica v7

The Primary/Replica column is who currently owns each partition, and the v7 is the partition map version — the generation number of the ownership map. When every node agrees on the same map version, ownership has converged; a node stuck on an older version is a node that hasn't caught up to the latest reshuffle. Asking a single node for detail (Get-ZrNode orders 10.0.0.10-n0) expands this into the exact primary and replica partition IDs it holds and when its role last changed — the view you reach for when a failover has just happened and you want to know where ownership landed.

Step 3 — watch it while it runs​

A snapshot tells you the store is up. A live view tells you what it's doing. Watch-ZrStoreMetrics polls the managers and redraws a per-node table on an interval until you cancel it:

Watch-ZrStoreMetrics -StoreName orders -RefreshSec 2

It lays the everyday counters out first — request rate, gets, puts, deletes, item count, memory usage — and pushes the advanced replication and sync diagnostics to the bottom, so the numbers you check at a glance are at the top. Error counters go red the moment their rate climbs above zero; throughput shades green as it rises. It also calls out role changes inline: if a failover flips a primary to a replica while you're watching, a [ROLE CHANGE] line prints with a timestamp — the live signal that a partition just changed hands.

One deliberate detail worth calling out, because it's the kind of thing that makes a tool usable at 2 a.m.: the view fetches the next frame before clearing the screen, and a transient empty poll keeps the last good frame rather than blanking. You get a steady table, not a flickering one, no matter how slow a manager is to answer.

Rounding out the read-only set, Get-ZrInfo reports the installed version, the install and data directories, both module locations, the host runtime — and probes each manager for reachability. A secured manager that answers 401 is reported as reachable but needing authentication, not as "down", so you can tell a network problem from a missing token at a glance.

Step 4 — stop, and the shape underneath it all​

Taking a store down is the lifecycle mirror of starting it:

# Fast whole-store stop (default: a hard kill, since every node goes down together)
Stop-ZrStore -Name orders

# Per-node drain and clean cluster-leave
Stop-ZrStore -Name orders -Graceful

The default is a fast hard kill, which is the right behaviour when the entire store is going down at once — there's no surviving cluster for a departing node to hand off to, so draining buys nothing. Use -Graceful when you want each node to drain and leave the cluster cleanly, which matters when you're stopping a subset rather than the whole thing.

Underneath every cmdlet in this post is one idea: a Zaris store has a desired shape — a canonical definition of the servers and partitions it's supposed to have — and the manager runs under a supervisor model that continuously drives the live processes toward it. New-ZrStore writes that shape; Start-ZrStore realises it; Get-ZrStore and Get-ZrNode report how close reality is to it. The same model is what makes elastic resizing a one-liner rather than a runbook — changing the desired shape with Add-ZrServer / Remove-ZrServer and letting the supervisor reconcile, which is its own elastic scaling story.

The honest summary​

The full day-one lifecycle is five commands and two reads:

New-ZrWorkspace -Name prod -Managers "10.0.0.10:7801","10.0.0.11:7801"
Use-ZrWorkspace prod
New-ZrStore -Name orders -ReplicationFactor 2 -BaseClusterPort 7811 -BaseClientPort 7861
Start-ZrStore -Name orders
Get-ZrStore orders # is it Running, and are all nodes up?
Get-ZrNode orders # who owns what, and did the map converge?
Watch-ZrStoreMetrics -StoreName orders -RefreshSec 2 # what's it doing now?

What makes this a control plane and not just a pile of scripts is that the cmdlets report live, aggregated truth — not what you asked for, but what the cluster actually is — and that the choices with real consequences are named in plain sight: replication factor against machine count, Async versus Sync, replica reads against staleness. Stand a store up, confirm it converged, and keep an eye on it — all without leaving the shell.