Skip to main content

No Auth Server: Offline-Verifiable Tokens, Roles, and Scopes in Zaris

· 11 min read
Clustron Team
Distributed Systems Engineering

Offline-verifiable tokens, roles, and scopes in Zaris

Most ways of putting auth in front of a distributed data store end up adding a second distributed system to run it: an auth service to call, a session database to replicate, a token-introspection endpoint every node has to reach on the hot path. You secured the store and bought yourself a new availability dependency and a new source of latency.

Zaris takes a different route, and it is a deliberately boring one: one signing key per cluster, and tokens that any node can verify by itself. A token is a JWT signed with the cluster's private key; every manager and every node holds the public half and checks the signature offline — no call to an auth server, no quorum, no database lookup. Security is off by default, and a cluster with it off behaves exactly like an unsecured one. This post is about what turns on when you switch it on: the token's shape, the role-and-scope model, the two planes where enforcement actually happens, and — because pretending otherwise is how people get a false sense of security — the operational edges that are easy to get wrong.

The token is the whole story​

Every Zaris token is a JWT. It names a subject — a user or an application — and carries one or more role @ scope grants. It's signed by the cluster's RS256 private key, which lives only on the managers; everything that needs to check a token holds only the public key.

That single design decision is what removes the auth server from the picture:

  • Any manager can issue a token. There's no primary auth node, no election, no shared mutable session state to coordinate. Issuing is a local signing operation.
  • Every manager and every node can verify a token with just the public key. Verification is a signature check plus an expiry and scope check — pure CPU, no network. The token is self-contained: the role and scope are embedded in it, so a node authorizing a connection needs nothing but the public key. There is no policy database on the node to keep in sync.

The practical payoff shows up on the data path. A client presents its token once, when the connection is established — the "Hello." The node verifies it, authorizes the connection for the store, and then trusts that connection for its lifetime. Tokens are never re-checked on the per-operation hot path, so turning security on costs you nothing per Get or Set. You pay for auth at connect time, once, and never again for that connection.

Roles and scopes​

Permissions are a small preset catalog of roles, each granted at a scope that says where it applies.

RoleScope levelWhat it grants
ClusterAdminclustereverything — all control-plane and data-plane permissions
StoreOperatorstorestart/stop the store, edit config, start/stop nodes, diagnostics, metrics
Observerstorediagnostics and metrics, read-only
DataWriterstoredata read / write / delete / scan
DataReaderstoredata read / scan

A scope is hierarchical, and a broader scope covers everything beneath it:

  • zaris:cluster — the whole cluster, for cluster-wide administration.
  • zaris:store:<name> — a single store, e.g. zaris:store:orders.
  • root (the empty scope) — covers everything; reserved for the bootstrap admin.

The containment is root ⊃ cluster ⊃ store, so a DataWriter granted at zaris:cluster can write to every store, while one granted at zaris:store:orders can write only there. You compose least-privilege out of these two axes: pick the smallest role that does the job, and pin it to the narrowest scope the identity needs.

Turning it on without locking yourself out​

The one thing a security system must never do is strand you outside your own cluster. Zaris handles this by making initialize and enable a single step: enforcement only comes on once the first admin exists.

Connect-ZrManager -Managers mgr-01:7801, mgr-02:7801, mgr-03:7801
Initialize-ZrSecurity -AdminSubject admin -AdminPassword (Read-Host -AsSecureString)
# creates the 'admin' login, generates the shared RS256 key on the first manager,
# copies the SAME key to the rest, turns enforcement on, and signs this session in.

The admin is a real login identity — a username and password, like a cloud root account — not a floating token you have to guard forever. After this, operators sign in with credentials and get a short-lived session token. Enforcement itself is a runtime toggle, persisted per manager, that flips without a restart or a config edit:

Enable-ZrSecurity     # control-plane enforcement ON  (requires an initialized admin)
Disable-ZrSecurity # OFF — keeps the key and policy, so you can re-enable instantly

Both fan out to every manager in the workspace, so the cluster stays consistent.

Two kinds of identity: people and applications​

Following the same pattern as AWS IAM, Kubernetes, and Vault, every token belongs to a registered identity — you cannot mint a token for an arbitrary name. There are exactly two kinds, plus the one bootstrap admin.

Users are people. They're created with a password (stored only as a salted PBKDF2 hash in users.json on each manager — never plaintext, no external database) and they log in to get a session token:

New-ZrUser -Name alice -Password (Read-Host -AsSecureString) `
-Role StoreOperator -Scope zaris:store:orders

Service accounts are applications. No password; creating one hands you a self-contained token the app presents and renews:

New-ZrServiceAccount -Name orders-svc -Role DataWriter `
-Scope zaris:store:orders -LifetimeMinutes 1440
# prints the application token once — store it in the app's secret manager.

New-ZrToken -Subject <name> still exists to issue or rotate a token, but the subject must already be a registered user or service account. There are no ad-hoc tokens for names that don't exist — that's the property that makes "who can touch this store?" an answerable question.

Using a token​

On the data plane, the token folds straight into the connection string — and you never have to bake a literal secret into it, because env: and file: indirection are first-class:

Connect-ZrStore -ConnectionString "zaris://node-a:7861,node-b:7861/orders?token=env:ZARIS_TOKEN"
Set-ZrItem -Key customer:42 -Value '{"name":"Ada"}'

From .NET it's the same string, or a TokenProvider in code when you want rotation:

var services = new ServiceCollection()
.AddClustronZarisConnectionString("orders",
"zaris://node-a:7861,node-b:7861/orders?token=env:ZARIS_TOKEN")
.BuildServiceProvider();

var client = await services
.GetRequiredService<IZarisClientProvider>()
.GetAsync("orders"); // token presented once, at the connection Hello

If the token is missing, expired, invalid, or not scoped to this store, the connection is refused with an Unauthorized status — you find out at connect, not three operations into a transaction. If you'd rather not mint and paste a token by hand, Connect-ZrStore can also exchange a username/password credential for one against the Management Service and renew it in the background.

Renewal that never drops a connection​

Here's where "verified only at connect" pays off twice. Because a node checks the token only when the connection is opened, an in-flight connection is never interrupted when its token expires. Renewal only has to make sure the client has a currently valid token the next time it opens a connection — on a reconnect, a new node, a node restart.

So the client reads its token through a provider on every (re)connect. Keep that provider returning fresh tokens and renewal is completely seamless:

var renewer = new ZarisRenewingTokenProvider(initialToken, "http://mgr-01:7801");
renewer.Start(); // renews itself at 80% of lifetime, before expiry

var options = new ZarisClientOptions {
StoreName = "orders",
Mode = ZarisClientMode.Remote,
TokenProvider = renewer.Token, // read on every (re)connect
};

The renew endpoint (POST /security/token/renew) is stateless — it verifies your current token offline against the shared key and returns a new one with the same subject and roles and a fresh expiry, so any manager can serve it and it survives restarts. An expired token can't be renewed (renew before expiry — the built-in provider does), and a revoked token or subject is refused. The guidance that follows from this is simple: prefer short lifetimes with automatic renewal over one long-lived token. You get the same seamlessness with far less exposure.

If you don't want client code in the loop at all, point the provider at an environment variable or a file and let an external rotator — a secret manager, a sidecar, a cron that re-mints — update it out of band:

options.TokenProvider = ZarisTokenProviders.FromEnvironment("ZARIS_TOKEN");
// or ZarisTokenProviders.FromFile("/etc/zaris/token");

The honest part: two planes means two switches​

This is the detail that bites people, so it gets its own section. "Security is on" is not one fact — it's two, because there are two enforcement points.

The control plane — the Management Service on port 7801, where you create stores, start and stop nodes, and manage grants — enforces as soon as you've initialized and enabled security. That's the switch Enable-ZrSecurity throws.

The data plane — the actual Zaris nodes that serve your keys — is a separate switch. A node only enforces tokens once it has the shared public key in its configuration, and you provision that per store:

Enable-ZrStoreSecurity -Store orders            # write the public key into every node's config
Enable-ZrStoreSecurity -Store orders -Restart # ...and rolling-restart to activate now

Until you do this, the store's nodes run an AllowAll authenticator and accept every connection — tokened or not. This is intentional: it lets you turn on the control plane first and roll data-plane enforcement out store by store, with a planned restart, instead of flipping the whole cluster at once. But it means that enabling the control plane does not, by itself, secure your data. The two matching traps:

  • Get-ZrSecurityStatus reports the control plane. Its Enabled / Initialized / Issuer fields describe management-plane enforcement. A True there does not prove that a given store's nodes are verifying tokens — that's Enable-ZrStoreSecurity's job, confirmed separately. If a "secured" store is still accepting untokened clients, the node config lacks the key or wasn't restarted.
  • Revocation isn't instant on the data path. Because data tokens are self-contained and checked only at connect, Revoke-ZrToken is honoured immediately by the control plane and by renewal (a revoked subject can't refresh), but an already-open data connection keeps working until it closes. The clean way to cut an application off on the data plane is to issue data tokens with a modest -LifetimeMinutes and let the current one expire, or to restart its connections after revoking. If you need near-instant cutoff, short lifetimes are the lever — not an expectation that the hot path re-checks.

Neither of these is a bug; they're the direct consequences of "verify offline, once, with no auth server." They're just the part you have to operate deliberately. The recommended rollout falls right out of them: initialize and enable the control plane, verify with Get-ZrSecurityStatus, issue operator and application tokens and confirm they work, then Enable-ZrStoreSecurity per store with a planned restart, and finally move applications onto short-lived tokens with automatic renewal.

Where it sits in the stack​

Token auth answers who may touch this store and what may they do. It composes with the other two halves of a locked-down Zaris and doesn't replace either:

  • Pluggable TLS trust encrypts the wire and authenticates the endpoints — it's the zariss:// transport under the token, orthogonal to who's holding the token.
  • The secure Helm walkthrough puts both together on Kubernetes end to end, which is where the connection-string and certificate details actually get exercised.

The thing worth keeping is the shape of the core idea. One RS256 key, JWTs that embed their own authority, verification that's pure offline CPU, and a token checked exactly once per connection. That's what buys you authentication and role-scoped authorization across an entire cluster without standing up an auth server, a session store, or a single extra network hop on the path your application actually runs on. The cost is the discipline this post spent its back half on: remember that the data plane is its own switch, and keep your tokens short-lived. Do that, and "secured" becomes a property you can state precisely rather than hope for.