Skip to main content

The Same Code, With or Without a Network: Zaris Embedded Mode

· 9 min read
Clustron Team
Distributed Systems Engineering

Zaris embedded in-process mode

Most distributed data stores make you choose, early and permanently, between two worlds. There is the "real" store — a cluster you deploy, point a client at over the network, and talk to through a wire protocol. And there is the "fake" store you reach for in unit tests: an in-memory dictionary, a mock, a hand-rolled stub that implements just enough of the interface to make the test compile. The two share nothing. Your test exercises the stub's behaviour, not the store's, and the gap between them is exactly where bugs hide.

Zaris collapses that choice. The same typed .NET client that talks to a distributed cluster over TLS also runs the store embedded inside your process, with no socket, no partition map, and no deployment — and it is the same store engine, not a stand-in. You select which one you get with a single connection string. This post is about how that works, the cases it's genuinely good for, and the hard line where embedded mode stops and a cluster begins.

One string is the whole switch​

A Zaris client is always configured from a connection string. The string carries the transport, the seed endpoints, the store name, and options. There are two shapes, and the difference between them is the difference between a cluster and an embedded store:

zaris://inproc/orders                     # embedded — store runs inside this process
zariss://zaris.mycorp.com:7863/orders # remote — connect to a cluster over TLS

The host token inproc is reserved. When the client sees it, there are no endpoints to dial, no transport to negotiate, no TLS handshake, and no token — the store is served from memory in the current process. Everything else about your code stays identical:

// Embedded: no server to start, no port, no network.
IZarisClient store = await ZarisClient.ConnectAsync("zaris://inproc/orders");

// Later, in production — the ONLY line that changed:
IZarisClient store = await ZarisClient.ConnectAsync("zariss://zaris.mycorp.com:7863/orders?ca=/etc/zaris/ca.crt");

In ASP.NET Core it is the same story through dependency injection — the connection string usually comes from configuration, so the switch is a config value, not a code change:

// Program.cs — one registration, driven entirely by the connection string.
builder.Services.AddClustronZaris("orders", builder.Configuration.GetConnectionString("orders")!);

Put zaris://inproc/orders in appsettings.Development.json and zariss://.../orders in production, and the same binary runs embedded on a developer's laptop and clustered in the cloud, with nothing recompiled.

Why "same client" is the point​

The reason this matters is in the type you get back. ConnectAsync returns an IZarisClient — or, for the full feature surface, IZaris — regardless of mode. Embedded mode is served by a class called ZarisInProcClient, and it implements the exact same interfaces as the cluster client:

public interface IZaris : IZarisClient
{
ILeasesClient Leases { get; }
ILocksClient Locks { get; }
ICountersClient Counters { get; }
IScanClient Scan { get; }
IWatchClient Watch { get; }
IPubSubClient PubSub { get; }
IHashClient Hashes { get; }
IListClient Lists { get; }
ISetClient Sets { get; }
ISortedSetClient SortedSets { get; }
IStreamClient Streams { get; }
}

Every one of those surfaces is present in embedded mode, backed by the same store engine the cluster uses. Key/value Get/Put/Delete, bulk and batch, multi-item transactions, TTL and expiry, optimistic compare-and-swap, Watch for reactive reads, native Streams with consumer groups, the replicated data structures — all of it resolves against the embedded store, not a mock.

That shared engine is what makes the embedded store a faithful test target. A compare-and-swap loop that relies on a conflict being reported precisely, a Watch subscription that must fire when a key changes, a transaction that must commit atomically or not at all — each behaves embedded the way it behaves clustered, because it is literally the same code path evaluating it. A KvStatus.Conflict embedded means what it means in production. A hand-written dictionary stub can't promise that.

What the embedded host actually is​

It's worth being precise about what zaris://inproc spins up, because the honest limits fall straight out of it. The embedded host is a fixed single-node topology that owns every segment. There is no partition map, because there is only one participant; there is no replication, because there is no second copy; there is no transport, because there is no second process to reach.

What does run is the real machinery you'd want in a test. The store provisions its segment up front, wires in the lease service, the watch service, and local metrics, and starts the TTL sweeper on its normal cadence — so a key you ExpireAsync with a 50 ms TTL actually disappears, on a timer, exactly as it would in a cluster. Expiry, eviction accounting, versioning, and watch notifications are all live. It is a complete store that happens to have a cluster size of one.

Within a single process, clients are shared by store name. Two ConnectAsync("zaris://inproc/orders") calls in the same process resolve to the same embedded store, so a value one writes, the other reads — handy when your test spins up a producer and a consumer and expects them to see each other. Different store names get different stores; different processes get entirely separate embedded stores (more on that limit below).

Where embedded mode earns its keep​

Unit and integration tests that exercise the real store. This is the headline use. Your repository, your cache layer, your saga coordinator — test them against zaris://inproc/test-{Guid}, get a fresh isolated store per test, and exercise the genuine semantics: real versioning for your CAS loops, real TTL expiry on a real timer, real transaction atomicity, real watch events. No Docker container to start in CI, no port to allocate, no teardown flakiness. The store is created on first connect and dies with the process.

Single-process applications that may grow up later. A desktop app, a background worker, a CLI tool, or a small service that needs a fast structured store today but might become a clustered deployment tomorrow. Start embedded. The day you need to share that data across instances, you change the connection string to point at a cluster — and because you coded against IZaris the whole time, nothing else moves.

Integrations that ship embedded by default. The Zaris ASP.NET Core session provider is a concrete example: wiring Zaris as your session store is one call, and in its sample it runs entirely embedded —

builder.Services.AddZarisSession(options =>
{
options.ConnectionString = "zaris://inproc/sessions";
});

— which means a developer can F5 the app and have working sessions with zero infrastructure, then move to zariss://.../sessions for a multi-instance deployment without touching the session code.

To use embedded mode, reference the Clustron.Zaris.InProc package alongside the client (it targets .NET 8). The package self-registers: the client discovers and loads the in-process factory on first use, so simply referencing it is enough. In a test assembly where you'd rather be explicit, new InProcBootstrap().Register(); wires it deterministically.

Where it honestly stops​

Embedded mode is a real store, but it is a single-node, in-memory store, and pretending otherwise would be the marketing fluff this blog tries to avoid. Three hard limits, straight from the design:

No durability. The embedded store holds data in process memory and nowhere else. There is no replication and no on-disk log in this mode — when the process exits, the data is gone. That is exactly what you want for a test (clean slate every run) and exactly what you must not rely on for anything you need to survive a restart. Embedded mode is not a database you back your production data with.

No high availability, no scale-out. A cluster of one tolerates the loss of zero nodes and scales to the memory and CPU of one process. All the reasons you'd reach for distributed Zaris — surviving a node failure without data loss, sizing replicas for durability, scaling throughput across cores and machines — simply don't apply to an embedded host, because there's only one of it. When you need those properties, you need the cluster.

No cross-process sharing. The in-process registry shares a store by name within one process. Two different processes that both open zaris://inproc/orders get two independent stores that know nothing about each other — there is no socket between them, by definition. The moment your data needs to be visible to a second process or a second machine, embedded mode is the wrong tool and a cluster is the right one.

None of these are bugs to be fixed; they're the definition of "embedded." The value isn't that the in-process store is secretly distributed — it's that the day you need distribution, the only thing you change is the string that says where the store lives.

The same promise, kept twice​

The pitch is small and concrete. You write your application against one client surface. In a test, or on a laptop, or inside a single-process tool, that surface is served by a real Zaris store running in your own memory — fast to create, faithful in its semantics, and gone when the process ends. In production, the identical code talks to a distributed cluster over TLS. The behaviour your tests pinned down is the behaviour you ship, because it was never a different implementation — only a different place for the store to live.

That's the whole idea behind zaris://inproc: the gap between "the store in my test" and "the store in production" is one connection string wide, and you get to decide, per environment, which side of it you're on.