Engineering notesNo. 10Multi-tenancyAll notes

Three ways to isolate a tenant in a NoSQL SaaS platform

For a multi-tenant SaaS platform on a document store, we weighed an account per tenant, a table per tenant and shared tables, and chose shared tables. Sharding tenants across databases and enforcing tenant scope in the data layer bought the isolation back.

Multi-tenancyNoSQLShardingSaaSData isolation

Multi-tenancy is the decision that makes SaaS economics work and the one that, done wrong, ends the company. Same application, same infrastructure, many customers' data — and an absolute requirement that no customer ever sees another's. On the customer success platform we build for a US software company, used by tens of thousands of people across five continents, that data lives in a document store.

There are three broad ways to arrange tenants in one. We evaluated all three explicitly rather than defaulting into one, and that is the part worth reproducing.

Separate accountsOption 1 · strongest

A distinct cloud account per tenant, each with its own namespace, its own footprint and its own billing view.

For: the cleanest separation available, and resource usage per tenant falls out naturally for metering. Against: cumbersome to administer, and impractical once tenants number in the hundreds — every new customer is an account provisioning exercise.

Table per tenantOption 2 · middle

One database, table names prefixed with a tenant identifier.

For: access policies can be applied at table level, metrics captured per table, and throughput scaled per tenant — so one heavy tenant cannot starve a quiet one. Against: the operational burden moves everywhere. Every tool, dashboard and support query needs to understand the naming scheme, and anything metering consumption gains a layer of indirection.

Shared tablesOption 3 · chosen

All tenants in common tables, partitioned by an index whose key is the tenant identifier. What would normally be the primary key becomes a secondary component of it.

For: one unified way to manage and migrate every tenant's data, and tenant-wide analytics become an ordinary query rather than a fan-out across tables. Against: the least granular control over access, performance and scaling; isolation depends on query discipline; and a problem with a shared table is a problem for everybody.

Choosing the least isolated option

On isolation alone the ranking is obvious, and we went the other way. That deserves an explanation rather than an assertion.

The advantages of options 1 and 2 are mostly about operating tenants separately — separate limits, separate metrics, separate access policies. The advantages of option 3 are about operating the platform as one thing — one migration, one schema change, one analytics query, one backup strategy.

For a product that ships frequently and needs to see across its whole customer base, the second set compounds. A schema change against one shared table is an afternoon. The same change against four hundred tenant-prefixed tables is a migration programme with a rollback plan, and it recurs every time the product evolves.

Option 2's isolation is paid for once per tenant. Option 3's simplicity is collected every time the product changes.

Buying the isolation back

Choosing shared tables does not mean accepting a single undifferentiated pool. We recovered the isolation at a different layer: sharding.

We distribute tenants across multiple databases, with all of one tenant's data contained in a single shard. A catalog maps tenants to shards. Within a shard the schema is shared; across shards, tenants are physically separated.

This works because of a property of SaaS access patterns that is easy to overlook: almost every request touches exactly one tenant. A user is logged into one account, looking at one organisation's data. Cross-tenant queries exist, but they are analytics rather than the hot path. So the natural partitioning key of the workload and the natural isolation boundary are the same thing, and sharding costs very little at request time.

Isolation is not one dial. Schema is shared for operational simplicity; physical separation happens at the shard; access control is enforced in a data layer queries cannot bypass. Three mechanisms, three concerns, none doing a job better handled elsewhere.
Recovered

Blast radius

A problem with one shard affects the tenants on it, not all of them. The single-point-of-failure objection to shared tables is answered structurally.
Recovered

Noisy neighbours

A heavy tenant can be moved to its own shard. The lever that option 2 gave per table now exists per shard.
Recovered

Restore granularity

Smaller databases are easier to operate. Restoring one tenant to a point in time means restoring one small shard, not a monolith.

What sharding costs, and what discipline it demands

Two things are genuinely harder, and both are ongoing rather than one-off.

Shard management is a real subsystem. It needs the tenant-to-shard catalog, procedures to add and remove shards, and — the one people underestimate — a way to move a tenant from one shard to another while it is in use. That last operation is not optional: tenants grow, distributions become uneven, and the ability to rebalance is what keeps the scheme working in year three.

Query discipline becomes a correctness requirement. This is the cost of shared tables that no amount of sharding removes. With per-tenant tables, forgetting the tenant scope produces an error — the table does not exist. With shared tables, forgetting it produces results: another customer's data, returned successfully, to a user who should never see it.

This is the same argument as declarative idempotency: when a property must hold everywhere, make it impossible to omit rather than expecting it to be remembered.

The generalisable part

Isolation is not a single dial. It is several independent properties — blast radius, performance containment, access control, restore granularity, data residency — and they can be satisfied at different layers. Treating them as one choice pushes teams toward the most isolated option, which is usually the most expensive to operate and often more isolation than any individual requirement actually demanded.

Decompose the requirement, then ask which layer serves each part most cheaply. Here the answer was: schema shared for operational simplicity, physical separation at the shard for blast radius and noisy neighbours, and enforced query scoping for access control. Three mechanisms, three concerns, none of them doing a job better handled elsewhere.

And the general caution: when you choose a less isolated option, name the mechanism that compensates — and make it structural. Shared tables with disciplined queries is a sound design. Shared tables with hopefully disciplined queries is a breach with a delay on it.

Engineering notes — No. 10 · 2026-10-01