Case studySaaS & customer successAll work

Two frameworks, one dashboard

A customer success platform competing with Salesforce, serving tens of thousands of users across five continents — and the architectural decisions that let a small team keep pace with a company a thousand times its size.

Client A US-headquartered customer success SaaS companyEngagement Seven engineers at launch · 2018 — presentPractice SaaS & customer success

The client sells a customer success platform to subscription businesses: software that tells an account team which customers are thriving, which are quietly disengaging, and which are about to leave. It competes directly with Salesforce and the established customer success vendors, and it is used from the Philippines and Singapore through India and Europe to the United States.

The product is fundamentally an analytics surface over messy data. It ingests event telemetry from the customer's own product, financial records, support tickets and CRM activity, then renders health scores, heat maps and trend dashboards that a senior manager can read in ten seconds. Customers integrate their own finance, HR and operations systems alongside third-party tools, and the dashboard reshapes itself around whatever is connected.

That combination — real-time analytics, deep third-party integration, global deployment, enterprise security expectations — is normally the output of a large engineering organisation. The client is not one.

What makes a platform harder than an application

  • Tenants share infrastructure but must never share data. Multi-tenancy is what makes the unit economics work and it is also the thing that, done wrong, ends the company.
  • Customers want the platform in their own cloud. Some enterprise buyers will not accept shared hosting, so the same build has to deploy privately without forking.
  • Usage has to be measured to be billed. Every API call through gateway and compute has to be attributed to a tenant, in a way that survives being queried a year later.
  • Third parties consume it directly. Mobile and web SDKs mean an external developer is an audience, with everything that implies — documentation, versioning, a sandbox, and the knowledge that a published interface cannot casually change.
  • Compliance varies by region, and so does acceptable latency, so geography becomes an architectural concern rather than a deployment detail.

An application has users. A platform has tenants, developers, regions, meters and versions — and each one is a design problem before it is a feature.

Documented in five views, not one diagram

The platform's architecture is written up against the 4+1 view model — the discipline of describing a system from several stakeholder perspectives rather than one, so that the picture a developer needs and the picture an operator needs are both present and consistent.

view 1

Logical

What the system does — functional decomposition into services and the responsibilities each holds.
view 2

Development

How the code is organised — layers, module boundaries, and the structures a developer works inside daily.
view 3

Process

Runtime behaviour — tenancy, geography, metering, throttling, upgrades, availability.
view 4

Physical

Where it runs — deployment topology for server, mobile, analytics and the build pipeline.
+1

Scenarios

The use cases that tie the other four together and validate that they agree.

This matters more than it sounds for a small team. Written views are what let one engineer make a change in an area another engineer owns without a meeting, and they are why an eight-year engagement has not accumulated the usual archaeology.

Halfway through, the front-end framework was wrong

The platform's dashboards were originally built in AngularJS. As the product matured it became clear that React's component model was a much better fit for what the UI actually was: a grid of independent widgets that switch on and off depending on which of the customer's systems are connected, each caching, refreshing and invalidating its own data.

That is the moment every long-lived product reaches, and it usually has two bad answers. Rewrite everything — months of work producing no new customer value, with a feature freeze the business will not accept. Or run both frameworks side by side in one page and accept a permanent mess.

The third answer

Micro front-ends. Each widget became an independently deployable front-end unit, connected to its own microservice on the back end. The Angular dashboards kept working while new widgets arrived in React, and the migration proceeded widget by widget on the product team's schedule rather than as a single event.

The structural insight is that widget-based micro front-ends give the front end what microservices give the back end — independent deployability, contained failure, and the ability for different parts of one screen to evolve at different speeds. For a dashboard product whose whole premise is configurable widgets, the architecture ends up mirroring the domain, which is usually the sign that it is right.

The migration was possible because the seams already matched the product — widgets were independently toggled before they were independently deployed. Without that correspondence, micro front-ends invent boundaries the domain does not have.

It finished

The migration is complete. The platform is entirely React, the old framework is gone, and at no point was there a feature freeze. The transition happened inside ordinary product work — a widget replaced here when it was being changed anyway, another there when a new integration needed it — until there was nothing left in the old framework to replace.

That is the part worth sitting with, because it is the part that distinguishes this from a rewrite that happened to go well. The migration was never a project. It had no plan with an end date, no dedicated team, no quarter in which the roadmap paused. It was a property of the architecture: once the seams existed, replacing a unit was cheaper than not replacing it, and the work happened as a side effect of building the product.

Three ways to separate tenants, and why the least isolated one won

Tenant isolation in a document store is a genuine fork in the road. All three options were evaluated explicitly rather than defaulted into — and the one chosen is not the one an isolation ranking would pick.

Separate accountsStrongest isolation

A distinct cloud account per tenant, each with its own namespace, footprint and billing view. The cleanest separation available, and resource usage per tenant falls out naturally for metering — but cumbersome to administer, and impractical once tenants number in the hundreds.

Per-tenant tablesThe middle option

One database, table names prefixed by tenant. Access policies apply at table level, metrics are captured per table, and throughput scales per tenant. The cost is that the operational burden spreads everywhere: every tool, dashboard and support query has to understand the naming scheme.

Shared tablesChosen

All tenants in common tables, partitioned by an index keyed on the tenant identifier. One unified way to manage and migrate every tenant's data, and tenant-wide analytics become an ordinary query rather than a fan-out. Against it: the least granular control, and isolation that depends on query discipline.

Choosing the least isolated option deserves an explanation. The first two are mostly about operating tenants separately — separate limits, metrics, policies. The third is about operating the platform as one thing — one migration, one schema change, one analytics query, one backup strategy.

For a product that ships frequently, the second set compounds. A schema change against one shared table is an afternoon. The same change across several hundred tenant-prefixed tables is a migration programme with a rollback plan — and it recurs every time the product evolves.

Per-tenant isolation is paid for once per tenant. Shared-schema simplicity is collected every time the product changes.

Buying the isolation back with shards

Shared tables does not mean one undifferentiated pool. The isolation is recovered at a different layer: tenants are distributed across shards, with all of one tenant's data held in a single shard, and a catalog mapping tenants to shards.

This works because of a property of the workload that is easy to overlook — almost every request touches exactly one tenant. A user is signed into one account looking at one organisation's data. Cross-tenant queries exist, but they are analytics rather than the hot path.

Isolation is not one dial. Schema is shared for operational simplicity; physical separation happens at the shard; access control is enforced in a data layer queries cannot bypass. Three mechanisms, three concerns, none doing a job better handled elsewhere.
Recovered

Blast radius

A problem with one shard affects the tenants on it, not all of them.
Recovered

Noisy neighbours

A heavy tenant can be moved to a shard of its own.
Recovered

Restore granularity

Recovering a single tenant means restoring one small database rather than a monolith.

What it costs is a real subsystem: the tenant-to-shard catalog, procedures to add and remove shards, and — the part most often underestimated — the ability to move a tenant between shards while it is in use.

Geography as an architectural concern

Running across regions was driven by three pressures, and it is worth separating them because they pull differently. Failover — a region can fail and load must move. Latency — non-static data served from far away feels broken regardless of how fast the query was. Compliance — some jurisdictions require that data physically remain within them.

Only the third is non-negotiable, but it is the one that forces onboarding, identity and operations to become region-aware rather than global. That is the expensive part, and it has to be designed in early or retrofitted painfully.

Metering, throttling and a sandbox

The unglamorous machinery that separates a platform from an application, and that almost nobody budgets for.

Metering attributes every API call to a user and tenant, centrally, so usage can be billed and usage patterns analysed. The design question is not whether to log but where: a purpose-built usage table gives fast, structured, per-tenant queries at the cost of write volume, while dumping to log storage and querying it later is cheaper to write and slower to answer. The platform took the first route, because usage questions are asked constantly and need to be cheap.

Throttling exists so that no tenant, and no experimenting developer, can degrade production for everyone else.

The sandbox is the piece most SaaS products skip, and it is a distribution decision rather than a technical one. Developers evaluate a platform by trying it, and will not complete a procurement conversation to do so. So the sandbox offers the full feature surface, implemented as an ordinary account with throttles, and is barred from production applications — full fidelity for evaluation, no path to becoming a free tier by accident.

AI inside the workflow, not beside it

The recent work is the AI layer, and its design principle is that it lives inside the account manager's existing workflow rather than as a separate assistant they have to remember to open.

  • Call transcripts summarised automatically, with action items extracted — turning a recording nobody rewatches into a record the next person can act on.
  • Sentiment analysis across conversations, surfacing the calls that went badly rather than leaving them buried among the ones that went fine.
  • Churn risk from unsupervised clustering — finding the shapes of accounts that leave, rather than scoring against a hand-written rule about login frequency.
  • Retrieval-augmented generation over the account's own history, so the assistant answers from this customer's record rather than from general knowledge.
  • Engagement recommendations delivered where the account manager already works, inside the CRM.

The clustering choice deserves a note. Churn prediction is usually framed as supervised classification, which requires a clean labelled history of who churned and why — something most customer success data does not have, because the reasons are recorded inconsistently or not at all. Clustering sidesteps the labelling problem: it finds the account shapes that precede departure without anyone having to define them first, which is both more honest about the data and more likely to surface a pattern nobody had thought to look for.

Client described rather than named, except where cleared.