Engineering notesNo. 02ObservabilityAll notes

Alerting on 3,000 EV chargers without alert fatigue

Sent one by one, alarms from 3,000 EV chargers teach operators to filter them out, smoke alarms included. The platform we build sends each alarm once, holds self-clearing ones such as dropped connections for fifteen minutes, and sweeps every minute for conditions no event reports.

AlertingIoT telemetryEvent vs temporalRule engines

We build the remote management platform for an EV charger manufacturer: three thousand chargers across several countries, each connector reporting diagnostics twice a minute. Hundreds of status messages arrive every minute, any of which may carry an alarm. Operations staff cannot watch a dashboard continuously, so they want to be told when something needs them.

The naive design is three lines long and everybody builds it first. An alarm arrives; look up which users care about that alarm on that charger; email them. It is correct, it is simple, and within a fortnight it has destroyed itself.

Two alarms that need opposite treatment

Here are two real alarms from the same fleet. A notification system that treats them identically is broken, and treating them identically is exactly what the naive design does.

Alarm type A

Smoke detected

There may be a fire inside a cabinet full of power electronics, in a car park, possibly next to a vehicle.

Notify instantly. A false positive is cheap. A delay is not.

Alarm type B

Communications failure

The charger's mobile connection dropped. This happens constantly, for ordinary reasons, and it usually comes back by itself within a few minutes.

Notify, and somebody's attention is wasted.

At fleet scale, type B fires dozens of times a day. Send each one and an operator's inbox fills with alerts about a condition that resolves itself before they have finished reading. They are not stupid, so they do the rational thing: they write a filter rule.

And now the smoke alarm does not reach anyone either, because it arrives in the folder nobody opens.

This is the failure mode. Not a missed detection — the system detected it perfectly. The alert was generated, delivered, and ignored, because the channel had already been poisoned by correct-but-useless notifications. An alerting system's real output is not alerts. It is attention, and attention is finite and easily spent.

The component nobody builds

The naive design asks one question: did an alarm fire? The useful design asks a second one: does this alarm, right now, warrant a human?

Those are different questions, and the second needs state that the first does not. So we put a distinct component between deciding an alert applies and sending it.

01

Trigger

An alarm arrives from a charger, or a timer fires.
02

Generator

Matches it against rules and decides an alert applies.
03

Suppressor

Decides whether it should be sent now — or once, or at all.
04

Notifier

Sends on the channel the rule specifies.

The suppressor enforces two guarantees that between them account for almost all alert fatigue:

One alarm, one notification. The charger reports twice a minute and keeps reporting the alarm for as long as the condition holds. Without suppression, a fault lasting an hour produces a hundred and twenty identical emails. This is the single most common way monitoring systems self-destruct, and it is embarrassing rather than difficult — it just has to be somebody's explicit job.

Persistence thresholds. A rule declares whether it fires on the event or only if the condition survives a defined period. Smoke fires immediately. A comms failure fires only if it is still there after fifteen minutes — by which point it is no longer a flaky cell tower, it is a charger that has genuinely dropped off the network.

Fifteen minutes is not a preference. It is a fact about cellular networks, and it belongs in the rule as domain knowledge set by someone who knows the fleet — not as a slider in a user's notification settings, where it would be guessed.

The second clock

Here is the part that is easy to get wrong, and it is the reason persistence thresholds are harder than they look.

A persistence rule cannot be evaluated when an event arrives, because the fact that matters is that nothing has happened for fifteen minutes. Non-events do not arrive. There is no message that says "the comms failure you saw earlier is still going."

An event-driven system, by construction, cannot detect the absence of change. So we added a second trigger that is not an event at all: a sweep that runs every minute, pulls the rules that care about duration, and checks them against current alarm state.

Why the sweep also has to look backwards
12:00:00Sweep runsNo alarm present on this charger. Nothing to do.
12:00:18Alarm raisedBetween ticks. The sweep is not watching.
12:00:41Alarm clearsAlso between ticks. Lifetime: 23 seconds.
12:01:00Sweep runsNo alarm present. If it queries only live alarms, this is indistinguishable from the previous tick — the event is invisible.

So the sweep queries live and recently-cleared alarms, not just live ones. Sampling a continuous system at discrete intervals loses whatever happens between samples, unless the sampler deliberately looks back over the interval it just skipped.

For a flapping charger — one raising and clearing repeatedly — this is the difference between a system that sees a pattern and one that sees nothing at all. The failure is silent, which is what makes it worth designing for rather than discovering.

The dashed arrow going nowhere is the mechanism. A monitoring system's value depends on what it declines to send, and the component that declines needs somewhere to live — otherwise the decision ends up scattered across the senders as conditionals nobody can audit.

What it costs

Almost nothing, built in from the start. The suppressor is a small amount of state and two guarantees. The second trigger is a scheduled job.

Retrofitting it is a different matter, because by then the notification logic is distributed across the handlers that send notifications, each with its own accumulated conditionals about when not to. That is the usual shape: alert suppression is rarely absent from mature systems, it is just scattered — a dozen local hacks that collectively implement a policy nobody has written down and nobody can change safely.

Making it one named component with one responsibility is the whole trick. It also makes the policy inspectable, which matters the first time somebody asks why they were not told about something.

The generalisable part

Two things transfer beyond charger fleets.

Design the suppression layer before the notification layer. Any system that interrupts a human needs an explicit answer to "should this one actually go out?", and that answer needs somewhere to live. If it does not have a component, it will end up as conditionals inside the sender, and it will be wrong in ways nobody can audit.

If you ever need to act on something not happening, you need a clock as well as an event stream. Timeouts, stalled workflows, absent heartbeats, conditions that persist — none of these generate events, and all of them are things systems are routinely asked to notice. Teams discover this late, usually after shipping an event-driven design and then finding that a requirement phrased as "alert me if it is still broken in fifteen minutes" has no event to hang on.

Then make the sweep look backwards over its own interval, because a discrete sampler that only reads the present will quietly miss everything that resolved between ticks.

Engineering notes — No. 02 · 2026-10-01