Engineering notesNo. 08Machine learningAll notes

Anomaly detection for EV charger faults nobody has named

We run a random cut forest over telemetry from 3,000 EV chargers to catch the faults no rule describes, by flagging units that behave unlike their peers. Each finding is paired with the record of how similar cases were fixed before.

Anomaly detectionUnsupervised learningIoT telemetryE-mobility & energyApplied AI

The rule engine in the remote management platform we build for an EV charger manufacturer handles faults that are understood. Cabinet over temperature, communications lost, a connector reporting a state it should not be in. Somebody characterised each of those once, wrote a rule, and the fleet has been covered for it ever since.

The category that rules cannot reach is the charger that is not alarming and not right. Nothing has crossed a threshold. Every individual reading is within tolerance. But the unit is drawing slightly more current than it used to for the same session, or its session completion time has crept up, or it fails marginally more often on one connector than the other — and in six weeks it will fail properly, during a busy afternoon, with somebody's car attached.

There is no rule for this, because nobody has written down what it looks like. Nobody has written it down because nobody knows.

Why supervised learning does not help

The instinct is to train a classifier: label historical telemetry with which units subsequently failed, learn the precursor pattern.

That needs labels, and in this domain the labels are bad in several compounding ways:

  • Failures are recorded as tickets, not as telemetry ranges. A ticket says a technician replaced a component on a Tuesday. It does not say when the degradation started, which is the thing to be predicted.
  • The interesting failures are rare. A fleet of three thousand units generates very few instances of any particular degradation mode, which is not much of a training set for any of them.
  • Successful prevention destroys the label. A unit that was recalibrated before it failed produces no failure to learn from — so the better maintenance gets, the worse the training data becomes.
  • New failure modes have no history at all. A hardware revision or firmware release can produce a degradation pattern that has never occurred before, and a supervised model will score it as ordinary.

A model that can only recognise failures already seen and labelled is a rule engine with extra steps.

The fleet is the training set

The unsupervised framing sidesteps all of it. So we asked a different question: not does this look like a failure, but does this look like the others.

We use a random cut forest over the telemetry. The useful properties for this problem are specific:

Property 1

No labels

It learns the shape of ordinary operation from the data itself. Nobody has to define normal, which is fortunate, because nobody could.
Property 2

Incremental

It accommodates a continuous feed from a changing fleet rather than requiring a fixed training window to be curated in advance.
Property 3

Multivariate

It scores the combination. A current draw that is fine and a duration that is fine can be jointly unusual, and that is often what a developing fault looks like.

The structural advantage is that a fleet is a naturally supervised population without anyone supervising it. Three thousand units doing the same job under varying conditions means normal is discoverable empirically — it is whatever the population does. A unit diverging from its peers is interesting regardless of whether anyone has a name for what is happening to it.

This is exactly the case the rule engine cannot serve, and the two are complementary rather than competing: rules encode what is known, the model surfaces what is not.

An anomaly is not a problem

The failure mode of this approach is important enough to design against from the start.

Unsupervised detection produces statistical outliers. Most of them are not faults. A charger in an unusually cold location is an outlier. A unit installed last week is an outlier. A site with genuinely atypical usage is a permanent outlier, correctly identified and entirely uninteresting.

Shipped raw to an operations team, that is a very sophisticated way of producing noise — and noise has a predictable consequence, which is that within a fortnight nobody looks at it. The model then joins the growing category of systems that are technically working and operationally ignored.

So anomaly detection needs the same discipline as alerting: suppression, persistence thresholds, and a bias toward the recurring over the momentary. A one-off outlier is weather. The same unit diverging from its peers for a week is a finding. Only the second deserves a human, and building that distinction is not optional decoration on top of the model — it is the part that determines whether the model is used.

Closing the loop

An anomaly score on its own tells an engineer that a unit is unusual, which is a poor handover. The score becomes actionable when it is joined to the accumulated record of how similar situations were resolved before — manuals, product documentation, and the history of past fixes, which we index so they can be retrieved alongside it.

The combination is what changes the triage conversation: not this unit is anomalous, but this unit is anomalous in a way that resembles a pattern resolved four times previously, by adjusting a parameter. The first requires an expert to interpret it. The second does not, which is the entire point, because the expert is the scarce resource.

The generalisable part

If you operate a population of similar things, peer comparison is a supervision signal you already have. Fleets of devices, instances of a service, stores in a chain, branches, machines on a line. You do not need labelled failures to know that one member is behaving unlike the rest, and in most of these domains labelled failures are exactly what you lack.

Two conditions for it to work. The population has to be genuinely comparable, or you will spend your time rediscovering that the cold site is cold — segment first, then compare within segment. And you must build the triage layer before you ship the scores, because an unsupervised detector without suppression is a noise generator, and the cost of that is not a wasted model but a team that has learned to ignore one more channel.

Engineering notes — No. 08 · 2026-10-01