The client manufactures EV chargers and sells them to charge point operators, who install them in car parks, forecourts and highway stops across several countries. Once a charger is in the ground it is somebody's revenue, and when it stops working it is somebody's angry customer standing in a car park.
Almost every fault historically needed a technician to physically visit the site. That is expensive everywhere and very expensive in high-income economies, where a single visit can cost more than the charger earns in a month. Multiply by three thousand units across multiple countries and the economics of the whole business rest on one number: how often somebody has to drive out.
The remote management platform exists to drive that number down. Not primarily to detect faults faster — detection was never the hard part — but to fix them from a data centre.
What makes this hard
- The telemetry never stops. Every connector reports diagnostics twice a minute, continuously, from three thousand units. Hundreds of status messages arrive every minute and each one may contain an alarm that matters.
- Most alarms are noise, and the noisy ones look identical to the serious ones at the moment they arrive. A charger that has lost its mobile signal for ninety seconds raises the same kind of event as one with a real fault.
- Commands go out to physical hardware in the field. Getting a remote instruction wrong does not produce a stack trace; it produces an overheating cabinet or a charger that will not restart.
- Field workflows cannot be disrupted. Technicians and operations staff have established ways of working, and a platform that demands they change is a platform they route around.
- Every millisecond of processing is multiplied by the message rate, so anything done per-message — a database lookup, say — becomes an architectural decision rather than an implementation detail.
An alerting system that cries wolf gets muted. Once it is muted, the faults it catches correctly no longer matter.
Five jobs, one source of truth
The remote management system does five things, and the fifth quietly makes the other four possible.
- Ingest the continuous diagnostic stream from every charger and connector.
- Display network status, health and configuration for operations staff and managers.
- Surface alarms in real time so staff can act on them.
- Execute commands remotely — recalibrate, reset, upgrade firmware.
- Hold the master data — the single source of truth for charge point operators, service providers, charging stations, chargers and connectors.
That last one is unglamorous and load-bearing. A rule that applies to "all chargers belonging to this operator" is only trustworthy if the system genuinely knows which chargers those are. Master data quality is the precondition for any automation on top of it.
If the cabinet hits 55°, speed the fans. At 80°, shut it down.
The centrepiece is a rule engine that issues commands to chargers without a human in the loop. An administrator composes a rule from three parts — a trigger (an alarm, or a parameter crossing a threshold), a condition expressed as if-then-else, and one or more actions.
The branching condition is what makes it useful rather than merely automatic, because the right response to a problem depends on how bad the problem is.
This is the difference between an alerting system and a control system. An alerting system would have paged somebody at 55° and again at 80°. The control system resolves the 55° case entirely and escalates only the case that genuinely needs judgement — which is the entire point, because the 55° case is far more common.
Rules are scoped to sets of chargers, and they can also run on a schedule rather than on an event. Scheduled rules handle the maintenance class of problem: check whether firmware is current across the fleet, and upgrade it where it is not — a task that would otherwise be a recurring manual campaign across three thousand units.
Five components, and the interesting one is the suppressor
Both the rule engine and the notification module are built from the same chain.
Triggers
An alarm arrives, a parameter crosses a threshold, or a timer fires for scheduled and persistence-based rules.Generator
Matches the incoming event against the rules in scope and decides what, if anything, should happen.Suppressor
Decides whether this should fire now — once only, or not until the condition has persisted long enough to be real.Actions
Issue a command to the charger, raise a CRM ticket, or send a notification by SMS or email.Why the suppressor carries the design
Two alarms illustrate why this component exists, and they need opposite treatment.
A smoke alarm must notify immediately. There is no acceptable delay, and a false positive is a price worth paying.
A communication failure alarm usually means the charger's mobile connection dropped briefly. It self-rectifies within a few minutes and requires nobody. Notify on it immediately and operations staff receive dozens of alerts a day for a condition that resolves itself — and within a fortnight they stop reading any of them.
So a rule declares which kind it is. Basic rules fire on the event. Advanced rules fire only if the condition persists past a threshold the administrator sets — fifteen minutes for a flaky connection, immediately for smoke. The suppressor also guarantees that one alarm produces one notification, not one per status message for as long as the condition lasts.
It is a small component. It is also the one that determines whether the whole system gets trusted or muted, and it is the part that a naive implementation leaves out entirely.
Two clocks, because one is not enough
Persistence rules cannot be evaluated when an event arrives, because the interesting fact is that nothing has changed for fifteen minutes. So the system runs on two triggers: event-driven for the immediate class, and a once-a-minute sweep for the persistence class, checking live and recently-cleared alarms against the rules that care about duration.
Nothing touches the database on the hot path
With hundreds of status messages a minute, any per-message database read becomes a scaling problem long before the fleet does. The design constraint was explicit from the start: minimise server load per message.
The answer is that rules are cached in memory and every parsed alarm or parameter is matched against the cached set. No database access on the hot path, small processing footprint, and load that scales with rule complexity rather than fleet size.
The question that was asked and answered
The design document poses it directly: should there be a queue to absorb load spikes and decouple notification processing from message ingestion?
The answer was no — because a queue already exists between the MQTT subscriber and the message processor. Adding a second one would have bought nothing but another component to operate and another place for messages to go missing.
That is worth noting less for the conclusion than for the habit. The obvious architectural move was considered, tested against what was already in place, and rejected in writing with a reason. Most systems accumulate queues precisely because nobody asks.
Learning what normal looks like, when nobody can define it
Rules handle the faults somebody already understands. The harder category is the fault nobody has characterised yet — the charger that is not alarming but is behaving unlike its peers.
That problem resists supervised learning, because it has no labels: there is no dataset of correctly-tagged degradations to train against. So the anomaly detection is unsupervised, using a random cut forest to establish what normal operation looks like across the fleet and flag departures from it. The model does not need anyone to define "normal" in advance, which is precisely why it can catch things nobody anticipated.
Separately, a retrieval-augmented pipeline indexes the material support engineers were already using — product manuals, design documents and the accumulated history of how past faults were resolved. When a fault appears, the system can surface how this pattern was handled before, rather than leaving each engineer to rediscover it.
The operational effect is on triage: pattern-based triaging cuts the back-and-forth between a first-line responder and the engineer who has seen the fault before.
Around those two sit three more models, all in production.
Maintenance before the fault
Usage history and diagnostic signatures flag the chargers at risk of overload and the components nearing failure, so a part can be replaced on a scheduled visit rather than an emergency one.Logs, grouped into patterns
Language models read charger logs and fault reports and group them, so a recurring fault appears as one pattern rather than a hundred separate tickets.Insights in plain language
Summaries of how sessions and sites are performing — what is going well, what is not — written for the managers who will never open a dashboard.The constraint worth naming: all of this was added without changing field workflows. The AI sits behind interfaces operations staff already used. A system that requires technicians to adopt new habits does not get adopted.
The rest of the estate
Remote management is the centre of a longer engagement across the client's operations.
| System | What it does |
|---|---|
| Battery swapping | IoT-based automated swap stations — a charged battery exchanged for a depleted one with no wait. RFID and app-based identification, automated payment, operations dashboard, and analytics predicting usage and failures. |
| Installation management | Workflow for commissioning both large commercial DC and residential AC units — survey, installation, commissioning, plus sales orders, asset and inventory tracking and warranty management. |
| Field engineering | Engineer assignment with route optimisation, ticket resolution, and a mobile application for technicians on site — the workflow that handles the visits the rule engine could not prevent. |
| CRM and commerce | Customer relationship management with warranty and reporting modules, enterprise resource planning integration, and an e-commerce marketplace. |