Why Local Control Must Keep Water Systems Running During Internet Outages
A water system cannot pause because a mobile tower is down, a fibre link is cut, or a cloud service is unreachable. Pumps still need to start. Tanks still need to hold set levels. Chlorine dosing still needs to stay within approved limits. Alarms still need to trigger before a fault becomes a spill, overflow, or supply failure.
That is why local control matters.
Remote monitoring, cloud dashboards, and central reporting all have a valuable place in modern water operations. They help teams see what is happening across many sites, compare performance, and respond faster. But for the equipment that keeps water moving and treatment processes stable, the first line of control must sit close to the asset itself.
When the internet goes down, the site must keep making safe decisions on its own.

Edge control keeps decisions close to the equipment
Edge control means placing the control logic at or near the site where the physical process happens. In a water system, that usually means a PLC, RTU, industrial controller, or smart telemetry unit connected directly to field instruments and equipment.
A typical edge controller may read signals from:
Level sensors in tanks, reservoirs, wet wells, and clear water storages
Flow meters on inlet, outlet, and dosing lines
Pressure transmitters on rising mains and reticulation zones
Water quality instruments such as turbidity, pH, conductivity, and chlorine analysers
Pump status signals, fault contacts, and motor current readings
Valve position feedback
Local alarm inputs such as high level, low level, intrusion, or power fail
The same controller can also send outputs to:
Start and stop pumps
Open and close valves
Adjust variable speed drives
Control dosing pumps
Activate beacons, sirens, or dial-out alarms
Run local standby sequences
Shift control to a backup asset when a primary asset fails
The key point is simple: the controller does not need the internet to decide what to do next.
A cloud platform may show the operator the current tank level, but the edge controller should decide when to start the pump. A dashboard may trend chlorine residual, but the dosing process should still follow its local control loop if the data link drops. A central alarm platform may send notifications, but the site should still record the alarm and trigger local protection.
This arrangement avoids a dangerous dependency. Communications links are useful, but they are not process control devices. They can fail for reasons outside the water operator’s control, including storms, power outages, carrier issues, damaged cables, network congestion, failed routers, misconfigured firewalls, and planned maintenance.
Edge control accepts that reality and designs around it.
Pumps must continue to run under local rules
Pumps are among the clearest examples of why local control matters. A pump station cannot wait for a cloud command every time a wet well reaches a high level or a storage reservoir falls below its start point.
Local pump control should handle routine actions without outside help. That includes:
Starting duty pumps when a level, pressure, or flow condition calls for it
Stopping pumps at a safe stop setpoint
Alternating duty and standby pumps to reduce uneven wear
Starting a backup pump if the duty pump fails
Locking out a pump when a fault condition is detected
Protecting against dry running, high pressure, overload, or excessive starts
Holding safe states during power loss and restart
In a sewerage pump station, a communications outage should not stop the controller from responding to a rising wet well. If the level reaches the duty start point, the duty pump should run. If that pump trips, the standby pump should run. If the level keeps rising, a high level alarm should activate locally and be stored for later transmission.
In a treated water supply system, a remote tank should continue to call for water based on level, pressure, and programmed limits. If the central monitoring system goes offline, the site should not drain a reservoir because it was waiting for a remote schedule to come back.
Local automation also reduces unnecessary operator intervention. Staff should not need to drive to a pump station during every network outage just to check whether the wet well is being controlled. They may still attend if the outage affects safety or response needs, but the asset should buy them time.
Control at the edge turns a communications failure into a visibility problem, not an immediate operations failure.

Alarms still need to protect the site
During an outage, alarms become more important, not less. Operators may lose live visibility, so the local system must keep detecting, acting on, and storing alarm events.
There are two broad types of alarm behaviour to think about.
The first is local protective action. These alarms should trigger an immediate local response, even if no one can see the alarm in the cloud at that moment.
Examples include:
High wet well level starting an emergency pump sequence
Low suction pressure stopping a pump to prevent damage
High turbidity diverting flow or stopping a process
Chemical tank low level triggering dosing protection
Motor overload locking out a pump
Reservoir high level closing an inlet valve
Intrusion or cabinet door alarms recording site access
The second is operator notification. These alarms need to reach people, but if the normal communication path is down, the system should use the options available.
Depending on the site and risk level, that may include local beacons, sounders, radio paths, SMS gateways, dial-out units, or alternate carriers. Some sites may only be able to store alarms until the link recovers. In those cases, the local controller should still timestamp the alarm, keep the sequence of events, and show the alarm on the local interface.
Alarm design also needs to avoid flooding operators when the link returns. If a network outage lasts several hours, a site may generate repeated status changes. A well-designed system will separate live active alarms from historical events, preserve the timing, and avoid sending a confusing burst of duplicate messages.
The aim is not just to know that something went wrong. The aim is to know what happened first, what the asset did in response, and whether the issue is still active.
Treatment processes cannot depend on a live connection
Treatment systems need stable control because water quality changes can happen faster than a remote operator can react, especially during an outage.
A treatment plant or dosing site may need to keep managing:
Coagulation and flocculation control
Filtration operation and backwash triggers
Chlorine, fluoride, pH, or other chemical dosing
UV system permissives and status monitoring
Clear water tank levels
Sludge handling or waste processes
Shutdown conditions when readings move outside safe limits
For many of these processes, local control is not optional. The controller must read instruments, compare values against setpoints, run interlocks, and adjust outputs in near real time.
A cloud platform can help review performance and compliance trends. It can help supervisors see which sites need attention. It can also support reporting and long-term improvement. But the live treatment sequence should not rely on round trips through the internet.
Take chemical dosing as an example. A local controller can monitor flow and dose proportionally. It can stop dosing if there is no flow. It can alarm if a dosing pump fails, if a chemical tank reaches low level, or if an analyser reading moves outside a safe range. Those actions should continue if the communications link fails at 2 am during heavy rain.
The same principle applies to filtration and backwash sequences. If a filter reaches a programmed headloss or run-time limit, the controller should know whether it is allowed to backwash, which valves to open, which pumps to run, and which interlocks must be satisfied. Losing remote visibility should not leave the filter in an unsafe or undefined state.
Good edge control also defines what happens when the process cannot continue safely. Sometimes the correct local decision is to stop, isolate, or hold a state until an operator attends. The important point is that this behaviour is planned, tested, and handled locally.

Buffered data protects the operational record
When communications fail, data should not vanish. The site should keep recording what happened locally and send it when the connection returns.
This is where buffered data becomes essential.
Buffered data is stored at the edge during an outage. The local controller, RTU, or gateway keeps a record of readings, events, alarms, and equipment states in its own memory. Once communications return, it forwards the stored data to the central system or cloud platform.
A useful buffer should capture enough detail to reconstruct the outage period. That may include:
Time-stamped level, pressure, flow, and water quality readings
Pump starts, stops, run hours, and fault events
Valve commands and position changes
Alarm activations, acknowledgements, and returns to normal
Communications loss and restoration times
Power fail and generator status events
Local operator actions at the site interface
The timestamps matter. If data arrives late but keeps the original event time, operators can still understand the true sequence. Without proper timestamps, the cloud platform may show a misleading cluster of events at the reconnection time.
The storage method matters as well. A short outage may only need a modest buffer. A remote site with unreliable coverage may need enough storage for longer interruptions. For critical assets, the buffer should survive power cycling where practical, so records are not lost during a restart.
Data quality also needs care. The system should identify gaps, mark delayed records clearly, and avoid overwriting important alarm history too quickly. If a sensor failed during the outage, the record should show the sensor fault rather than pretending the value was normal.
Buffered data supports operations, maintenance, compliance, and incident review. It helps answer practical questions after the link comes back.
Did the pumps keep up with inflow?
Did the wet well reach spill level?
Did chlorine residual stay within the target range?
Which pump failed first?
How long was the generator running?
Was the outage a communications issue, a power issue, or both?
Without local buffering, teams may only see a blank space in the trend. In water operations, a blank space is rarely good enough.
Automatic reconnection should be predictable and safe
A communications outage does not end the moment a modem finds signal again. The recovery process needs to be orderly.
Automatic cloud reconnection should happen without an operator manually restarting hardware at site. Once the link is available, the edge device should re-establish its secure connection, confirm the session, and begin sending current status and buffered records.
A good reconnection process will:
Reconnect without interrupting local control
Send current status quickly so operators can see the live state
Upload buffered data in the correct order
Preserve original timestamps
Avoid duplicate alarms and repeated event records
Resume normal reporting intervals after the backlog clears
Flag any missing data or buffer limits reached during the outage
The system should also avoid creating new risk during recovery. For example, a remote command queued before the outage may no longer be appropriate once the link returns. The site conditions may have changed. Pending commands should either expire after a defined time or require fresh confirmation.
This is especially important for commands that start pumps, change dosing mode, reset faults, or move valves. Live remote control should reflect the current site state, not stale instructions from hours earlier.
Reconnection also needs to handle partial recovery. A link may drop in and out before it becomes stable. The edge device should not reboot repeatedly, lose data, or switch control modes back and forth in a way that affects the process. It should keep controlling locally and treat communications as a service that may come and go.
Cloud systems are still valuable when the edge is in charge
Strong local control does not make cloud systems less useful. It makes them safer to use.
Cloud platforms are well suited to tasks that do not need millisecond response at the asset. They can bring many sites into one view, store long-term data, support reporting, and help teams spot patterns across a network.
They are useful for:
Fleet-wide pump performance review
Energy and run-time reporting
Alarm dashboards and escalation
Maintenance planning
Compliance records
Remote configuration where suitable controls are in place
Trend comparison across sites
Long-term water quality review
The best architecture gives each layer the right job.
Control layer | Best role | What happens during an internet outage |
Field devices | Measure and actuate | Sensors and equipment continue to provide local signals |
Edge controller | Run process logic and protection | Pumps, alarms, interlocks, and treatment sequences keep operating |
Communications network | Move data and commands | Link may be unavailable or intermittent |
Cloud platform | Display, store, report, and analyse | Receives buffered data when the connection returns |
This layered approach also helps with cybersecurity and resilience. Critical logic stays inside the local control environment. Remote access can be limited, logged, and managed. If external connectivity is unavailable or intentionally disabled during an incident, the water process still has a defined mode of operation.
For operators, the result is calmer incident handling. A communications outage still needs attention, but it does not automatically become a process emergency.
Designing local control for outage conditions
Local control should not be treated as a backup afterthought. It should be designed, documented, and tested from the beginning.
Key design questions include:
Which assets must keep running without communications?
Which control loops must sit fully at the edge?
What setpoints and interlocks are required locally?
Which alarms need local action, local indication, remote notification, or all three?
How much buffered data is needed for likely outage durations?
What happens if power and communications fail at the same time?
How does the system reconnect and upload stored records?
Which remote commands should expire if they cannot be delivered straight away?
How will operators know whether displayed data is live, delayed, or restored from buffer?
Testing is just as important as design. A practical commissioning test should include disconnecting communications and watching the site continue in local mode. Pumps should start and stop. Alarms should activate. Treatment logic should keep running within its defined limits. Data should buffer. The system should reconnect without manual intervention and upload the missing records.
These tests often reveal small but important issues, such as a cloud-only alarm, a missing timestamp, a setpoint that cannot be changed locally, or a controller that stops logging when the modem fails.
Finding those issues during commissioning is far better than finding them during a storm, heatwave, flood, or major carrier outage.

Reliable water operations start at the edge
Internet-connected water systems can give operators better visibility than older standalone sites ever could. They can reduce travel, improve reporting, and help maintain assets across large service areas. But connectivity should never be the single point holding a water process together.
The safest design assumes that communications will fail at some stage.
When that happens, local control keeps the essentials running. Pumps follow level and pressure rules. Alarms protect the asset and record events. Treatment processes continue under local interlocks. Data is buffered with accurate timestamps. When the connection returns, the system reconnects automatically and fills in the record without disrupting the process.
That is the standard water systems should be built around: cloud-connected when available, locally controlled at all times.
.png)


