Knowledge Center

Network Monitoring: The Signals That Give Warning

Sarah Jane Sep 08, 2026 5 min read

Network monitoring is usually implemented as a dashboard of green lights, which tells you that things are up. What you actually want is warning that something is degrading, and that comes from a small number of specific signals.

Up or down is the least useful thing to watch

A device that is down is already causing a problem someone has noticed.

The value is in the signals that change before that: errors accumulating on a port, a link running at the wrong speed, a power supply that has failed leaving no redundancy, or optical light levels drifting.

That is the same principle covered in our hardware monitoring guide — monitor the things that degrade gradually, because those are where warning is possible.

Port errors are the earliest signal

The most useful thing a switch reports and the most commonly ignored.

Interfaces count errors: malformed frames, discards, collisions and similar. A healthy port accumulates essentially none. A port whose error count climbs steadily has a physical problem — a marginal cable, a bad termination, a dirty fibre connector, or a duplex mismatch.

The important property is that errors appear long before the link fails. Traffic still passes, applications still work, and performance is quietly worse because of retransmissions.

What to watch: error counters increasing over time rather than their absolute value, since counters accumulate since last reset. A rate of change is meaningful; a total is not.

On copper, the usual causes are covered in our cabling guide — untwisted pairs at termination, bend radius, and proximity to power. On fibre, contamination is the leading cause, as our fibre guide covers.

Optical light levels

The most diagnostic figure available on fibre links, and most switches report it per port.

Transmit power low or absent points at the local transceiver.

Good transmit, low receive at the far end points at the path between — contamination, a bend, a damaged patch lead, or distance beyond what the module supports.

Receive power fluctuating suggests an intermittent physical problem, frequently a marginal connection or a cable under stress.

Trending these rather than reading them once is what turns them into warning. A link whose receive power has drifted down over months is telling you something before it stops.

Speed and duplex mismatches

A specific fault worth calling out because it presents as a performance problem rather than an error.

A link negotiating to a lower speed than expected, or a duplex mismatch, produces poor throughput and high error counts while the interface reports itself as up.

Users report slowness; monitoring reports green. Checking negotiated speed against expected speed across your ports finds these quickly, and it is worth doing after any cabling work.

Also worth remembering that a Wi-Fi 6 device on older access points behaves like an older device — both ends must support a standard for it to apply, as our coverage guide covers.

Environmental and power on network equipment

Switches report more about themselves than people use.

Temperature. Network equipment in a branch cabinet or an unventilated cupboard runs hot, and heat shortens life. Alert on it.

Power supply and fan status. On modular equipment, a failed supply means redundancy is gone — a state that should be treated as urgent rather than logged and forgotten.

PoE budget consumed. Budget is shared across the switch rather than per port, so it can run out before ports do. Watching consumption against budget prevents the surprise when the next access point does not power up. Our switch guide covers this.

Configuration change is worth watching too

Not strictly monitoring, and it belongs in the same conversation.

A significant share of network incidents follow a change rather than a failure. Knowing that a configuration changed, and being able to compare it against what it was, turns an investigation into a comparison.

Two practical points. Keep configuration backups, and keep them somewhere that survives the loss of the device — our documentation guide covers why storage location decides whether documentation helps.

And have out-of-band access, because the changes that cause the worst incidents are the ones that remove your ability to reach the device.

The alert path decides whether any of it matters

The same conclusion as hardware monitoring generally.

Alerts must reach a person rather than an unread mailbox, must work out of hours, and must reach someone who can act. A dashboard nobody opens at the weekend is not protection.

And guard against alert fatigue: alert on things requiring action, set thresholds against your actual lead times rather than defaults, and fix or suppress recurring alerts nobody acts on — because an alert nobody acts on trains everyone to ignore alerts generally.

Test the path by triggering something deliberately. An untested alert path is an assumption, exactly like an untested backup.

Starting from nothing

You probably already own most of the capability. Managed switches report port errors, negotiated speed, optical levels, temperature, power state and PoE consumption without buying anything.

Start there: collect port error counters and trend them, alert on power and temperature, check negotiated speeds against expected, and route it all somewhere a person reads. The gap is usually configuration and the alert path rather than sensors.

Common questions

What should I monitor beyond up and down?

Port error counters trending upward, negotiated versus expected speed, optical light levels on fibre, temperature, power supply and fan state, and PoE budget consumed. A device that is down is already a problem someone noticed.

Why do port errors matter if traffic still passes?

Because they appear long before a link fails. Applications still work while performance is quietly worse from retransmissions, and a climbing error count indicates a physical problem — a marginal cable, bad termination, dirty fibre connector or duplex mismatch.

Should I watch error totals or the rate?

The rate of change. Counters accumulate since last reset, so a total tells you little. A counter climbing steadily over time is the meaningful signal.

Why do users report slowness when monitoring is green?

Frequently a speed or duplex mismatch. The interface reports itself up while throughput is poor and errors accumulate. Check negotiated speed against expected across your ports, particularly after cabling work.

Do I need to buy a monitoring system?

Not to start. Managed switches already report port errors, negotiated speed, optical levels, temperature, power state and PoE consumption. The gap is usually configuration and the alert path rather than sensors.

Most of this is already reported by your switches — the work is collecting it and routing alerts somewhere a person reads.

Sarah Jane

Sarah Jane

Senior IT Hardware Specialist · TechSellerUSA
Sarah helps businesses and IT teams source the right enterprise hardware at wholesale prices. View profile →