How to Tell If a Hard Drive Is Failing: The Four Values
Every drive fails eventually, and the useful question is whether yours is failing now or just old. This covers the indicators that actually predict failure and the ones that do not.
The four values that predict drive failure
Drives report health data continuously. Four values carry most of the signal.
Reallocated sector count. Sectors the drive has retired because they failed. Zero is healthy. A small stable number is usable. A number that is growing is a drive actively degrading, and that is the strongest single predictor.
Pending sector count. Sectors the drive has flagged as suspect but not yet retired. Any non-zero value deserves attention, because these become reallocated sectors or read errors.
Uncorrectable error count. Reads the drive could not recover. Should be zero on an enterprise drive.
Command timeouts. Operations the drive failed to complete in time. Rising timeouts on an otherwise healthy-looking drive frequently indicate a cable or backplane problem rather than the drive.
Our drive health guide covers reading these on a server.
What does not predict failure
Three things people watch that carry less signal than they think.
Power-on hours alone. A drive at 30,000 hours with clean health values is in better condition than one at 8,000 hours with growing reallocated sectors. Hours measure time powered, not work done or wear incurred. Our hours guide covers judging the figure properly.
Temperature, within normal range. Sustained high temperature shortens life, but a drive running warm in a well-cooled chassis is not a warning.
Noise. Drives make noise. A sudden change in noise is worth investigating; steady operating noise is not a symptom.
The pattern that matters most
Failures cluster at two points: early in life, from manufacturing defects, and late in life, from wear. Between them is a long, relatively flat period.
Two consequences that change how you plan.
A drive that has survived its first months is in the safest part of its life, which is why a tested used drive with moderate hours is a reasonable buy. Our burn-in guide covers why testing catches the early cluster.
Drives installed together fail together. Same batch, same age, same workload, same environment. When one drive in an array reaches end of life, the others are close behind. That is the single most important thing to understand about array planning.
Why that matters during a rebuild
The dangerous moment is not the first failure. It is the rebuild.
A rebuild puts sustained heavy read load on every remaining drive in the array, at exactly the point when those drives are the same age as the one that just failed. Large drives make this worse because rebuilds take longer.
Three responses.
RAID 6 or equivalent double parity at large capacities, so a second failure during rebuild is survivable. Our RAID guide covers the trade-offs.
Hot spares, so the rebuild starts immediately rather than when someone notices.
And a backup, because RAID is not backup. Our backup guide covers the principle.
Planning replacement rather than reacting
Four habits.
Monitor the four values, and alert on reallocated and pending sectors changing rather than on absolute thresholds. Our monitoring guide covers routing alerts somewhere a person reads.
Record drive hours at installation, so you know the age spread across each array.
Hold spares for arrays that matter, in the correct part number, sector format and class. Our spares guide covers this.
Consider staged replacement on ageing arrays β replacing drives over time rather than waiting for failures spreads the age range and reduces the chance of several failing at once.
Frequently asked questions
How do I know if a hard drive is failing?
Watch reallocated sector count, pending sectors, uncorrectable errors and command timeouts. A growing reallocated sector count is the strongest single predictor. Alert on the values changing rather than on absolute thresholds.
Do power-on hours predict drive failure?
Poorly on their own. A drive at 30,000 hours with clean health values is in better condition than one at 8,000 hours with growing reallocated sectors. Hours measure time powered, not wear.
Why do several drives fail around the same time?
Because drives installed together share batch, age, workload and environment. When one reaches end of life, the others are close behind. It is the most important thing to understand about array planning.
Why is a rebuild the dangerous moment?
It puts sustained heavy load on every remaining drive, at exactly the point when those drives are the same age as the one that failed. Large capacities make it worse because rebuilds take longer.
Should I replace all drives in an ageing array at once?
Staged replacement over time is frequently better. It spreads the age range across the array and reduces the chance of several failing together, without the cost and disruption of replacing everything at once.
Are rising command timeouts a drive problem?
Not always. Rising timeouts on a drive whose other values look healthy frequently indicate a cable or backplane problem rather than the drive itself. Check the connection before replacing the drive.
Send us the health data from a drive you are unsure about and we will tell you honestly whether it needs replacing now or watching.




