Reading Drive Health Data: Which Attributes Matter
Drive health data is reported by every enterprise drive and read by almost nobody until something has already failed. Read properly, it gives weeks of warning; read carelessly, it produces false alarms and ignored ones.
Raw values are not the useful part
The first thing that confuses people.
Health attributes are reported with a raw value and a normalised value, and the raw numbers are vendor-specific. A raw figure that looks alarming on one drive is normal on another, because manufacturers encode different things in the same field.
So the useful readings are:
Whether an attribute has crossed its threshold, which is the drive’s own judgement about itself.
Whether a value is changing over time, which is the strongest signal available.
A single reading tells you very little. The same attribute checked monthly tells you a great deal — and this is why recording a baseline matters, the same principle our performance guide applies.
The attributes that actually predict failure
Most attributes are noise for this purpose. These are the ones worth watching.
Reallocated sectors. Sectors found bad and remapped to spares. A small number is normal on many drives. A count that is climbing is the clearest warning there is, because it means the drive is actively finding new defects.
Pending sectors. Sectors the drive suspects are bad but has not yet remapped, because it has not been able to write to them. Frequently more urgent than reallocated, since these represent data that may not be readable.
Uncorrectable sectors. Sectors that could not be recovered. Any non-zero value warrants attention.
Reported uncorrectable errors and command timeouts. These slow the whole array, because operations wait on the struggling drive — which is why one degrading drive makes everything feel slow.
Spin retry count on mechanical drives, indicating difficulty starting.
The single most useful habit: watch reallocated and pending counts for change rather than for absolute value.
Attributes that are informative but not predictive
Worth reading and worth not panicking about.
Power-on hours. Tells you age, not condition. High hours on a drive with clean error counts is a healthy drive that has been working. This matters when buying refurbished — hours are expected, error counts are the question. Our grading guide covers what to ask a seller.
Temperature. Useful as a trend. A drive running consistently hotter than its neighbours suggests an airflow problem rather than a drive problem — our airflow guide covers why.
Load cycle count. Can rise quickly on drives with aggressive power management, which is a configuration matter rather than a fault.
Start-stop count. Largely irrelevant on drives that run continuously.
On SSDs, read different things
Flash fails differently, so the attributes differ.
Wear indicator or life remaining — how much of the rated endurance has been consumed. This is the primary health measure, and it is predictable in a way mechanical wear is not. Our endurance guide covers DWPD and TBW.
Total data written, which lets you calculate the actual write rate and project when endurance will be reached.
Available spare blocks, declining as blocks are retired.
A useful consequence: because endurance is measurable and consumption is trackable, SSD replacement can genuinely be planned rather than reacted to. Mechanical drive failure is far less predictable.
What the data does not tell you
Being honest about the limits, because overconfidence here causes its own problems.
Drives fail without warning. A meaningful proportion of failures show nothing beforehand — electronics failures and mechanical events that happen suddenly. Clean health data is not a guarantee.
It does not replace redundancy or backups. Health monitoring buys warning time where warning is available. Our backup guide covers why RAID is not backup, and our RAID guide covers rebuild risk.
A controller may hide it. Drives behind a hardware RAID controller frequently do not present health data to the operating system directly — you read it through the controller utility instead, which is one of the arguments for passthrough where a software layer needs it. Our controller guide covers this.
Acting on what you find
A workable policy.
Climbing reallocated or pending counts — plan replacement now, while the array is healthy and you choose the timing. Replacing a degrading drive deliberately is far better than replacing it during a rebuild.
Any uncorrectable errors — treat as failing.
Correlated ageing across an array — drives installed together reach wear-out together, so several drives showing early signs is a stronger message than one. Our lifespan guide covers why staggering replacement matters.
SSD wear indicator approaching its limit — schedule replacement against the measured write rate.
And route the alerts somewhere a person reads out of hours. Health data nobody looks at is the same as no health data — the point our monitoring guide keeps returning to.
Common questions
Why do raw values look alarming?
Because they are vendor-specific — manufacturers encode different things in the same field, so a figure that looks bad on one drive is normal on another. Watch whether thresholds are crossed and whether values are changing over time.
Which attributes actually predict failure?
Reallocated and pending sectors, watched for change rather than absolute value, plus uncorrectable errors and command timeouts. A climbing reallocated count is the clearest warning available.
Are high power-on hours a problem?
Not in themselves. Hours tell you age, not condition — high hours with clean error counts is a healthy drive that has been working. On refurbished stock, hours are expected and error counts are the real question.
Is health data different on SSDs?
Yes. The primary measure is the wear indicator or life remaining, alongside total data written and available spare blocks. Because endurance is measurable and consumption trackable, SSD replacement can genuinely be planned rather than reacted to.
Why can I not see health data on my drives?
Drives behind a hardware RAID controller frequently do not present it to the operating system directly — read it through the controller utility instead. That is one of the arguments for passthrough where a software layer needs the data.
This is a general guide to interpreting drive health data rather than a diagnosis — if counts are climbing on a production array, plan replacement while you still choose the timing.




