smartctl -H said PASSED right up to the day I pulled it, which is the most useful thing I can tell you about smartctl -H.

The overall health assessment is a threshold test. A drive fails it when an attribute crosses the manufacturer's limit. Everything below that limit is PASSED, including a drive whose reallocated sector count has been climbing steadily for two and a half months.

Mine went 0, 0, 0, then 8, then 24, then 56, then 104 over eleven weeks. The threshold was somewhere north of 1000. It would have passed for a long time yet.

What made it visible was recording the numbers rather than the verdict. A cron job writes the raw values of attributes 5 (reallocated sectors), 187 (reported uncorrectable), 197 (current pending) and 198 (offline uncorrectable) to a file once a day. Nothing clever — one line per disk per day. The pattern was obvious within a week of looking at it as a series instead of a snapshot.

The replacement itself was uneventful, which is the whole argument for doing it early. zpool replace, seven hours of resilver on a mirror with the array still serving files, no downtime, no restore, no stress. The failed-drive version of that same afternoon involves a degraded pool, a resilver you are watching nervously, and the non-zero chance that the second drive picks that exact window to find a bad sector of its own.

A disk with 104 reallocated sectors is not dead. It may never die. But it costs about ninety pounds to stop thinking about it, and the alternative is to keep a running tally in your head of how worried to be.

Old drive is in a drawer, wiped, labelled with the date and the sector count. If it turns out to run for another five years I will report back and feel foolish.