Replacing disks at 200 reallocated sectors removed 88% of triple-disk RAID failures at one vendor
Published 2026-10-10
In Ma et al. (FAST 2015), about 1 million SATA disks in backup systems from one storage vendor showed that the count of reallocated sectors (RS) correlates strongly with whole-disk failure. Replacing a disk once its count passed 200 eliminated about 88% of the RAID failures caused by three-disk errors, which were 80% of all RAID failures, or about 70% of all disk-related incidents. The same paper says the signal is incomplete: in simulation, a threshold below 200 caught 52% to 70% of impending whole-disk failures, with 0.8% to 4.5% false positives. Google's FAST 2007 data and Backblaze's 2016 SMART tally point the same way on reallocated sectors, and both also show failed drives with no warning.
- Population
- About 1 million SATA disks, 6 models (each at least 30,000 disks), in backup systems from one vendor
- Window
- Logs from June 2008 over up to 5 years; log lengths from 60 months down to 21 months by model
- Failure definition
- Whole-disk failure: lost connection, operation timeout, or failed write
- Predictor
- Count of reallocated sectors (RS): sectors the disk remapped to a spare after a write error
- Cross-checks
- Google FAST 2007 (more than 100,000 ATA disks) and Backblaze, October 2016 (67,814 disks)
Numbers
| Measure | Value | Unit | Population or scope | Source |
|---|---|---|---|---|
| Triple failures among RAID failures, before the rule | 80 | percent of RAID failures | Vendor's backup systems; 5% other hardware faults and 15% user errors and unknown | FAST 2015 |
| Triple failures avoided after replacing disks at 200 RS | about 88 | percent of triple-failure incidents | Deployed on models A-1, A-2, and B-1 | FAST 2015 |
| All disk-related incidents avoided | about 70 | percent of incidents | Same deployment, normalized to the pre-deployment monthly average | FAST 2015 |
| Deployed replacement threshold | 200 | reallocated sectors | Median time to failure fell below 3 days beyond this count | FAST 2015 |
| Impending failures caught, thresholds below 200 | 52 to 70 | percent of failures within 60 days | Simulation on 100,000 A-2 disks | FAST 2015 |
| False positives, thresholds below 200 | 0.8 to 4.5 | percent of working disks | Simulation on 100,000 A-2 disks | FAST 2015 |
| Failed disks found in their fourth year, A-1, A-2, B-1 | 63, 66, and 64 | percent of failed disks | Each model's failed disks | FAST 2015 |
| Growth of error counts in the second year | 25 to about 300 | percent increase | Disks with at least one sector error; C-2 lowest and A-2 highest | FAST 2015 |
| Triple failures left after the rule, from hard-to-predict causes | 12 | percent of triple failures | Sudden failures or several mildly unreliable disks | FAST 2015 |
| RAID groups flagged as vulnerable by ARMOR | 80 | percent of vulnerable RAID-6 groups | Simulation: 5,000 healthy groups and 500 failed groups | FAST 2015 |
| Total triple-failure coverage with ARMOR added | 98 | percent of triple failures | Simulation only; not deployed | FAST 2015 |
| Size of the vendor study | about 1 million disks; 6 models | disks; models | Each model has at least 30,000 disks; logs of 21 to 60 months | FAST 2015 |
| Google: drives with a reallocation count above zero | about 9 | percent of drives | More than 100,000 ATA disks, December 2005 to August 2006 | Google, FAST 2007 |
| Google: failure likelihood within 60 days after the first reallocation | over 14 | times the rate for drives with none | Same population | Google, FAST 2007 |
| Backblaze: failed versus operational drives with one or more of five SMART counts above zero | 76.7 versus 4.2 | percent of drives | 67,814 drives in 2016; 23.3% of failed drives showed none | Backblaze 2016 |
This table is a short extract of printed figures, not a copy of the papers and not the logs. Rows 1 to 12 are from Ma et al. (FAST 2015); rows 13 to 15 are from the two cross-check sources and are different measures on different fleets.
Method
The figures are copied from the USENIX PDF of Ma, Douglis, Lu, Sawyer, Chandra, and Hsu, FAST 2015 (pages 241 to 256), not refit. The authors analysed error logs sent back from EMC Data Domain backup systems, about 1 million SATA disks in 6 models, with log lengths of 21 to 60 months from June 2008. A whole-disk failure is defined by the systems as a lost connection to the disk, an operation that exceeds the timeout threshold, or a failed write. They compare decile distributions of error counts on failed and working disks. They test the predictor, called PLATE, by simulation on 100,000 disks of one model (a disk is predicted to fail if its count passes a threshold and it fails within 60 days; a disk that passes the threshold and survives the 60 days is a false positive), and then report what happened after the rule was deployed in production. A second component, ARMOR, scores a whole RAID group by combining per-disk failure probabilities and was only simulated, on 5,000 healthy and 500 failed groups. Two other sources check the direction of the signal: Pinheiro, Weber, and Barroso, FAST 2007 (a Google population) and a Backblaze post of 6 October 2016. Neither tests the 88% figure.
Limits
The 88% comes from the vendor's own deployment, compared before and after, and is normalized to the average number of RAID failures per month before deployment. The paper says it cannot release error rates for disk groups or disk models, so absolute failure rates are not available, and disk models are anonymized. The systems are write-heavy backup systems, which the authors say may make write errors more common than read errors. The deployed rule covered three of the six models (A-1, A-2, and B-1) and had been running for nearly a year. The simulated catch rate and false-positive rate come from one model, and the introduction prints 'up to 65% with up to 2.5% false alarms' while the evaluation section prints 52% to 70% with 0.8% to 4.5% for thresholds below 200. About 20% of RAID-group failures come from user errors, hardware faults, and unknown causes that a sector count cannot predict. The 88% is the share of triple failures avoided, and it is larger than the share of single failures predicted because avoiding one of three disk failures is enough to save a RAID-6 group. ARMOR was not deployed, so its 98% total coverage is a simulation. The paper reports that the proactively replaced disks were tested by the vendor's specialists, who saw no noticeable number of false positives, but it does not give the figure. The Backblaze post does not publish the disk counts behind its percentages in the text I opened, and Google's population is consumer-grade ATA disks from 2005 and 2006.
What the paper found
The answer. Counting reallocated sectors and acting on the count worked: after the vendor started replacing disks whose count passed 200, RAID failures caused by triple-disk errors dropped by about 88%, equal to about 70% of all disk-related incidents. Triple failures had been 80% of RAID failures, and the rest were 5% other hardware faults (adapters, cables, shelves) and 15% user errors and unknown causes. The paper explains the large drop: preventing one of the three disk failures is enough to save a RAID-6 group, so catching only part of the failing disks removes most triple failures. The remaining 12% of triple failures came from sudden failures or from several somewhat unreliable disks that each stayed under the threshold.
Why reallocated sectors. The authors compared several sector-error counts and report RS as strongly correlated with whole-disk failure: failed disks tend to have more RS than working ones. They reason that a reallocation is the last resort after other recovery has failed, so it filters out the temporary errors that also show up on working disks. Disks also failed at similar ages: 63% of A-1, 66% of A-2, and 64% of B-1 failed disks were found in the fourth year, and 68% of C-2 failures came in the second year, which raises the chance that several disks in one group fail close together. Among disks that had at least one sector error, the average count in the second year was 25% (C-2) to about 300% (A-2) higher than in the first, and about 5% of A-2 disks had sector errors in the first 30 months, then 10% more in the next 6 months.
The threshold and its cost. The authors picked 200 RS because replacing a disk could take up to 3 days, and beyond 200 RS the median time to failure fell below 3 days. They also wanted a false-positive rate below 1%, so that working disks are not replaced for no reason. In simulation on 100,000 A-2 disks, thresholds below 200 caught 52% to 70% of impending whole-disk failures (failing within 60 days), with 0.8% to 4.5% false positives. The failures it misses are mostly caused by hardware faults, user errors, and unknown reasons that are not visible as sector errors.
What the cross-checks say. Google's population of more than 100,000 consumer ATA disks found about 9% with a reallocation count above zero, drives more than 14 times likelier to fail within 60 days after the first reallocation, and over 56% of failed drives with no count on any of four strong SMART signals including reallocation. Backblaze's 2016 post lists reallocated sectors (SMART 5) among five attributes it watches and reports that 76.7% of failed drives, against 4.2% of operational drives, had one or more of the five above zero, so 23.3% of failed drives gave no SMART warning. All three agree that a reallocated sector is a strong warning and that it is not a complete one. The numbers differ because the fleets, the failure definitions, and the sets of attributes differ.
The medium behind these counts is thehard disk drivecard. The warning-sign question is also covered in theSMART failure signals analysis.
Sources
- 01RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures, Ma, Douglis, Lu, Sawyer, Chandra, and Hsu, FAST 2015 (USENIX PDF, pages 241 to 256) · accessed 2026-10-10
- 02Failure Trends in a Large Disk Drive Population, Pinheiro, Weber, and Barroso, FAST 2007 (USENIX HTML proceedings) · accessed 2026-10-09
- 03What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 · accessed 2026-10-10