Skip to content

Paper note

Replacing disks at 200 reallocated sectors removed 88% of triple-disk RAID failures at one vendor

Published 2026-10-10

In Ma et al. (FAST 2015), about 1 million SATA disks in backup systems from one storage vendor showed that the count of reallocated sectors (RS) correlates strongly with whole-disk failure. Replacing a disk once its count passed 200 eliminated about 88% of the RAID failures caused by three-disk errors, which were 80% of all RAID failures, or about 70% of all disk-related incidents. The same paper says the signal is incomplete: in simulation, a threshold below 200 caught 52% to 70% of impending whole-disk failures, with 0.8% to 4.5% false positives. Google's FAST 2007 data and Backblaze's 2016 SMART tally point the same way on reallocated sectors, and both also show failed drives with no warning.

Population
About 1 million SATA disks, 6 models (each at least 30,000 disks), in backup systems from one vendor
Window
Logs from June 2008 over up to 5 years; log lengths from 60 months down to 21 months by model
Failure definition
Whole-disk failure: lost connection, operation timeout, or failed write
Predictor
Count of reallocated sectors (RS): sectors the disk remapped to a spare after a write error
Cross-checks
Google FAST 2007 (more than 100,000 ATA disks) and Backblaze, October 2016 (67,814 disks)

Numbers

Reallocated sectors (RS) and RAID failures, as printed in Ma et al., FAST 2015 (about 1 million SATA disks, 6 models, backup systems from one vendor, June 2008 onward), rows 1 to 12, and in two cross-check sources, rows 13 to 15. RS is the count of sectors remapped to a spare after a write error. A triple failure is three simultaneous whole-disk or sector failures in one RAID-6 group. Qualifiers such as 'about' and 'up to' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Triple failures among RAID failures, before the rule80percent of RAID failuresVendor's backup systems; 5% other hardware faults and 15% user errors and unknownFAST 2015
Triple failures avoided after replacing disks at 200 RSabout 88percent of triple-failure incidentsDeployed on models A-1, A-2, and B-1FAST 2015
All disk-related incidents avoidedabout 70percent of incidentsSame deployment, normalized to the pre-deployment monthly averageFAST 2015
Deployed replacement threshold200reallocated sectorsMedian time to failure fell below 3 days beyond this countFAST 2015
Impending failures caught, thresholds below 20052 to 70percent of failures within 60 daysSimulation on 100,000 A-2 disksFAST 2015
False positives, thresholds below 2000.8 to 4.5percent of working disksSimulation on 100,000 A-2 disksFAST 2015
Failed disks found in their fourth year, A-1, A-2, B-163, 66, and 64percent of failed disksEach model's failed disksFAST 2015
Growth of error counts in the second year25 to about 300percent increaseDisks with at least one sector error; C-2 lowest and A-2 highestFAST 2015
Triple failures left after the rule, from hard-to-predict causes12percent of triple failuresSudden failures or several mildly unreliable disksFAST 2015
RAID groups flagged as vulnerable by ARMOR80percent of vulnerable RAID-6 groupsSimulation: 5,000 healthy groups and 500 failed groupsFAST 2015
Total triple-failure coverage with ARMOR added98percent of triple failuresSimulation only; not deployedFAST 2015
Size of the vendor studyabout 1 million disks; 6 modelsdisks; modelsEach model has at least 30,000 disks; logs of 21 to 60 monthsFAST 2015
Google: drives with a reallocation count above zeroabout 9percent of drivesMore than 100,000 ATA disks, December 2005 to August 2006Google, FAST 2007
Google: failure likelihood within 60 days after the first reallocationover 14times the rate for drives with noneSame populationGoogle, FAST 2007
Backblaze: failed versus operational drives with one or more of five SMART counts above zero76.7 versus 4.2percent of drives67,814 drives in 2016; 23.3% of failed drives showed noneBackblaze 2016

This table is a short extract of printed figures, not a copy of the papers and not the logs. Rows 1 to 12 are from Ma et al. (FAST 2015); rows 13 to 15 are from the two cross-check sources and are different measures on different fleets.

Method

The figures are copied from the USENIX PDF of Ma, Douglis, Lu, Sawyer, Chandra, and Hsu, FAST 2015 (pages 241 to 256), not refit. The authors analysed error logs sent back from EMC Data Domain backup systems, about 1 million SATA disks in 6 models, with log lengths of 21 to 60 months from June 2008. A whole-disk failure is defined by the systems as a lost connection to the disk, an operation that exceeds the timeout threshold, or a failed write. They compare decile distributions of error counts on failed and working disks. They test the predictor, called PLATE, by simulation on 100,000 disks of one model (a disk is predicted to fail if its count passes a threshold and it fails within 60 days; a disk that passes the threshold and survives the 60 days is a false positive), and then report what happened after the rule was deployed in production. A second component, ARMOR, scores a whole RAID group by combining per-disk failure probabilities and was only simulated, on 5,000 healthy and 500 failed groups. Two other sources check the direction of the signal: Pinheiro, Weber, and Barroso, FAST 2007 (a Google population) and a Backblaze post of 6 October 2016. Neither tests the 88% figure.

Limits

The 88% comes from the vendor's own deployment, compared before and after, and is normalized to the average number of RAID failures per month before deployment. The paper says it cannot release error rates for disk groups or disk models, so absolute failure rates are not available, and disk models are anonymized. The systems are write-heavy backup systems, which the authors say may make write errors more common than read errors. The deployed rule covered three of the six models (A-1, A-2, and B-1) and had been running for nearly a year. The simulated catch rate and false-positive rate come from one model, and the introduction prints 'up to 65% with up to 2.5% false alarms' while the evaluation section prints 52% to 70% with 0.8% to 4.5% for thresholds below 200. About 20% of RAID-group failures come from user errors, hardware faults, and unknown causes that a sector count cannot predict. The 88% is the share of triple failures avoided, and it is larger than the share of single failures predicted because avoiding one of three disk failures is enough to save a RAID-6 group. ARMOR was not deployed, so its 98% total coverage is a simulation. The paper reports that the proactively replaced disks were tested by the vendor's specialists, who saw no noticeable number of false positives, but it does not give the figure. The Backblaze post does not publish the disk counts behind its percentages in the text I opened, and Google's population is consumer-grade ATA disks from 2005 and 2006.

What the paper found

The answer. Counting reallocated sectors and acting on the count worked: after the vendor started replacing disks whose count passed 200, RAID failures caused by triple-disk errors dropped by about 88%, equal to about 70% of all disk-related incidents. Triple failures had been 80% of RAID failures, and the rest were 5% other hardware faults (adapters, cables, shelves) and 15% user errors and unknown causes. The paper explains the large drop: preventing one of the three disk failures is enough to save a RAID-6 group, so catching only part of the failing disks removes most triple failures. The remaining 12% of triple failures came from sudden failures or from several somewhat unreliable disks that each stayed under the threshold.

Why reallocated sectors. The authors compared several sector-error counts and report RS as strongly correlated with whole-disk failure: failed disks tend to have more RS than working ones. They reason that a reallocation is the last resort after other recovery has failed, so it filters out the temporary errors that also show up on working disks. Disks also failed at similar ages: 63% of A-1, 66% of A-2, and 64% of B-1 failed disks were found in the fourth year, and 68% of C-2 failures came in the second year, which raises the chance that several disks in one group fail close together. Among disks that had at least one sector error, the average count in the second year was 25% (C-2) to about 300% (A-2) higher than in the first, and about 5% of A-2 disks had sector errors in the first 30 months, then 10% more in the next 6 months.

The threshold and its cost. The authors picked 200 RS because replacing a disk could take up to 3 days, and beyond 200 RS the median time to failure fell below 3 days. They also wanted a false-positive rate below 1%, so that working disks are not replaced for no reason. In simulation on 100,000 A-2 disks, thresholds below 200 caught 52% to 70% of impending whole-disk failures (failing within 60 days), with 0.8% to 4.5% false positives. The failures it misses are mostly caused by hardware faults, user errors, and unknown reasons that are not visible as sector errors.

What the cross-checks say. Google's population of more than 100,000 consumer ATA disks found about 9% with a reallocation count above zero, drives more than 14 times likelier to fail within 60 days after the first reallocation, and over 56% of failed drives with no count on any of four strong SMART signals including reallocation. Backblaze's 2016 post lists reallocated sectors (SMART 5) among five attributes it watches and reports that 76.7% of failed drives, against 4.2% of operational drives, had one or more of the five above zero, so 23.3% of failed drives gave no SMART warning. All three agree that a reallocated sector is a strong warning and that it is not a complete one. The numbers differ because the fleets, the failure definitions, and the sets of attributes differ.

The medium behind these counts is thehard disk drivecard. The warning-sign question is also covered in theSMART failure signals analysis.

Sources

  1. 01RAIDShield: Characterizing, Monitoring, and Proactively Protecting Against Disk Failures, Ma, Douglis, Lu, Sawyer, Chandra, and Hsu, FAST 2015 (USENIX PDF, pages 241 to 256) · accessed 2026-10-10
  2. 02Failure Trends in a Large Disk Drive Population, Pinheiro, Weber, and Barroso, FAST 2007 (USENIX HTML proceedings) · accessed 2026-10-09
  3. 03What SMART Stats Tell Us About Hard Drives, Backblaze, 6 October 2016 · accessed 2026-10-10