Skip to content

Paper note

Hard drives were 82% of 290,000 datacenter hardware failure tickets, and 2% of failed servers produced over 99% of the failures

Published 2026-10-11

In Wang, Zhang, and Xu (DSN 2017), over 290,000 hardware failure tickets from four years at one large Internet company, hard drives were 81.84% of all failures, against 3.06% for memory and 0.31% for SSDs. The failures were very uneven: 2% of the servers that ever failed produced over 99% of the failures, and 35 of 1,411 days (2.48%) had over 500 hard drive failures. In one case 32% of a product line's servers reported hard drive SMART alerts in one night. Schroeder and Gibson (FAST 2007) found disks were 20% to 50% of hardware replacements in three other systems, and Ford et al. (OSDI 2010) found 37% of Google node failures came in bursts, so the dominance of drives and the clustering both show up elsewhere.

Population
One large Internet company, dozens of datacenters, hundreds of thousands of servers
Window
Four years of failure tickets
Failure definition
Any ticket in the fixing or unrepaired classes, including SMART and read or write error warnings; false alarms excluded
Batch definition
A day with more than N (100, 200, or 500) failures of one component class
Cross-checks
Replacement logs of HPC1, COM1, and COM2, FAST 2007; Google storage cells, OSDI 2010

Numbers

Hardware failure ticket figures as printed in Wang et al., DSN 2017 (one large Internet company, four years), rows 1 to 11, and in two cross-check sources, rows 12 to 14. Qualifiers such as 'over' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
Hardware failure tickets analysedover 290,000ticketsOne large Internet company, dozens of datacenters, hundreds of thousands of servers
Length of the record4yearsSame company
Hard drives as a share of all failures81.84percent of failuresFailure tickets excluding false alarms
Memory as a share of all failures3.06percent of failuresSame tickets
SSDs as a share of all failures0.31percent of failuresSame tickets
Failures in out-of-warranty hardware that are not handled25percent of failures (lower bound)Same tickets
Share of failed servers that produced over 99% of all failures2percent of servers that ever failedSame tickets
Fixed components that never repeat the same failure85percent (lower bound)Same tickets
Days with over 500 hard drive failures2.48percent of days35 of 1,411 days
Servers of one product line reporting hard drive SMART alerts in a single night32percent of that product line's serversNovember 2015 case study; 99% detected within about 6 hours
Mean time to respond to a failure ticket42.2daysTickets that led to a repair order
Disks as a share of all hardware replacements20 to 50percent of replacementsThree systems: HPC1 30%, COM2 50%, COM1 nearly 20%
Node failures that are part of a burst of at least 2 nodes37percent of failuresTens of Google storage cells, 120 second window
Older data blocks that fail a checksum on scrubbing1 in 10^6 to 10^7fraction of older blocksGoogle file system, background scrubbing

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 11 are from Wang et al. (DSN 2017). Row 12 is from Schroeder and Gibson (FAST 2007) and rows 13 and 14 from Ford et al. (OSDI 2010); they are cross-checks with different measures.

Method

The figures are copied from the author PDF of Wang, Zhang, and Xu, IEEE/IFIP DSN 2017, not refit. The data are all hardware failure operation tickets (FOTs) from the failure management system of one large Internet company over four years, covering dozens of datacenters. About 90% of the tickets come from agents that read system logs and poll device status, and about 10% are entered by hand as miscellaneous. The paper counts every ticket that is fixed or left unrepaired as a failure and excludes false alarms. A batch day is a day with more than N failures of one component class. Two other sources check the surrounding claims: Schroeder and Gibson, FAST 2007 (replacement logs of HPC1, COM1, and COM2) for the share of disks among hardware replacements, and Ford et al., OSDI 2010 (Google storage cells) for failures that cluster in time. Neither tests the 2% and 99% result.

Limits

The numbers are one company's own fleet, hardware, and ticketing rules, and the company is not named in the paper. A failure here is a ticket, and the authors say there is no agreement on whether a drive with SMART alerts or occasional read and write errors is failed, so the 81.84% share mixes fatal stops with warnings; it is not an annual failure rate and cannot be compared with rates per drive-year. The paper has no device counts per class in the text I read, so the shares say how many tickets there were, not how likely a given drive was to fail. The response time of 42.2 days is the time operators took to act on a ticket, not a repair time. The paper does not say why the November 2015 batch happened. The Schroeder and Gibson shares are from replacement logs of other organizations and the Ford figure is for node failures in a distributed file system, so those numbers use other definitions and are not comparable with the tickets.

What the paper found

The answer. In Wang et al., hard drives dominate the failure tickets: 81.84% of all failures, against 3.06% for memory, 1.74% for power, 1.23% for RAID cards, 0.67% for flash cards, and 0.31% for SSDs. Another 10.20% are miscellaneous tickets that operators entered by hand, and the operators suspect about 25% of those are hard drive related. About half of the hard drive tickets relate to SMART values or prediction error counts, which are warnings and not fatal stops, so the share is a share of tickets, not of dead drives. In Schroeder and Gibson, disks were also the most commonly replaced part in two of three systems, 30% in HPC1 and 50% in COM2, and nearly 20% in COM1, so the main direction matches while the size does not.

Failures are concentrated. In Wang et al., 2% of the servers that ever failed contribute more than 99% of all failures, and one server in a web product line reported over 400 failures from a failing RAID card battery unit, with each automatic recovery marking the problem solved before it came back. Even so, repeats are not the norm: over 85% of fixed components never repeat the same failure, and about 4.5% of servers that ever failed had repeating failures. Ford et al. found the same kind of concentration at the disk level, where checksum mismatches found by scrubbing, 1 in 10^6 to 10^7 older blocks, were concentrated on a small number of disks.

Failures also come in batches. During 2.48% of days, 35 of 1,411, the company saw over 500 hard drive failures in a single day. In a November 2015 case, thousands of servers, 32% of one product line, reported hard drive SMART alerts, 99% of them between 21:00 and 3:00; the operators replaced about 28% of the drives and decommissioned the 70% or more that were out of warranty, and the cause is not known. Ford et al. found that 37% of node failures in Google cells were part of a burst of at least 2 nodes within 120 seconds, so shared causes such as batches, switches, or power are not rare.

What this does not show. Over 1/4 of failures were in out-of-warranty hardware and were not handled at all, the mean time operators took to respond to a ticket was 42.2 days against a median of 6.1 days, and the mix of tickets depends on what the detectors report. Treat the figures as how one large fleet's failure queue looks, not as the failure probability of a drive.

Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.

Sources

  1. 01What Can We Learn from Four Years of Data Center Hardware Failures?, Wang, Zhang, and Xu, IEEE/IFIP DSN 2017, pages 25 to 36 (author PDF, copy on the netman.aiops.org reading list) · accessed 2026-10-11
  2. 02Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 4 on other components) · accessed 2026-10-11
  3. 03Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF) · accessed 2026-10-11