Hard drives were 82% of 290,000 datacenter hardware failure tickets, and 2% of failed servers produced over 99% of the failures
Published 2026-10-11
In Wang, Zhang, and Xu (DSN 2017), over 290,000 hardware failure tickets from four years at one large Internet company, hard drives were 81.84% of all failures, against 3.06% for memory and 0.31% for SSDs. The failures were very uneven: 2% of the servers that ever failed produced over 99% of the failures, and 35 of 1,411 days (2.48%) had over 500 hard drive failures. In one case 32% of a product line's servers reported hard drive SMART alerts in one night. Schroeder and Gibson (FAST 2007) found disks were 20% to 50% of hardware replacements in three other systems, and Ford et al. (OSDI 2010) found 37% of Google node failures came in bursts, so the dominance of drives and the clustering both show up elsewhere.
- Population
- One large Internet company, dozens of datacenters, hundreds of thousands of servers
- Window
- Four years of failure tickets
- Failure definition
- Any ticket in the fixing or unrepaired classes, including SMART and read or write error warnings; false alarms excluded
- Batch definition
- A day with more than N (100, 200, or 500) failures of one component class
- Cross-checks
- Replacement logs of HPC1, COM1, and COM2, FAST 2007; Google storage cells, OSDI 2010
Numbers
This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 11 are from Wang et al. (DSN 2017). Row 12 is from Schroeder and Gibson (FAST 2007) and rows 13 and 14 from Ford et al. (OSDI 2010); they are cross-checks with different measures.
Method
The figures are copied from the author PDF of Wang, Zhang, and Xu, IEEE/IFIP DSN 2017, not refit. The data are all hardware failure operation tickets (FOTs) from the failure management system of one large Internet company over four years, covering dozens of datacenters. About 90% of the tickets come from agents that read system logs and poll device status, and about 10% are entered by hand as miscellaneous. The paper counts every ticket that is fixed or left unrepaired as a failure and excludes false alarms. A batch day is a day with more than N failures of one component class. Two other sources check the surrounding claims: Schroeder and Gibson, FAST 2007 (replacement logs of HPC1, COM1, and COM2) for the share of disks among hardware replacements, and Ford et al., OSDI 2010 (Google storage cells) for failures that cluster in time. Neither tests the 2% and 99% result.
Limits
The numbers are one company's own fleet, hardware, and ticketing rules, and the company is not named in the paper. A failure here is a ticket, and the authors say there is no agreement on whether a drive with SMART alerts or occasional read and write errors is failed, so the 81.84% share mixes fatal stops with warnings; it is not an annual failure rate and cannot be compared with rates per drive-year. The paper has no device counts per class in the text I read, so the shares say how many tickets there were, not how likely a given drive was to fail. The response time of 42.2 days is the time operators took to act on a ticket, not a repair time. The paper does not say why the November 2015 batch happened. The Schroeder and Gibson shares are from replacement logs of other organizations and the Ford figure is for node failures in a distributed file system, so those numbers use other definitions and are not comparable with the tickets.
What the paper found
The answer. In Wang et al., hard drives dominate the failure tickets: 81.84% of all failures, against 3.06% for memory, 1.74% for power, 1.23% for RAID cards, 0.67% for flash cards, and 0.31% for SSDs. Another 10.20% are miscellaneous tickets that operators entered by hand, and the operators suspect about 25% of those are hard drive related. About half of the hard drive tickets relate to SMART values or prediction error counts, which are warnings and not fatal stops, so the share is a share of tickets, not of dead drives. In Schroeder and Gibson, disks were also the most commonly replaced part in two of three systems, 30% in HPC1 and 50% in COM2, and nearly 20% in COM1, so the main direction matches while the size does not.
Failures are concentrated. In Wang et al., 2% of the servers that ever failed contribute more than 99% of all failures, and one server in a web product line reported over 400 failures from a failing RAID card battery unit, with each automatic recovery marking the problem solved before it came back. Even so, repeats are not the norm: over 85% of fixed components never repeat the same failure, and about 4.5% of servers that ever failed had repeating failures. Ford et al. found the same kind of concentration at the disk level, where checksum mismatches found by scrubbing, 1 in 10^6 to 10^7 older blocks, were concentrated on a small number of disks.
Failures also come in batches. During 2.48% of days, 35 of 1,411, the company saw over 500 hard drive failures in a single day. In a November 2015 case, thousands of servers, 32% of one product line, reported hard drive SMART alerts, 99% of them between 21:00 and 3:00; the operators replaced about 28% of the drives and decommissioned the 70% or more that were out of warranty, and the cause is not known. Ford et al. found that 37% of node failures in Google cells were part of a burst of at least 2 nodes within 120 seconds, so shared causes such as batches, switches, or power are not rare.
What this does not show. Over 1/4 of failures were in out-of-warranty hardware and were not handled at all, the mean time operators took to respond to a ticket was 42.2 days against a median of 6.1 days, and the mix of tickets depends on what the detectors report. Treat the figures as how one large fleet's failure queue looks, not as the failure probability of a drive.
Background on drive types is on thedrive technologycard. Warning signs before a failure are also covered in theSMART failure signals analysis.
Sources
- 01What Can We Learn from Four Years of Data Center Hardware Failures?, Wang, Zhang, and Xu, IEEE/IFIP DSN 2017, pages 25 to 36 (author PDF, copy on the netman.aiops.org reading list) · accessed 2026-10-11
- 02Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 4 on other components) · accessed 2026-10-11
- 03Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF) · accessed 2026-10-11