Skip to content

Paper note

In an 18-month Berkeley study of 368 disks, 7 data disks failed against 13 enclosure backplane failures; disks were 90% of parts but about 4% of errors

Published 2026-10-11

In Talagala and Patterson (UC Berkeley, 1999), 18 months of operation of a prototype with 20 PCs and 368 SCSI data disks, 34 components were replaced. Only 7 of the 368 data disks (1.9%) failed, the lowest percentage of any failing component, against 13 of 46 disk enclosures (28.3%, backplane faults) and 6 of 24 IDE system disks (25.0%). Disks were about 90% of the components but only around 4% of 688 logged error instances, while SCSI timeouts and parity errors were 49% (87% without network errors), and no restart was caused by a data disk. Ford et al. (OSDI 2010) later found a disk MTTF of 10 to 50 years against 4.3 months for a node, and Schroeder and Gibson (FAST 2007) found disks were 20% to 50% of hardware replacements, a count not a rate.

Population
One prototype: 20 PCs, 368 SCSI data disks, 24 IDE system disks, 46 enclosures
Window
18 months of replacements; 6 months of error logs
Failure definition
A component that was replaced (absolute failure)
Error instance
Log messages of one category within 10 seconds of each other
Cross-checks
Replacement logs of HPC1, COM1, and COM2, FAST 2007; Google storage cells, OSDI 2010

Numbers

Component failure and error figures as printed in Talagala and Patterson (UC Berkeley, 1999; 18 months, 368 disks), rows 1 to 12, and in two cross-check sources, rows 13 to 15. Qualifiers such as 'over' and 'around' are the sources' wording.
MeasureValueUnitPopulation or scopeSource
SCSI data disks in the prototype368disks20 PCs, 3.2 TB, 8.4 GB disks, hardware from 1996
Length of the failure record18monthsPrototype operation
Components replaced in the failure record34componentsAll component types
SCSI data disks replaced1.9percent of 368 disks in 18 months7 of 368 disks
Disk enclosure backplane failures28.3percent of 46 enclosures in 18 months13 of 46 enclosures
IDE system disks replaced25.0percent of 24 disks in 18 months6 of 24 disks
Enclosure power supplies replaced3.26percent of 92 supplies in 18 months3 of 92 supplies
Error instances in the logs688error instances6 months of logs, 16 or 17 machines
Share of errors that were name-lookup and file-mount errorsover 40percent of error instances (lower bound)Same logs
Share of errors that were SCSI timeouts or parity errors49 to 87percent of error instances49% of all errors; 87% once network errors are removed
Share of errors that came from data disksaround 4percent of error instances (about)Same logs
Restarts caused by data disk or IDE disk errors0restarts of 7316 nodes, 6 months
Disks as a share of hardware replacements20 to 50percent of replacementsHPC1 30%, COM2 50%, COM1 nearly 20%
Disk mean time to failure10 to 50yearsTens of Google storage cells
Node mean time to failure4.3monthsSame cells

This table is a short extract of printed figures, not a copy of the papers and not the data. Rows 1 to 12 are from Talagala and Patterson (1999). Row 13 is from Schroeder and Gibson (FAST 2007) and rows 14 and 15 from Ford et al. (OSDI 2010); they are cross-checks with different measures.

Method

The figures are copied from the author PDF of Talagala and Patterson, UC Berkeley technical report UCB/CSD-99-1042, not refit. The prototype was a web-accessible image collection on 20 PCs running FreeBSD 2.2 with 368 8.4 GB SCSI data disks in cooled enclosures and 24 IDE system disks. A component counts as failed when it was replaced ('absolute failure'), over 18 months. Error logs from the last 6 months for 16 or 17 of the machines were filtered, grouped into error instances when messages of one category fell within 10 seconds, and sorted by hand into eleven categories. Two other sources check the surrounding claims: Schroeder and Gibson, FAST 2007 (replacement logs of three other systems) and Ford et al., OSDI 2010 (Google storage cells). Neither tests the prototype's counts.

Limits

The numbers are one small research prototype built from 1996 parts, with a few hundred disks, and its reliability says little about modern drives. Counts are small: 7 disk failures and 1 failure each for several other component types, so the percentages are noisy and the paper warns that components of different numbers cannot be compared directly. The 18-month figures are not annualized in the paper. The error-instance table lists fewer data disk errors than the text's 'around 4%', and the text and a table caption disagree on whether 16 or 17 machines were logged; this note uses the text's wording. Failures are replacements, not fail-stop events. Schroeder and Gibson count replacements in other systems and Ford et al. count failures in a distributed file system, so those numbers use other definitions and are not comparable with the prototype's percentages.

What the paper found

The answer. In Talagala and Patterson, the data disks were the most reliable part of the system, and the parts around them were not. Of 368 SCSI data disks, 7 (1.9%) were replaced in 18 months, the lowest percentage among failing component types. The disk enclosures, which hold the SCSI bus backplane, had 13 failures in 46 (28.3%), the IDE system disks 6 in 24 (25.0%), and enclosure power supplies 3 in 92 (3.26%), so in total 34 components were replaced, nearly two a month. The authors suspect the IDE disks fared worse because they sat in ordinary PC chassis rather than enclosures built for cooling and low vibration.

Errors point at the cabling and the bus. Of 688 error instances in 6 months of logs, over 40% were name-lookup and file-mount errors caused by dependence on external servers, and SCSI timeouts and parity errors were 49% of all errors, rising to 87% once the network errors are set aside. Data disk errors were only around 4% of errors although disks were about 90% of components, and most were recovered errors. The machines with the most SCSI errors traced to an enclosure replacement and likely loose cables. No restart of a node was caused by a data disk or IDE disk error; the largest causes were external network failures and a single power outage, 22% of restarts.

Warnings and sharing. The authors saw SCSI timeout and parity errors escalate over time on machines 1, 3, and 9, and suggest disk and SCSI failures look predictable, while a failed external server was reported as 16 error instances at once. Their message is that independent-failure assumptions are too simple and that single points of failure and shared components dominate what operators see.

Later fleets agree on the direction. Ford et al. measured a disk mean time to failure of 10 to 50 years against 4.3 months for a node in Google's cells, and found nodes drove availability. Schroeder and Gibson found disks were the most commonly replaced part in two of three systems (30% in HPC1, 50% in COM2) and nearly 20% in COM1, but warn that this reflects how many disks there are, not that disks are less reliable. The systems, years, and measures differ, so the sizes are not comparable, but the lesson is the same: the disk is rarely the weakest link per unit.

Background on drive types is in the drive technology index. Warning signs before a failure are also covered in the SMART failure signals analysis.

Sources

  1. 01An Analysis of Error Behavior in a Large Storage System, Talagala and Patterson, UC Berkeley technical report UCB/CSD-99-1042, February 1999 (author PDF, 18 pages) · accessed 2026-10-11
  2. 02Disk failures in the real world: What does an MTTF of 1,000,000 hours mean to you?, Schroeder and Gibson, FAST 2007 (USENIX PDF, section 4) · accessed 2026-10-11
  3. 03Availability in Globally Distributed Storage Systems, Ford et al., USENIX OSDI 2010 (USENIX PDF, section 3) · accessed 2026-10-11