Search
All content
Technology
Conventional Magnetic RecordingCMR lays down magnetic tracks side by side without overlap. Modern drives still use perpendicular magnetic recording under the hood; vendors label the product behavior “CMR” to distinguish it from shingled (SMR) layouts.
Technology
Shingled Magnetic RecordingSMR increases HDD areal density by writing new tracks partially overlapping previous ones—like roof shingles—so each pass narrows the writable track width. Reads stay random-access, but rewriting a sector often requires reading and rewriting entire shingled bands, so vendors expose SMR through drive-managed firmware or host-visible zoning (ZAC/ZBC).
Technology
Perpendicular Magnetic RecordingPMR stores each bit with its magnetization perpendicular to the disk surface. Western Digital’s November 2007 white paper dates a 2005 HGST field test, a 2006 commercial introduction, and the first terabyte desktop drive, the Deskstar 7K1000, to January 2007. The paper describes trailing-shield writers and granular CoCrPt oxide media with a soft underlayer as the stack that made that density practical.
Technology
Heat-Assisted Magnetic RecordingHAMR briefly heats a nanoscale spot on the platter with a laser-assisted writer so the recording head can flip bits in high-coercivity media that would resist conventional perpendicular writes. The heat pulse is short; vendors use it to push areal density beyond prior PMR/CMR limits on the same 3.5-inch form factor.
Technology
Nearline Hard Disk DriveSeagate’s Exos X18 datasheet describes one drive sold into this class: 7200 RPM, SATA or SAS, helium, CMR, 24×7 duty, 12–18 TB, 4.16 ms average latency, and 170/550 random IOPS, for hyperscale and scale-out stores including Hadoop and Ceph. Toshiba’s 30 March 2026 release says it is sampling 3.5-inch nearline drives at 30–34 TB with host-managed SMR, FC-MAMR, and helium, also rated 24/7, and that CMR nearline reaches 28 TB. That release names HAMR only as a later plan.
Technology
Microwave-Assisted Magnetic RecordingMAMR uses microwave energy—often from a spin-torque oscillator near the write head—to help flip bits in high-anisotropy media without the full thermal cycle of HAMR. Toshiba brands its production-oriented approach Flux Control MAMR (FC-MAMR™); research variants such as MAS-MAMR add localized microwave switching and remain on the laboratory path.
Technology
NVM ExpressNVM Express (NVMe) is a standardized register and command-set interface for PCI Express–attached non-volatile memory subsystems. It replaces the single-queue, high-latency AHCI patterns designed for spinning disks with multiple submission/completion queues, lightweight commands, and semantics tuned for flash and future NVM.
Technology
Hard Disk DriveA hard disk drive stores bits on spinning platters coated with magnetic material, read and written by heads on a moving actuator. Decades of mechanical, interface, and recording-technology layers—CMR, SMR, HAMR, helium fill, dual actuators—now sit behind the same logical block abstraction exposed over SATA, SAS, or USB bridges.
Technology
Floppy DiskThe floppy disk is a removable flexible magnetic disk sealed in a jacket, first shipped by IBM in 1971 as the read-only 8-inch 23FD diskette drive (code name Minnow) that held about 80 KB, roughly 3,000 punched cards, and was built to load microcode into IBM mainframe controllers.
Technology
GMR Spin-Valve Read HeadA GMR spin-valve head is a hard-disk read sensor built from two magnetic layers separated by a thin non-magnetic spacer, with one layer pinned by an antiferromagnet so that the other rotates in the field of a bit and changes the sensor's electrical resistance; IBM shipped the first one in the 16.8 GB Deskstar 16GP, announced in November 1997 and approved for shipment on 24 December 1997.
Technology
Removable Cartridge DiskA removable cartridge disk is a magnetic disk drive whose recording disk travels in a cartridge that the user swaps in and out, and the 1982-1999 generation split into two designs: flexible media held next to the head by an air film (Iomega Bernoulli Box, 1982; Zip, shipping March 1995) and rigid hard-disk cartridges (SyQuest SQ306, first shipped 8 August 1982; Iomega Jaz, December 1995).
Technology
Magneto-Optical Disc and MiniDiscA magneto-optical (MO) disc is a rewritable disc on which a focused laser heats a spot of a perpendicularly magnetised film until a small applied magnetic field can reverse it, and on which the same laser later reads the bit through the magneto-optic Kerr effect; the technology was demonstrated in 1967, sold as 5.25-inch computer drives from 1988, and was also the recordable medium of the Sony MiniDisc (1992-2013).
Technology
Magnetic Drum and Core Memory (pre-RAMAC storage)Before the IBM RAMAC disk drive of 1956, computers kept programs and data on magnetic drums (a rotating cylinder with a fixed head for each track, first built in the mid-1940s and delivered in a production computer in 1950) and, from 1953, in magnetic core memory (a grid of ferrite rings giving random access to every bit), and the two technologies split a job that disk drives later took over in part.
Technology
Longitudinal to Perpendicular Recording Transition (1975-2010)Hard-drive makers moved from longitudinal to perpendicular recording between December 2004 (Toshiba's first announcement) and about 2010 because thinner longitudinal media were running into a thermal-stability limit, a change that rested on Shun-ichi Iwasaki's work at Tohoku University from the mid-1970s.
Technology
PRML Read Channel (Partial-Response Maximum-Likelihood, 1970-2000)PRML is a way of reading a hard drive that, instead of finding each magnetic peak on its own, equalizes the weak readback signal to a known partial-response shape and uses the Viterbi algorithm to pick the most likely bit sequence; IBM shipped the first disk drive with it, the 0681, in 1990.
Technology
Magnetic Bubble Memory (1967-1986)Magnetic bubble memory stored each bit as the presence or absence of a tiny magnetic domain (a bubble) in a thin garnet film, moved around fixed loops by a rotating magnetic field, so it kept data without power and had no moving parts; it was invented at Bell Labs in the late 1960s, sold by Texas Instruments from 1977 and Intel from 1979, and lost the market in the early 1980s.
Technology
Compact Cassette and DAT as Data Media (1963-2009)The Philips Compact Cassette (1963) became a widely used data medium for home computers once the 1975 Kansas City standard defined how to record bits as audio tones, and the 3.81 mm tape of the cassette lineage returned in the late 1980s as Digital Audio Tape (DAT), whose rotating-head recording was adapted by HP and Sony into the Digital Data Storage (DDS) backup format of 1989.
Technology
Holographic Data Storage (1963-2012)Holographic data storage records whole pages of bits at once as optical interference patterns inside the thickness of a photosensitive medium rather than as marks on a surface, an idea analysed in theory in 1963, developed in research programmes through the 1990s, and taken to prototype drives and an Ecma standard (ECMA-377, 2007) before the companies involved stopped trading.
Technology
Optical Disc Lineage: CD, CD-ROM, DVD, Blu-ray and the HD DVD Format War (1979-2008)The mainstream optical disc lineage is a sequence of 12 cm discs in which each generation kept the disc size and cut the laser wavelength and track pitch (CD at 1982, DVD unified in 1995 and standardised in 1997, Blu-ray announced in 2002), with two format disputes along the way: SD versus MMCD before DVD, and HD DVD versus Blu-ray before the 2008 withdrawal of HD DVD.
Technology
SCSI and ATA Disk Interface History: SASI, SCSI, IDE/ATA, Serial ATA and SAS (1979-2009)Two interface families put the disk controller inside the drive and defined a standard way for a host to talk to it: SCSI, which grew out of Shugart's SASI and became ANSI X3.131-1986, and ATA (IDE), which Compaq, Western Digital and Control Data's Imprimis division developed from 1986 and ANSI standardised as X3.221 in 1994; both later moved from parallel cables to serial links, as Serial ATA (specification 1.0 in 2001) and Serial Attached SCSI (INCITS 376, approved in November 2003).
Analysis
3D NAND retention: layer count is not a reliability ratingSeparate cell architecture, early retention loss, neighbor-state interference and wear before turning NAND chip measurements into SSD reliability claims.
Analysis
AFR is a rate, not a one-year failure probabilityWhy exposure, zero-event uncertainty and the choice of clock matter when interpreting drive failure rates, with reproducible synthetic examples and a FAST 2007 methods caveat.
Analysis
Atomic rename is not a durable commitFile contents, directory entries and commit acknowledgements have different persistence boundaries. Linux documentation and two USENIX studies explain what a rename does not guarantee.
Analysis
Missing SMART ages and the limits of drive-age comparisonsA full Q4 2025 raw-data audit finds 18 TB drives in a mislabeled capacity group and shows why excluding 106 missing-age drive-days also excludes 19 failures.
Analysis
Why a valid checksum can still return the wrong dataSeparate byte integrity, block identity and freshness: lost writes, misdirected writes and parity pollution, with five reproducible synthetic examples.
Analysis
Azure disk-error prediction caught 36.5% of faulty disks at a 0.1% false-positive rate, and random cross-validation said 91.6%In Xu et al. (USENIX ATC 2018), Microsoft's disk-error predictor, CDEF, found 36.50% of faulty disks on one test set when it was allowed to wrongly flag only 0.1% of healthy disks (29.67% to 41.09% across three test sets), with about 10,000 healthy disks per 3 faulty ones. Scored with a random split, the same task gave 91.64%, which is why the authors insist on testing forward in time. Adding system-level signals such as OS events to SMART raised the average catch rate at that false-positive rate from 27.6% to 30.3%, and to 35.8% with feature selection. In production at Azure the authors report about 63,000 fewer minutes of VM downtime a month. The data are one cloud system in one month, and labels come from engineers' root-cause analysis of service problems.
Analysis
Hard drives were 82% of 290,000 datacenter hardware failure tickets, and 2% of failed servers produced over 99% of the failuresIn Wang, Zhang, and Xu (DSN 2017), over 290,000 hardware failure tickets from four years at one large Internet company, hard drives were 81.84% of all failures, against 3.06% for memory and 0.31% for SSDs. The failures were very uneven: 2% of the servers that ever failed produced over 99% of the failures, and 35 of 1,411 days (2.48%) had over 500 hard drive failures. In one case 32% of a product line's servers reported hard drive SMART alerts in one night. Schroeder and Gibson (FAST 2007) found disks were 20% to 50% of hardware replacements in three other systems, and Ford et al. (OSDI 2010) found 37% of Google node failures came in bursts, so the dominance of drives and the clustering both show up elsewhere.
Analysis
Discard is not zeroing: what TRIM, UNMAP and fstrim actually promiseA discard request, zero-filled readback and reclaimed physical space are different observations. Protocol and Linux documentation show how to interpret each without false assurances.
Analysis
Enterprise SSD replacements ran at 0.22% a year across 1.4 million NetApp drives, with firmware the biggest swingIn Maneas et al. (FAST 2020), about 1.4 million SSDs in NetApp enterprise storage systems had an average annual replacement rate (ARR) of 0.22%, with models ranging from 0.07% to nearly 1.2%. That is well below the 4% to 10% of Google's flash drives removed within four years and the 2% to 9% a year the same paper quotes for hard disks, though the definitions of replacement differ. Within the NetApp data, the earliest firmware versions of some models were replaced up to 8 to more than 10 times as often as later ones, and the chance of a second replacement in the same RAID group within a week was 9.39% against 0.0504% in a random week. One third of replacements were predictive and another third were severe SCSI errors. Method limits are on the note.
Analysis
80% of 315 verified fail-slow drives at Alibaba were caused by software scheduling, and 216 of them were in two clustersIn Perseus (Lu et al., FAST 2023), a fail-slow detector run on 248K Alibaba drives for 10 months found 304 fail-slow drives, and of 315 verified fail-slow drives in the authors' benchmark, 252 (80.0%) were caused by ill-implemented software scheduling and 63 (20.0%) by hardware, with 216 of the 252 in just two clusters of open-channel SSDs that shared one flaw. Isolating the drives cut node-level 99.99th percentile write latency by 48%. Gunawi et al. (FAST 2018), a separate set of 101 fail-slow reports from 12 institutions, found 39% of root causes were external factors such as configuration, environment, power, and temperature, and that 17% of incidents took months to detect. In both, a drive that runs slowly is often slow for a reason outside the drive.
Analysis
Network is the largest fail-slow hardware type: 29 of 112 occurrences in 101 reportsNetwork is the largest hardware type in Table 3 of Gunawi et al., FAST 2018: 29 of 112 root-cause occurrences, ahead of CPU 25, disk 23, SSD 22, and memory 13. Those 112 occurrences come from 101 fail-slow reports at 12 institutions, incidents reported from 2000 through 2017, and a report can name more than one cause. Device errors account for 40 occurrences and firmware issues for 20. Temperature, power, environment, and configuration sum to 44 of 112 (39.3%), which rounds to the proceedings’ statement that 39% of root causes are external factors, not 39% of the 101 reports.
Analysis
What fails most often: unreadable data on disks, uncorrectable reads on SSDs, and bus errors on the linkThe most frequent field fault is not a dead drive but data that cannot be read: 3.45% of 1.53 million hard disks developed at least one latent sector error in 32 months, and 20.3–90.5% of Google's flash drives (ten models, first 4 years; 3 for eMLC) had at least one uncorrectable error and 19.1–62.7% a final read error, while 4.2–34.1% of Facebook's SSD platforms reported at least one uncorrectable error. Link faults show up separately as CRC counts (about 2% of Google's disks), and the public counters miss a large share of failed drives: over 56% of failed drives in the Google study and 23.3% in Backblaze's 2016 log showed no signal on the SMART attributes each used.
Analysis
Why retrying a failed fsync does not prove durabilityA successful retry can coexist with stale bytes on storage. Two USENIX studies, Linux writeback-error semantics, and a reproducible four-case model explain the boundary between an observed error and repaired data.
Analysis
Where do hard drives fail mechanically (heads, spindle, vibration, helium), and which public signals see it?Public sources let you name four mechanical failure points of a hard drive and say which counters watch each, but none of the sources opened gives the share of real-world failures each one causes. Head-disk contact: a Berkeley simulation of a 2.5-inch drive found the head hit the disk at a 400 G, 0.5 ms positive shock but not at 300 G, and at a 2.0 ms negative shock of 1,600 G but not 1,500 G. Vibration: in a 1.8-inch drive test, throughput dropped to zero at 300 to 1,000 Hz below 15 g and recovered when the shaking stopped. Helium: an HGST paper found disk flutter rose fast as air replaced helium and levelled off near 80% air, and Backblaze saw one helium drive's SMART 22 fall to 94 to 99 with no other fault. Spindle: Seagate's documented counters are retries in the last 8 spin-ups and a count of mechanical start failures. Vendor counters (head bitmap, shock, free-fall, high-fly writes, helium pressure) exist, but the sources describe what they count, not how often they warn before a failure.
Analysis
LDPC in NAND: the read cost of confidenceHow soft information is acquired, why decoder structure changes the result, and which error-rate denominators an SSD reliability claim must disclose.
Analysis
Adding performance and location data lifted 10-day disk failure prediction to 0.95 MCC across 380,000 disksIn Lu et al. (FAST 2020), models fed only SMART attributes missed many failing disks, and adding performance metrics and rack location raised the best model, a CNN-LSTM, to 0.95 MCC and 0.95 F-measure for a 10-day prediction horizon on 380,000 hard disks from 64 sites over about two months, against 0.77 MCC for the next best method, a random forest. Location alone helped by less than 10% in MCC, and only when performance data were present. The result comes from one operator's data center and one short window, so it is a measure of what is possible there, not a general error rate. Backblaze (2016) and Google (FAST 2007) both report that a large share of failed drives show no SMART warning, which is the same gap from other fleets.
Analysis
Field disk replacements averaged 3.01% a year, 3.4 times the 0.88% a 1,000,000-hour MTTF impliesIn Schroeder and Gibson (FAST 2007), disk replacement logs from large production systems (more than 100,000 drives, up to five years of use) gave a weighted average annual replacement rate of 3.01% for drives under five years old, 3.4 times the 0.88% annual failure rate that a datasheet MTTF of 1,000,000 hours implies. Replacement rates also kept rising with age instead of settling after year one, and the time between replacements was far from the exponential model that most reliability arithmetic assumes. A replacement is not a confirmed failure, so read the 3.4 times as a gap between datasheet arithmetic and field replacement practice. Google's FAST 2007 paper reports a 1.7% to over 8.6% range, and the Backblaze Q2 2026 report gives a lifetime fleet rate of 1.41%.
Analysis
NAND paired pages: why a later write can damage earlier dataTwo NAND pages can share cells. Distinguish interrupted programming, neighbor interference and the protection promised by an SSD.
Analysis
NAND read-retry: when a successful SSD read takes longerSeparate NAND recovery from host retries, retention loss from read disturb, and chip measurements from simulated SSD performance.
Analysis
No clear SMART correlation with fail-slow NVMe SSDs: 4,584 of 779,978 drives, 10 later in the failure ticketsSMART attributes showed no clear correlation with fail-slow labels on the Alibaba NVMe SSDs in Lu et al. (USENIX ATC 2022). Under the 5-minute bar, Table 7 counts 4,584 slow drives out of 779,978 (0.59%). Table 5's 1.41% is the unweighted average of eight models, not that fleet share. The printed 6.05 times versus HDDs matches (1.41 - 0.20) / 0.20, not the quotient 1.41 / 0.20 (7.05). Only 10 of the 4,584 slow drives (around 0.22%) later appear in the failure tickets.
Analysis
What the NVMe health log can see: five warning bits, a few counters, and the failures that show up elsewhereThe NVMe SMART / Health Information log is a 512-byte snapshot of five current-state warning bits (spare low, temperature, reliability degraded, read-only, backup power failed) plus counters for media errors, unsafe shutdowns, error-log entries, temperature time and wear, and it has no field for link, boot, lost-device or slow-I/O faults. In an Alibaba fleet of over one million NVMe SSDs, 99.97% of valid critical-warning readings were zero, only 14.15% of failure tickets with a recorded symptom were drives whose SMART values crossed a threshold, and 40 groups of drives showed no clear correlation between SMART attributes and fail-slow behaviour.
Analysis
RAID write hole: why a bitmap is not a write journalA bitmap tracks dirty regions; PPL repairs a narrower parity hazard; a journal preserves replay data. A reproducible XOR example separates these guarantees.
Analysis
Replacing disks at 200 reallocated sectors removed 88% of triple-disk RAID failures at one vendorIn Ma et al. (FAST 2015), about 1 million SATA disks in backup systems from one storage vendor showed that the count of reallocated sectors (RS) correlates strongly with whole-disk failure. Replacing a disk once its count passed 200 eliminated about 88% of the RAID failures caused by three-disk errors, which were 80% of all RAID failures, or about 70% of all disk-related incidents. The same paper says the signal is incomplete: in simulation, a threshold below 200 caught 52% to 70% of impending whole-disk failures, with 0.8% to 4.5% false positives. Google's FAST 2007 data and Backblaze's 2016 SMART tally point the same way on reallocated sectors, and both also show failed drives with no warning.
Analysis
When the cable, backplane or adapter fails instead of the disk: what link faults look like in logs and field dataIn NetApp's 2008 field study of about 39,000 storage systems (about 1,800,000 disks, 44 months, January 2004 to August 2007), physical interconnect failures (host adapters, cables, shelf power and backplanes) made up 27 to 68% of storage subsystem failures, against 20 to 55% for the disks themselves, and in one class the whole subsystem failed at 4.6% a year while its disks failed at 0.9%. A link fault reaches the operator as a disk that vanishes or times out, so it is often replaced as a bad disk. The public signals that separate the two are specific: on Linux, an ICRC flag in the ATA error register and SError bits such as BadCRC point to the cable, connector or power path, a CRC count in SMART (attribute 199) counts interface errors, and a drive that is missing with no media errors points the same way. What these signals cannot do is name the failing part, because a bad cable, a bad port and a marginal power supply write the same bits; the kernel documentation says only 'often'.
Analysis
Disk scrubbing: detection time is not a durability guaranteeWhat latent sector errors, staggered scan order, Linux MD checks and ZFS scrubs actually establish, and why error-discovery timestamps complicate reliability models.
Analysis
Media error or link error? What SCSI sense codes and Linux ATA log lines can and cannot tell apartA SATA drive that fails to read a sector and a SATA cable that corrupts a transfer are separated by two bits in the ATA error register: UNC (0x40, a media error, reported only after the drive's own retries) maps to SCSI sense key 03h MEDIUM ERROR with ASC/ASCQ 11h/04h, while ICRC with ABRT (0x84) maps to sense key 0Bh ABORTED COMMAND with 47h/00h and, in Linux, to the log text 'ATA bus error' (Emask 0x10). The separation is lossy: the Linux translation table collapses every link CRC into one SCSI parity code although the T10 list has separate codes for data-phase CRC (47h/01h) and UDMA CRC (08h/03h), a bus error clears the device and media bits of the same command, and a timeout (Emask 0x4) says only that the controller got no answer. Linux changes link speed after more than 3 bus or timeout errors in 10 minutes, but media errors are not counted by any of those rules.
Analysis
Sector errors predicted a week ahead let a scrubber find them 1.7 to 1.8 times faster on hard disks, using accelerated scrubs 2% of the timeIn Mahdisoltani et al. (USENIX ATC 2017), random-forest classifiers trained on drive monitoring data predicted a week ahead whether a disk would develop sector errors. For two hard disk models from Backblaze's public data, limiting false alarms to 10% of error-free weeks caught 90% and 95% of the weeks with errors, and limiting false alarms to 2% still caught 70% to 90%. Three Google SSD models were harder: 50% to 70% at 10% false alarms. When a simulated scrubber doubled its speed after a prediction and did so for at most 2% of the time, it found errors 1.7 to 1.8 times faster on hard disks and 1.4 to 1.5 times faster on two of the SSD models. The scores are weekly interval predictions on a held-out quarter of the data, and the text I read does not say the split was by time. Bairavasundaram et al. (2007) and Pinheiro et al. (2007) confirm that sector errors are common and that scrubbing finds most of them.
Analysis
0.86% of nearline disks and 0.065% of enterprise disks developed a checksum mismatch in 41 monthsIn Bairavasundaram et al. (FAST 2008), 3,088 of about 358,000 nearline disks (0.86%) and 767 of 1.17 million enterprise-class disks (0.065%) developed at least one silent checksum mismatch over 41 months, in field logs covering 1.53 million disks. That is about 400,000 mismatched blocks in all, but the burden is lopsided: the median affected disk had 3 mismatches, the mean was 104, and the worst disk had 33,000. The paper's rates are per disk and come from one vendor's support logs, so they do not say how often a file is corrupted. An independent 2007 CERN report found 22 mismatching files among 33,700 checked, about 1 in 1,500, using a different method and a different population.
Analysis
Over 56% of failed drives had no count on the four strong SMART signalsOver 56% of failed drives in the Google production population of more than 100,000 consumer-grade serial and parallel ATA drives (December 2005–August 2006) had no count in any of the four strong SMART signals (scan errors, reallocation count, offline reallocation, and probational count), so models based only on those signals can never predict more than half of the failed drives.
Analysis
SMART zeros, missing attributes and failure-day selectionA complete Q4 2025 raw-data audit separates 66 failed drives with five measured zeros from 170 with partial zero-only readings, and explains why complete-case filtering changes the manufacturer mix.
Analysis
What does a shingled (SMR) hard drive failure look like in logs, and what does SMART miss?The public SMR failures are mostly not worn-out media but a write-cache overflow: a drive-managed SMR drive absorbs writes in a small conventional area and, once that fills, slows by one to three orders of magnitude, and in some reports it then returns errors and is dropped from the array. Measured slowdowns: 13.2 MiB/s against 209.3 MiB/s for 32 KiB writes on a 4 TB WD Red EFAX (Ars Technica, 15.9:1), nearly 230 h against under 17 h for a 4-drive RAIDZ resilver (ServeTheHome), and about 0.1 MB/s for more than 25 min in a prototype host-aware SMR drive (University of Minnesota paper). In the one failing resilver whose logs were published (OpenZFS issue 10214) SMART overall health said PASSED and the reallocated, pending, uncorrectable and CRC counters were all 0, while the ATA error log counted 7,487 errors (the 24 it still showed were all IDNF) and ZFS counted 47.7 K write errors. Only one lab (iXsystems) and some users saw the failed state; ServeTheHome and Ars, testing the same drive model, saw slowness only, and a CMR drive on a shared controller logged the same error text.
Analysis
In Alibaba's SSD fleet, 12.9% of failures were in the same node and 18.3% in the same rack within 30 minutes of another failureIn Han et al. (FAST 2021), a study of nearly one million SSDs of 11 models in Alibaba data centers over 2018 and 2019, 12.9% of the roughly 19,000 failures were intra-node failures and 18.3% were intra-rack failures, meaning another SSD in the same node or rack failed within 30 minutes. Once two SSDs in a node had failed, the chance of a further failure in that group was 26.3% to 64.3%, against a fleet annual failure rate of 1.16%. The strongest SMART attribute had a rank correlation of only 0.23 with these failures. Maneas et al. (FAST 2020, NetApp) found a drive replacement was 180 times more likely in the week after another in the same RAID group, and Ford et al. (OSDI 2010, Google) found 37% of node failures in bursts, so failures that cluster in space and time appear in all three fleets.
Analysis
38% of failed datacenter SSDs showed none of the four SMART symptoms; the ones that did failed 3 to 20 times as oftenIn Narayanan et al. (SYSTOR 2016), a study of over half a million SSDs in Microsoft datacenters over nearly 3 years, 62% of devices that fail-stopped had shown at least one of four SMART symptoms (uncorrectable or CRC data errors, sector reallocations, program or erase failures, and SATA downshifts) and 38% had shown none. Devices with a symptom failed 3 to 20 times as often as devices without one, with data errors the strongest at up to 20 times. Of the models whose failure rate could be compared with vendor figures, two consumer models exceeded the published 0.61% to 0.73% annual failure rate. A multi-factor classifier reached 87% precision and 71% recall, but it was scored with 5-fold cross-validation on one operator's data. Maneas et al. (FAST 2020) and Backblaze (2016) show the same pattern in other fleets: symptoms raise risk and a large share of failures give no warning.
Analysis
What an unexpected power cut does to SSDs, and why power-loss protection is hard to check from outsideHistorical power-fault tests exposed SSD data and ordering failures, including on devices labelled with PLP. Those outcomes do not identify a failed capacitor or establish a current-product failure rate.
Analysis
SSD firmware bugs that fail drives at a fixed power-on-hour count: what the public advisories show and what could warn youA firmware bug tied to power-on hours looks, in every public signal the vendors describe, like a healthy drive until the hour counter reaches a fixed value, and then like a dead one: HPE says SAS SSDs with firmware older than HPD8 fail at 32,768 power-on hours (3 years 270 days 8 hours) and that neither drive nor data can be recovered afterwards, and HPE, Dell and Cisco say a second defect kills drives at 40,000 hours (HPE: 4 years 206 days 16 hours). The only advance warning the vendors name is the age counter itself, read as power-on hours with smartctl or sg_logs, together with the firmware revision; Cisco states the 40,000-hour defect does not touch the wear specification, so wear indicators give no hint. Because the trigger is accumulated runtime, HPE warns that drives deployed at the same time will likely fail nearly simultaneously, which is how a redundant array can lose more drives than its RAID level tolerates. A separate Cisco advisory describes a third defect at 28,224 hours (about 3.2 years) that makes a drive unresponsive until a power cycle, then recurs after roughly six weeks. These are vendor advisories, not field postmortems: none of the opened sources gives a count of failed drives or a log from a real incident.
Analysis
34.4% of RASR failures are not the SSD: 20.5% human mistakes, 13.9% cables, 31.2% labeled Failed DeviceIn Xu et al., USENIX ATC 2019, 34.4% of Reported As “SSD-Related” (RASR) failures are not caused by the SSD device. That 34.4% is human mistakes (20.5%) plus faulty cables (13.9%), inferred from which repair worked. It is not the complement of a measured device failure: Table 11 labels 31.2% Failed Device because replacing the SSD is the last resort, 11.9% Transient, and 22.5% Undetermined (FSCK 16.5% plus Data Check 6.0%). The paper prints 20.5 and 34.4. It does not print 22.5 or the addition 20.5 + 13.9.
Analysis
SSD sanitization: why TRIM and a successful command are not enoughFile deletion, deallocation and sanitization establish different things. A FAST study, NIST SP 800-88r2 and NVMe status semantics explain what evidence to collect before an SSD leaves your control.
Analysis
Do hot SSDs fail more, and what do the thermal-throttle counters actually show?In the one large public field study that measured SSD temperature (Meza et al., SIGMETRICS 2015, Facebook servers, snapshot of November 2014), heat raised failure rates only on the platforms whose SSDs rarely throttled: at 30 to 40 °C all platforms showed similar or slightly rising failure rates, and above 40 °C two platforms (A and B, almost no throttling) showed rising rates, two (C and E, aggressive throttling) were less temperature-sensitive, and two (D and F) showed falling rates that the authors link to young drives in the early-failure period; some controllers start acting from around 80 °C and in the extreme case shut the SSD down. For hard disks, Google's 2007 study of more than 100,000 drives found no consistent rise of failures with temperature at moderate temperatures. What an operator can read from outside is narrow: the NVMe health log keeps minutes above the warning and critical thresholds and, on drives that implement them, transition counts and seconds spent in thermal management, but a zero can mean either never or not implemented, and no opened source gives a failure rate for a given count.
Analysis
How do worn-out SSDs actually die, and do Percentage Used and Available Spare warn you?Worn-out SSDs did not die the same way in the one public endurance run that logged it: in The Tech Report's experiment (published 12 March 2015) six consumer SSDs wrote from about 700 TB to just over 2.4 PB before failing, four gave warning through SMART life indicators or software messages, two (Samsung 840 and 840 Pro) died without warning while their SMART reserves still looked healthy, and none of the six was usable afterwards: they bricked on a power cycle or were no longer detected. The NVMe health log offers two life gauges, Percentage Used and Available Spare, but its own definitions limit them: Percentage Used is a vendor estimate that may exceed 100 and at 100 'may not indicate an NVM subsystem failure', and at Available Spare below its threshold an alert 'may occur'. In Google's FAST 2016 field data the chance of an uncorrectable error grew roughly linearly with program-erase cycles, with no sharp rise at the rated limit, so a gauge reaching 100 is not a cliff edge.
Analysis
Google measured a disk MTTF of 10 to 50 years but a node MTTF of 4.3 months; 37% of node failures came in burstsIn Ford et al. (OSDI 2010), one year of data from tens of Google storage cells of 1000 to 7000 nodes each, the mean time to failure was 10 to 50 years for a disk, 10.2 years for a rack, and 4.3 months for a node, so node failures, not disk failures, drove availability. 37% of failures were part of a burst of at least 2 nodes. In the authors' model, a 10% cut in the disk failure rate raised stripe availability by less than 1.5%, while a 10% cut in the node failure rate raised data availability by 18%. Maneas et al. (FAST 2020, NetApp SSDs) and Schroeder and Gibson (FAST 2007, HPC1 disks) also found strong clustering of drive replacements in time, so a failure rate alone does not describe risk.
Analysis
What do public storage-incident postmortems show about drives that fail together?Four public postmortems show drives failing together or in bulk because they shared something (a controller, a lot, a firmware), and the reports separate what the logs showed from what the operator only suspects. InfiniCLOUD (October 2026): 6 of 11 disks in a ZFS RAIDZ3 pool sat on one controller; one disk failed at 01:43 JST on 1 October, and a fourth had failed by 19:21 JST on 2 October, 41 h 38 min later, beyond the 3-disk tolerance; the cause of the last three failures is still unknown. Echelon VPS (February 2026): both NVMe drives of one mirror, same manufacturing lot and nearly equal write volume, hit a firmware assertion 4 min apart and took 218 instances offline for 94 min. GitHub (January 2016): a 2 h 6 min outage after a power blip, when over 25% of servers could not see their own drives after rebooting, a known firmware issue. Seagate (January 2009): a firmware event-log bug hung drives at power-up when the log counter was at entry 320 or 320 + n x 256, and the first fix was withdrawn after it disabled some 500 GB drives. None of the four reports publishes the raw logs, and three of the four name no drive or firmware version.
Analysis
What LTO TapeAlert can tell you about a failing tape, and what it cannotTapeAlert, the 64-flag status page that LTO drives return through SCSI LOG SENSE page 2Eh, can say that an unrecoverable error happened and can point at the cartridge (flag 4, Media) or leave the question open (flags 5 and 6, where IBM's text says the cause is uncertain between tape and drive), but it cannot settle which of the two it was: IBM's action text prescribes a swap test with a known-good cartridge and drive. It also cannot report damage in data that has not been read, because the LTO consortium paper notes that uncorrectable errors stay latent until the data is read. The error budget behind the flags is small: the specified end-of-life uncorrectable bit error rate is 1 in 10^19 user bits for LTO-8 and 1 in 10^20 for LTO-9, which the paper converts to one event per 12.5 exabytes read for LTO-9, against 1 in 10^15 for the HDD reference it uses. These are calculated specifications, not field failure rates, and the same flag number means different things on a drive and on a library (flag 4 is Media on an LTO drive and Hardware Fault on a StorageTek SL150 library).
Analysis
ZNS resource limits: why closing a zone may not unblock writesOpen and active zones consume different budgets. A worked example explains ZNS write stalls, usable capacity, Zone Append ordering and the limits of published performance claims.
Timeline
Storage technology timelineMilestones in magnetic, optical and solid-state storage.