Skip to content

Stat analysis

What do public storage-incident postmortems show about drives that fail together?

Published 2026-10-11

Four public postmortems show drives failing together or in bulk because they shared something (a controller, a lot, a firmware), and the reports separate what the logs showed from what the operator only suspects. InfiniCLOUD (October 2026): 6 of 11 disks in a ZFS RAIDZ3 pool sat on one controller; one disk failed at 01:43 JST on 1 October, and a fourth had failed by 19:21 JST on 2 October, 41 h 38 min later, beyond the 3-disk tolerance; the cause of the last three failures is still unknown. Echelon VPS (February 2026): both NVMe drives of one mirror, same manufacturing lot and nearly equal write volume, hit a firmware assertion 4 min apart and took 218 instances offline for 94 min. GitHub (January 2016): a 2 h 6 min outage after a power blip, when over 25% of servers could not see their own drives after rebooting, a known firmware issue. Seagate (January 2009): a firmware event-log bug hung drives at power-up when the log counter was at entry 320 or 320 + n x 256, and the first fix was withdrawn after it disabled some 500 GB drives. None of the four reports publishes the raw logs, and three of the four name no drive or firmware version.

Question
What do public storage-incident reports say about how disks failed together, and which signals showed it first?
Incidents
InfiniCLOUD zeze node (ZFS RAIDZ3, October 2026); Echelon VPS FRA6 cluster 12 (NVMe RAID-10, February 2026); GitHub primary data centre (January 2016); Seagate Barracuda 7200.11 family (January 2009).
Failure layers
Disk internal fault plus shared controller; same-lot NVMe firmware assertion; drives not detected after power cycle; firmware hang at power-up.
Not covered
Rates of such events, any ranking of causes, and any recovery of data from affected drives.

InfiniCLOUD zeze node: first symptom to unreadable pool

Clock times are JST as published. Elapsed minutes (min) and hours (h) are derived by subtraction from 01:43 on 1 October 2026, the first logged symptom.
Time (JST)Elapsed since first symptomEvent as publishedSource
1 Oct 01:430 minA disk on the controller (HBA) stopped responding to commandsInfiniCLOUD report, 2026
1 Oct 01:452 minOne disk failed and was disconnected; pool kept running on 10 of 11 disksInfiniCLOUD report, 2026
1 Oct 01:5310 minMonitoring flagged an anomaly; no data read errors yet, so service was judged unaffectedInfiniCLOUD report, 2026
1 Oct 06:364 h 53 minTroubleshooting found system files unreadable; 61 files unreadable by end of day, all 10 remaining disks shown DEGRADEDInfiniCLOUD report, 2026
2 Oct 19:2141 h 38 minFourth disk unavailable; pool state UNAVAIL (3-disk RAIDZ3 tolerance exceeded)InfiniCLOUD report, 2026
2 Oct 21:0343 h 20 minData recovery judged virtually impossible; plan switched to assuming total lossInfiniCLOUD report, 2026

Echelon VPS FRA6 cluster 12: mirror pair lost in four minutes

Clock times are UTC on 11 February 2026 as published. Elapsed minutes (min) are derived from 04:12, the first drive dropping.
Time (UTC)Elapsed since first dropEvent as publishedSource
04:120 minFirst drive stopped answering NVMe admin commands; array degradedEchelon VPS post-mortem, 2026
04:164 minMirror partner did the same; array offline; 218 instances lost their diskEchelon VPS post-mortem, 2026
04:3119 minCause narrowed to firmware, not backplane; restore from cluster replica chosenEchelon VPS post-mortem, 2026
05:4694 minAll 218 instances running; data current to 04:11Echelon VPS post-mortem, 2026

Four incidents side by side: what was shared, what the report leaves unknown

Counts and durations are as published unless marked derived. Percent is share of servers; min is minutes; h is hours.
IncidentShared componentScale and durationLeft unknown or unpublishedSource
InfiniCLOUD, Oct 2026One HBA carrying 6 of 11 pool disks; controller on older firmware4 disks failed in 41 h 38 min (derived); 61 files unreadable on day 1Why the 3 disks failed on 2 Oct (controller or disks); firmware flaw only under investigation; no cable errors were loggedInfiniCLOUD report, 2026
Echelon VPS, Feb 2026Same manufacturing lot, same firmware, near-equal write volume2 drives 4 min apart; 218 instances; 94 min (derived from timeline)Drive vendor, model, firmware revision and the assertion text are not publishedEchelon VPS post-mortem, 2026
GitHub, Jan 2016Same hardware class; firmware issue after power cycleOver 25% of servers rebooted; outage 2 h 6 minVendor, model, firmware version and number of drives not publishedGitHub incident report, 2016
Seagate 7200.11 family, Jan 2009Firmware event-log pointer bug plus a factory test fill patternHang when log counter at 320 or 320 + n x 256 at power-up; drives built through Dec 2008; builds from 12 Jan 2009 unaffectedShare of drives affected not given ("some percentage"); press says hundreds of 500 GB drives hit by the withdrawn SD1A updateSeagate field update, 2009
Seagate fix, 16 to 21 Jan 2009Firmware update SD1AUpdate released 17 Jan per The Register; withdrawn from the Knowledge Base 19 Jan for validationSeagate statement not published at press time; count of bricked drives is a forum-based estimateThe Register, 2009

Which signal each report says showed the problem

Only signals named in the reports are listed. A dash means the report does not say.
IncidentSignal named in reportWhat the report does not saySource
InfiniCLOUDDisk returned internal-failure responses to every command; controller processing delays logged just before; ZFS pool state DEGRADED then SUSPENDED then UNAVAIL; monitoring alert 10 min after first symptomWhether SMART or any health counter warned earlier; the raw ZFS or kernel log linesInfiniCLOUD report, 2026
Echelon VPSDrive stopped answering NVMe admin commands; array degraded alert firedWhether the NVMe health log showed anything before; the log line of the firmware assertionEchelon VPS post-mortem, 2026
GitHubRemote console screenshots showed boot failures because physical drives were no longer recognised; uptime comparison showed which servers rebootedAny drive-side log; which firmware versionGitHub incident report, 2016
SeagateDrive hangs at power-up and the host BIOS no longer detects it; drive enters failsafe after an assert failureAny SMART attribute that precedes it; the notice says the condition persists through power cyclesSeagate field update, 2009

Reading the numbers

The clearest pattern is that redundancy was counted in disks but the failure domain was bigger than a disk. InfiniCLOUD's report says six of its eleven RAIDZ3 disks hung off one controller, so a controller problem could take out more disks than three-way parity tolerates. Echelon's report says its mirror halves came from the same lot with almost the same write volume and hit the same firmware assertion four minutes apart. GitHub's servers shared a hardware class and a firmware issue, and clusters lost all members although the members were spread over different racks. In each case the reports describe a shared cause, not a coincidence of independent faults.

The reports are also clear about the limits of early signals. InfiniCLOUD's monitoring raised an alert 10 minutes after the first symptom, but no data read errors were logged then and the service was judged unaffected; the unreadable system files were found 4 h 53 min after the first symptom. ZFS recorded errors against all ten remaining disks as DEGRADED because it could not tell which disk held the data it could not rebuild, which the report says does not mean ten disks had failed. A pool state is therefore a statement about the pool, not a per-disk diagnosis.

Firmware shows up as a trigger or a spreader in all four, but the evidence differs. Seagate's notice gives a precise mechanism (event-log entry 320 or 320 + n x 256 at power-up, plus a factory test fill pattern) and a date after which new drives were not affected (12 January 2009). GitHub names only 'a known firmware issue' and a cold power-drain that made the disks visible again. Echelon says its affected firmware revision was replaced fleet-wide within eleven days on 214 racks. InfiniCLOUD says it is still investigating whether an older controller firmware's error recovery, such as aborting commands or resetting a disk, spread the problem. Only Seagate's mechanism is documented to the level of a counter value.

What remains unknown is as important as the timelines. No report publishes raw SMART, NVMe health log, SCSI sense or kernel log lines, so this page cannot say which public counter would have warned first. Three of the four name no drive or firmware version. InfiniCLOUD's own report says the cause of the three disk failures on 2 October is still uncertain and could lie in the controller or in the disks. Seagate's fix itself failed: The Register reports that update SD1A was withdrawn after it disabled some 500 GB drives, which is a reminder that a firmware remedy is a change with its own failure modes.

Related: SATA and SAS link failures, SSD firmware failures at fixed power-on hours, what fails most often on HDDs, SSDs and links and the SMART failure-signal note.

Method

Everything here was read on 2026-10-11 from pages and documents opened that day: InfiniCLOUD's incident report of 8 October 2026; Echelon VPS's post-mortem of 13 February 2026; GitHub's January 28th incident report (3 February 2016); Seagate's Urgent Field Update of 16 January 2009 (PDF, copy hosted by National Instruments); and The Register's article of 21 January 2009. Times and counts are copied from those texts. Elapsed times in the tables are subtractions of the published clock times, with the unit shown, and are marked as derived. Where a report states a cause as a belief rather than a confirmed log fact, the page says so.

Limits

Each of InfiniCLOUD, Echelon and GitHub is the operator's own account, and no second independent source on those three incidents was found, so their timelines are not cross-checked; the Seagate case has the vendor's notice and a trade-press report. None of the four publishes raw logs, and the Echelon, GitHub and InfiniCLOUD reports do not name the drive vendor, model or firmware version (InfiniCLOUD says only that the controller ran an older firmware with a newer one available). InfiniCLOUD's investigation was still open on 8 October 2026 and its cause statements are marked as beliefs there; it also lists the incident as ongoing. Echelon's figures (218 instances, 94 minutes, 214 racks) are self-reported, and the post does not say how the 'same lot' fact was established. GitHub's report says 'a known firmware issue' without a vendor or count of drives. The Register's 'hundreds of drives' for the withdrawn SD1A update is a press estimate drawn from a support-forum thread; Seagate's notice gives no number of affected drives and says only some percentage were susceptible. Four incidents are a small, self-selected sample of published reports and say nothing about how common such events are.

Sources

  1. 01InfiniCLOUD: Incident Report: Data loss due to storage failure on the zeze node (posted 8 October 2026) · accessed 2026-10-11
  2. 02Echelon VPS: Post-mortem: two NVMe drives, one mirror, four minutes (13 February 2026) · accessed 2026-10-11
  3. 03GitHub Blog: January 28th Incident Report (3 February 2016, updated 4 January 2019) · accessed 2026-10-11
  4. 04Seagate Urgent Field Update: Drive Hang after Power Cycle (dated 16 January 2009; copy hosted by National Instruments, PDF) · accessed 2026-10-11
  5. 05The Register: Seagate firmware fix bricks Barracudas (21 January 2009) · accessed 2026-10-11