What do public storage-incident postmortems show about drives that fail together?
Published 2026-10-11
Four public postmortems show drives failing together or in bulk because they shared something (a controller, a lot, a firmware), and the reports separate what the logs showed from what the operator only suspects. InfiniCLOUD (October 2026): 6 of 11 disks in a ZFS RAIDZ3 pool sat on one controller; one disk failed at 01:43 JST on 1 October, and a fourth had failed by 19:21 JST on 2 October, 41 h 38 min later, beyond the 3-disk tolerance; the cause of the last three failures is still unknown. Echelon VPS (February 2026): both NVMe drives of one mirror, same manufacturing lot and nearly equal write volume, hit a firmware assertion 4 min apart and took 218 instances offline for 94 min. GitHub (January 2016): a 2 h 6 min outage after a power blip, when over 25% of servers could not see their own drives after rebooting, a known firmware issue. Seagate (January 2009): a firmware event-log bug hung drives at power-up when the log counter was at entry 320 or 320 + n x 256, and the first fix was withdrawn after it disabled some 500 GB drives. None of the four reports publishes the raw logs, and three of the four name no drive or firmware version.
- Question
- What do public storage-incident reports say about how disks failed together, and which signals showed it first?
- Incidents
- InfiniCLOUD zeze node (ZFS RAIDZ3, October 2026); Echelon VPS FRA6 cluster 12 (NVMe RAID-10, February 2026); GitHub primary data centre (January 2016); Seagate Barracuda 7200.11 family (January 2009).
- Failure layers
- Disk internal fault plus shared controller; same-lot NVMe firmware assertion; drives not detected after power cycle; firmware hang at power-up.
- Not covered
- Rates of such events, any ranking of causes, and any recovery of data from affected drives.
InfiniCLOUD zeze node: first symptom to unreadable pool
| Time (JST) | Elapsed since first symptom | Event as published | Source |
|---|---|---|---|
| 1 Oct 01:43 | 0 min | A disk on the controller (HBA) stopped responding to commands | InfiniCLOUD report, 2026 |
| 1 Oct 01:45 | 2 min | One disk failed and was disconnected; pool kept running on 10 of 11 disks | InfiniCLOUD report, 2026 |
| 1 Oct 01:53 | 10 min | Monitoring flagged an anomaly; no data read errors yet, so service was judged unaffected | InfiniCLOUD report, 2026 |
| 1 Oct 06:36 | 4 h 53 min | Troubleshooting found system files unreadable; 61 files unreadable by end of day, all 10 remaining disks shown DEGRADED | InfiniCLOUD report, 2026 |
| 2 Oct 19:21 | 41 h 38 min | Fourth disk unavailable; pool state UNAVAIL (3-disk RAIDZ3 tolerance exceeded) | InfiniCLOUD report, 2026 |
| 2 Oct 21:03 | 43 h 20 min | Data recovery judged virtually impossible; plan switched to assuming total loss | InfiniCLOUD report, 2026 |
Echelon VPS FRA6 cluster 12: mirror pair lost in four minutes
| Time (UTC) | Elapsed since first drop | Event as published | Source |
|---|---|---|---|
| 04:12 | 0 min | First drive stopped answering NVMe admin commands; array degraded | Echelon VPS post-mortem, 2026 |
| 04:16 | 4 min | Mirror partner did the same; array offline; 218 instances lost their disk | Echelon VPS post-mortem, 2026 |
| 04:31 | 19 min | Cause narrowed to firmware, not backplane; restore from cluster replica chosen | Echelon VPS post-mortem, 2026 |
| 05:46 | 94 min | All 218 instances running; data current to 04:11 | Echelon VPS post-mortem, 2026 |
Four incidents side by side: what was shared, what the report leaves unknown
| Incident | Shared component | Scale and duration | Left unknown or unpublished | Source |
|---|---|---|---|---|
| InfiniCLOUD, Oct 2026 | One HBA carrying 6 of 11 pool disks; controller on older firmware | 4 disks failed in 41 h 38 min (derived); 61 files unreadable on day 1 | Why the 3 disks failed on 2 Oct (controller or disks); firmware flaw only under investigation; no cable errors were logged | InfiniCLOUD report, 2026 |
| Echelon VPS, Feb 2026 | Same manufacturing lot, same firmware, near-equal write volume | 2 drives 4 min apart; 218 instances; 94 min (derived from timeline) | Drive vendor, model, firmware revision and the assertion text are not published | Echelon VPS post-mortem, 2026 |
| GitHub, Jan 2016 | Same hardware class; firmware issue after power cycle | Over 25% of servers rebooted; outage 2 h 6 min | Vendor, model, firmware version and number of drives not published | GitHub incident report, 2016 |
| Seagate 7200.11 family, Jan 2009 | Firmware event-log pointer bug plus a factory test fill pattern | Hang when log counter at 320 or 320 + n x 256 at power-up; drives built through Dec 2008; builds from 12 Jan 2009 unaffected | Share of drives affected not given ("some percentage"); press says hundreds of 500 GB drives hit by the withdrawn SD1A update | Seagate field update, 2009 |
| Seagate fix, 16 to 21 Jan 2009 | Firmware update SD1A | Update released 17 Jan per The Register; withdrawn from the Knowledge Base 19 Jan for validation | Seagate statement not published at press time; count of bricked drives is a forum-based estimate | The Register, 2009 |
Which signal each report says showed the problem
| Incident | Signal named in report | What the report does not say | Source |
|---|---|---|---|
| InfiniCLOUD | Disk returned internal-failure responses to every command; controller processing delays logged just before; ZFS pool state DEGRADED then SUSPENDED then UNAVAIL; monitoring alert 10 min after first symptom | Whether SMART or any health counter warned earlier; the raw ZFS or kernel log lines | InfiniCLOUD report, 2026 |
| Echelon VPS | Drive stopped answering NVMe admin commands; array degraded alert fired | Whether the NVMe health log showed anything before; the log line of the firmware assertion | Echelon VPS post-mortem, 2026 |
| GitHub | Remote console screenshots showed boot failures because physical drives were no longer recognised; uptime comparison showed which servers rebooted | Any drive-side log; which firmware version | GitHub incident report, 2016 |
| Seagate | Drive hangs at power-up and the host BIOS no longer detects it; drive enters failsafe after an assert failure | Any SMART attribute that precedes it; the notice says the condition persists through power cycles | Seagate field update, 2009 |
Reading the numbers
The clearest pattern is that redundancy was counted in disks but the failure domain was bigger than a disk. InfiniCLOUD's report says six of its eleven RAIDZ3 disks hung off one controller, so a controller problem could take out more disks than three-way parity tolerates. Echelon's report says its mirror halves came from the same lot with almost the same write volume and hit the same firmware assertion four minutes apart. GitHub's servers shared a hardware class and a firmware issue, and clusters lost all members although the members were spread over different racks. In each case the reports describe a shared cause, not a coincidence of independent faults.
The reports are also clear about the limits of early signals. InfiniCLOUD's monitoring raised an alert 10 minutes after the first symptom, but no data read errors were logged then and the service was judged unaffected; the unreadable system files were found 4 h 53 min after the first symptom. ZFS recorded errors against all ten remaining disks as DEGRADED because it could not tell which disk held the data it could not rebuild, which the report says does not mean ten disks had failed. A pool state is therefore a statement about the pool, not a per-disk diagnosis.
Firmware shows up as a trigger or a spreader in all four, but the evidence differs. Seagate's notice gives a precise mechanism (event-log entry 320 or 320 + n x 256 at power-up, plus a factory test fill pattern) and a date after which new drives were not affected (12 January 2009). GitHub names only 'a known firmware issue' and a cold power-drain that made the disks visible again. Echelon says its affected firmware revision was replaced fleet-wide within eleven days on 214 racks. InfiniCLOUD says it is still investigating whether an older controller firmware's error recovery, such as aborting commands or resetting a disk, spread the problem. Only Seagate's mechanism is documented to the level of a counter value.
What remains unknown is as important as the timelines. No report publishes raw SMART, NVMe health log, SCSI sense or kernel log lines, so this page cannot say which public counter would have warned first. Three of the four name no drive or firmware version. InfiniCLOUD's own report says the cause of the three disk failures on 2 October is still uncertain and could lie in the controller or in the disks. Seagate's fix itself failed: The Register reports that update SD1A was withdrawn after it disabled some 500 GB drives, which is a reminder that a firmware remedy is a change with its own failure modes.
Related: SATA and SAS link failures, SSD firmware failures at fixed power-on hours, what fails most often on HDDs, SSDs and links and the SMART failure-signal note.
Method
Everything here was read on 2026-10-11 from pages and documents opened that day: InfiniCLOUD's incident report of 8 October 2026; Echelon VPS's post-mortem of 13 February 2026; GitHub's January 28th incident report (3 February 2016); Seagate's Urgent Field Update of 16 January 2009 (PDF, copy hosted by National Instruments); and The Register's article of 21 January 2009. Times and counts are copied from those texts. Elapsed times in the tables are subtractions of the published clock times, with the unit shown, and are marked as derived. Where a report states a cause as a belief rather than a confirmed log fact, the page says so.
Limits
Each of InfiniCLOUD, Echelon and GitHub is the operator's own account, and no second independent source on those three incidents was found, so their timelines are not cross-checked; the Seagate case has the vendor's notice and a trade-press report. None of the four publishes raw logs, and the Echelon, GitHub and InfiniCLOUD reports do not name the drive vendor, model or firmware version (InfiniCLOUD says only that the controller ran an older firmware with a newer one available). InfiniCLOUD's investigation was still open on 8 October 2026 and its cause statements are marked as beliefs there; it also lists the incident as ongoing. Echelon's figures (218 instances, 94 minutes, 214 racks) are self-reported, and the post does not say how the 'same lot' fact was established. GitHub's report says 'a known firmware issue' without a vendor or count of drives. The Register's 'hundreds of drives' for the withdrawn SD1A update is a press estimate drawn from a support-forum thread; Seagate's notice gives no number of affected drives and says only some percentage were susceptible. Four incidents are a small, self-selected sample of published reports and say nothing about how common such events are.
Sources
- 01InfiniCLOUD: Incident Report: Data loss due to storage failure on the zeze node (posted 8 October 2026) · accessed 2026-10-11
- 02Echelon VPS: Post-mortem: two NVMe drives, one mirror, four minutes (13 February 2026) · accessed 2026-10-11
- 03GitHub Blog: January 28th Incident Report (3 February 2016, updated 4 January 2019) · accessed 2026-10-11
- 04Seagate Urgent Field Update: Drive Hang after Power Cycle (dated 16 January 2009; copy hosted by National Instruments, PDF) · accessed 2026-10-11
- 05The Register: Seagate firmware fix bricks Barracudas (21 January 2009) · accessed 2026-10-11