SSD firmware bugs that fail drives at a fixed power-on-hour count: what the public advisories show and what could warn you
Published 2026-10-10
A firmware bug tied to power-on hours looks, in every public signal the vendors describe, like a healthy drive until the hour counter reaches a fixed value, and then like a dead one: HPE says SAS SSDs with firmware older than HPD8 fail at 32,768 power-on hours (3 years 270 days 8 hours) and that neither drive nor data can be recovered afterwards, and HPE, Dell and Cisco say a second defect kills drives at 40,000 hours (HPE: 4 years 206 days 16 hours). The only advance warning the vendors name is the age counter itself, read as power-on hours with smartctl or sg_logs, together with the firmware revision; Cisco states the 40,000-hour defect does not touch the wear specification, so wear indicators give no hint. Because the trigger is accumulated runtime, HPE warns that drives deployed at the same time will likely fail nearly simultaneously, which is how a redundant array can lose more drives than its RAID level tolerates. A separate Cisco advisory describes a third defect at 28,224 hours (about 3.2 years) that makes a drive unresponsive until a power cycle, then recurs after roughly six weeks. These are vendor advisories, not field postmortems: none of the opened sources gives a count of failed drives or a log from a real incident.
- Failure trigger
- Accumulated power-on hours reaching a fixed value in firmware: 32,768 h, 40,000 h, or 28,224 h (three advisories).
- Visible symptom
- Drive reports 0 GB and goes offline (Dell, Cisco, 40,000 h) or becomes unresponsive until power-cycled (Cisco, 28,224 h).
- Warning available
- Only the age counter and firmware revision; Cisco states the 40,000 h defect does not affect the wear specification.
- Not covered
- Failure counts, other vendors' drives, consumer SSDs, and the root cause inside the supplier's firmware.
Public timeline of the advisories
| Date (YYYY-MM-DD) | Event | Who | Cross-check (second source) | Source |
|---|---|---|---|---|
| 2019-11-22 | Original release of HPE bulletin for the 32,768-hour defect; HPD8 firmware released that day for the eight late-2015 models | HPE | Broadcom KB lists the HPE advisories | HPE bulletin a00092491 |
| 2019-12-09 | HPD8 released for the remaining affected models (2017 and 2018 launch series) | HPE | Same bulletin, revision 3 entry | HPE bulletin a00092491 |
| 2020-03-26 | iTWire reports HPE advisory for four models failing at 40,000 hours, fixed by HPD7; says the issue is unrelated to the November 2019 one | iTWire quoting HPE | Dell article (same hour count) | iTWire, 26 Mar 2020 |
| 2020-04-23 | Cisco FN70545 initial release for SSDs failing at 40,000 hours (fixed firmware C405) | Cisco | HPE and Dell: same 40,000 h figure | Cisco FN70545 |
| 2020-10-23 | HPE reissues the 32,768-hour bulletin with no content change, to make sure customers act | HPE | - | HPE bulletin a00092491 |
| 2020-11 | Cisco says it learned the scope of customer impact of a separate 28,224-hour issue | Cisco | Not listed under FN70545 | Cisco SSD FAQ |
The trigger values and what they convert to
| Quantity | Unit | Value as printed | Computed | Fixed firmware | Source |
|---|---|---|---|---|---|
| HPE defect 1: failure point | power-on hours | 32,768 (3 years 270 days 8 hours) | 1,365.3 days; 3.74 years | HPD8 | HPE bulletin a00092491 |
| HPE, Dell and Cisco defect 2: failure point | power-on hours | 40,000 (HPE: 4 years 206 days 16 hours; Cisco: 4.5 years) | 1,666.7 days; 4.57 years | HPD7 (HPE), D417 (Dell), C405 (Cisco) | iTWire, 26 Mar 2020 |
| Cisco separate defect: failure point | power-on hours | 28,224 (about 3.2 years) | 1,176 days; 3.22 years | update per product field notice | Cisco SSD FAQ |
| Cisco separate defect: recurrence after power cycle | weeks | approximately 6 | about 1,008 hours | - | Cisco SSD FAQ |
What each public signal shows, before and after
| Signal | Before the trigger | After the trigger | Cross-check (second source) | Source |
|---|---|---|---|---|
| Power-on hours (smartctl selftest-log 'Lifetime' hours; sg_logs page 0x15 'Accumulated power on minutes', unit: minutes) | Readable; the only counter the vendors name for judging how close a drive is | Not needed: drive has failed | HPE: Smart Storage Administrator, or a support call, gives hours | Cisco FN70545 |
| Firmware revision (below HPD8, HPD7, D417 or C405) | Identifies exposed drives; Cisco calls the firmware version the key identifier for its 28,224 h issue | Same | Dell and Broadcom list the same cut-off revisions | Dell KB 000177640 |
| Wear / endurance specification | Cisco: the 40,000 h defect does not impact the wear specification | Unaffected; drive still fails | - | Cisco FN70545 |
| Capacity reported by the drive | Normal | 0 GB available, drive goes offline (Dell, Cisco); Dell adds a format-corruption state reported as 03/31/00 after a power cycle | Cisco FN70545: same 0 GB symptom | Dell KB 000177640 |
| Array state (RAID, disk groups) | Healthy | Disk groups going offline (Broadcom); data loss if more drives fail than the RAID level tolerates (HPE) | Dell: update only with the array optimal | Broadcom KB 329029 |
Reading the numbers
Start with the counter, not the health indicators. The three advisories share one shape: nothing degrades gradually, and a fixed number of accumulated hours is the trigger. HPE prints 32,768 hours (3 years 270 days 8 hours, which is 1,365.3 days) and, for a second defect, 40,000 hours; Cisco prints 28,224 hours for a third. Since the failure point is a number, the useful public check is arithmetic: read the power-on hours and the firmware revision and compare them with the advisory.
Why one failed drive is a warning for the rest. HPE writes that drives deployed at the same time will likely fail nearly simultaneously, and that after the failure neither the drive nor the data can be recovered. Cisco's FAQ adds a limit to that reasoning: the clock is accumulated runtime, not calendar date, so units from one shipment that were powered on later are further from the trigger. A RAID set built from one batch can therefore cross the threshold together, and HPE spells out the consequence: data loss in non-redundant RAID 0, and in redundant modes if more drives fail than the mode supports.
Health and wear indicators are the wrong instrument. Cisco states the 40,000-hour index error does not impact the wear specification. The failing check is described in one sentence: a test that the index passes N when the value can go to N+1. A drive with ample endurance left is still exposed, so a low wear figure is not evidence of safety.
What the dead drive looks like. Dell and Cisco both report 0 GB of available capacity and an offline drive; Dell adds that after a power cycle the drive is left in a format-corruption state it reports as 03/31/00. The Cisco 28,224-hour defect behaves differently: the drive stops responding, a power cycle restores it, and it fails again after about six weeks (computed: roughly 1,008 hours). That second pattern could be mistaken for an intermittent link or controller fault; the sources do not show logs, so this is an inference, and the firmware revision and hour count are the vendors' stated identifiers.
Unknown: how many drives failed, in what time window, and what a host log showed in the minutes before the first failure. The vendor documents give trigger values, fixed firmware and symptoms, not incident counts or kernel logs, and the supplier's own root-cause analysis is not public in the opened sources.
Related: what fails most often on HDDs, SSDs and links, the NVMe health-log note, the SCSI sense and kernel-log note and the SMART failure-signal note.
Method
Everything on this page was read on 2026-10-10 from documents opened that day: HPE customer bulletin a00092491 (32,768-hour defect, revision history 2019-11-22 to 2024-02-23); the iTWire report of 26 March 2020 that quotes HPE's 40,000-hour advisory; Dell support article 000177640; Cisco Field Notice FN70545; Cisco's FAQ page on a separate 28,224-hour SSD issue; and a Broadcom (VMware) knowledge-base article that lists the Dell, HPE and Cisco advisories together. Each table row names a second source in the column Cross-check where one exists. Dates and numbers are copied as printed; values marked 'computed' are derived on this page: hours divided by 24 for days, by 8,760 (365 days) for years, and 6 weeks times 168 hours per week.
Limits
No count of failed drives, no failure rate and no incident log is given by any opened source, so this page makes no claim about how many drives failed or how often. The HPE 40,000-hour bulletin itself and HPE's SAS SSD advisory page were not opened (the advisory page returned a server error when fetched); the 40,000-hour HPE statements come from an iTWire article quoting the bulletin. The supplier of the drives is not named in the vendor documents; iTWire calls SanDisk 'rumoured' and says Dell's note named it, which this page does not verify. HPE describes its 32,768-hour and 40,000-hour defects as unrelated (per iTWire), and Cisco lists its 28,224-hour issue apart from FN70545; this page does not claim a common root cause beyond Cisco's one-sentence description of the 40,000-hour index check. The Dell sense code is printed as 03/31/00; reading it as sense key, ASC and ASCQ is this page's interpretation, not Dell's wording. The Dell article shows no publication date in the opened text. The Broadcom article is a support note that repeats the vendors' advisories and adds no independent measurement. All sources are vendors that sell or support the affected drives. The page covers diagnosis from public status only; nothing here concerns encryption, drive security features, or recovering anyone else's media.
Sources
- 01HPE Customer Bulletin a00092491 (rev. 8, 2024-02-23): critical firmware upgrade to prevent SAS SSD failure at 32,768 hours · accessed 2026-10-10
- 02iTWire, 26 Mar 2020: HPE warns firmware bug will kill four SSD models after 40,000 hours · accessed 2026-10-10
- 03Dell Support KB 000177640: Enterprise SSDs LTxx00MO/RO/WM fail at 40,000 power-on hours · accessed 2026-10-10
- 04Cisco Field Notice FN70545 (v2.0, 2023-06-30): SSD will fail at 40,000 power-on hours · accessed 2026-10-10
- 05Cisco: Solid State Drive issue on certain products (3.2 years / 28,224 power-on hours), FAQ · accessed 2026-10-10
- 06Broadcom (VMware) KB 329029: SSDs experience unexpected failures at 32k/40k power-on hours · accessed 2026-10-10