Research notes / Reliability methods
AFR is a rate, not a one-year failure probability
10 October 2026 / Methods review and synthetic worked examples
A drive failure rate needs a clock and a denominator. An annualized rate is not automatically the probability that your drive fails next year, and zero observed failures is not evidence of zero risk.
Start with drive-time, not the final headcount
For the exposure convention used in Backblaze's Drive Stats, count observed failures and divide by the accumulated drive-days, then scale to a year. Drives can enter and leave a fleet during observation; its final headcount cannot represent all that exposure. [1]
AFR (%) = failures / drive_days × 365 × 100
Here a year is explicitly 365 days. This yields failures per 100 drive-years. As an arithmetic example, one failure in 100 drive-days gives 365%. That is a sparse exposure estimate, not a claim that 365% of a fixed set of drives can fail once.
A probability conversion needs an additional model
For a constant individual hazard of lambda per year, exponential survival gives a one-year probability of 1 - exp(-lambda). If an AFR is used as that hazard estimate, lambda is AFR divided by 100. This assumption is separate from computing the rate. [2]
- A rate of 1% per year corresponds to 0.995% one-year probability under that model.
- A rate of 10% per year corresponds to 9.516% one-year probability under that model.
- A rate of 100% per year corresponds to 63.212% one-year probability under that model.
The approximation is close at small rates and diverges as rates rise. A pooled rate is not automatically an individual hazard: population mix, age and operating conditions still need examination.
Zero events: the exposure still carries information
For a Poisson count with zero events, a one-sided 95% upper bound is -ln(0.05), about 2.996 expected events. The upper endpoint of the equal-tailed two-sided 95% interval uses -ln(0.025), about 3.689. These are different conventions. Scale count bounds by 36500 / drive_days to express them as AFR percentages. [3]
| Example | Failures | Drive-days | AFR (%) | 95% interval (%) |
|---|---|---|---|---|
| Zero events, short exposure | 0 | 365 | 0 | 0 to 368.89 |
| Zero events, 100 times more exposure | 0 | 36,500 | 0 | 0 to 3.69 |
| One event, sparse exposure | 1 | 100 | 365 | 9.24 to 2,033.65 |
| Large-count numerical regression | 1,000 | 36,500,000 | 1 | 0.94 to 1.06 |
The two zero-event examples differ only in exposure. With 100 times more observed drive-time, the upper rate bound becomes 100 times smaller. Neither row establishes reliability outside its assumed observation process.
The overlooked distinction: a cluster has a different clock
Schroeder and Gibson's FAST 2007 study examined replacement records covering more than 100,000 disks. Its sections 2.4 and 5.3 explicitly distinguish time between replacements anywhere in a cluster from the lifetime of an individual disk. The HPC1 analysis used detection timestamps, which the other datasets did not supply as precisely. [4]
The paper found decreasing hazard for cluster replacement intervals alongside increasing replacement rates with device age. These are not contradictory results: the random variable and the clock differ. A longer quiet interval across a cluster does not demonstrate that aging makes each drive safer. Replacement logs also reflect operational decisions, not a uniform physical failure test.
This is a historical field result, not a fitted model for today's Backblaze fleet. It motivates checking dependence and age effects before using a stationary, independent Poisson model. Our earlier MTTF and field-replacement note covers the broader study.
What we checked in Hesela's own calculations
Three inherited interval generators used a recurrence initialized with exp(-lambda). A regression input of 1,000 events produced a count interval near 742.94 to 745.13: floating-point underflow had broken the calculation. We replaced the duplicated code with validated chi-square quantiles; the result is approximately 938.97 to 1,063.95.
We then checked all 132 existing count/exposure/interval records in three published tables against the corrected calculation, including their rounding. They passed: this audit found a large-count implementation defect, not a reason to alter those published numbers. Rechecking arithmetic does not revalidate the original extraction or prove that its statistical model fits.
Zero, one and large counts, invalid inputs and unit conversions are covered by regression tests. The runtime versions and source hashes accompany the machine-readable worked examples.
Reproduce and interpret
Download the example generator, calculation module and pinned requirements into one directory. In an isolated Python environment, install those requirements and run the generator. It uses no device access, random sampling or external data downloads.
For an empirical result, report the event rule, cohort, observation dates, missing records and exposure convention. Separate a confidence interval conditional on a model from uncertainty about whether that model applies. Do not turn overlapping or non-overlapping intervals into an automatic significance test.
Corpus: AFR, drive-day, hazard rate and MTBF. Existing tables: model intervals, age exposure and older-drive model composition.
Primary sources and limits
- Backblaze: 10 Stories From 10 Years of Drive Stats Data, sections 3-4
- NIST e-Handbook: Exponential distribution (8.1.6.1)
- NIST e-Handbook: Constant repair rate model and confidence bounds (8.4.5.1)
- Schroeder and Gibson: Disk failures in the real world, FAST 2007, sections 2.4 and 5.3
- NIST e-Handbook: Failure (or hazard) rate (8.1.2.3)
Sources accessed 10 October 2026. No raw archive extraction, fleet experiment or empirical paper replication was performed in this work. All new worked-example inputs are synthetic.