Vendor disk failure rates: Myth or metric?
Disk problems contribute to 20% to 55% of storage subsystem failures
April 4, 2008 12:00 PM ETComputerworld - The statistics of mean time between failures (MTBF) and average failure rate (AFR) have gotten lots of attention lately in the storage world, especially with the release of three much-discussed studies devoted to the topic in the last year. And for good reason: Vendor-stated MTBFs have risen into the 1 million-to-1.5 million-hour range, equaling 114 to 170 years, a lifespan that no one is seeing in the real world.
Three studies over the past year on MTBF include the following:
- Google Inc.'s "Failure Trends in a Large Disk Drive Population"
- Carnegie Mellon University's "Disk Failures in the Real World"
- University of Illinois' "Are Disks the Dominant Contributor for Storage Failures?"
Indeed, "how do these numbers help a person who wants to evaluate drives?" says Steve Smith, a former EMC Corp. employee and an independent management consultant in Bellevue, Wash. "I don't think they can.
Even storage system maker NetApp Inc. acknowledges in a response to an open letter on the StorageMojo blog that failure rates are several times higher than reported. "Most experienced storage array customers have learned to equate the accuracy of quoted drive-failure specs to the miles-per-gallon estimates reported by car manufacturers," the company says. "It's a classic case of 'Your mileage may vary' -- and often will -- if you deploy these disks in anything but the mildest of evaluation/demo lab environments."
Study results
The upshot of the recent studies can be summarized this way: Users and vendors live in very different worlds when it comes to disk reliability and failure rates.Consider that MTBF is a figure that's reached through stress-testing and statistical extrapolation, Harris says. "When the vendor specs a 300,000-hour MTBF -- which is common for consumer-level SATA drives -- they're saying that for a large population of drives, half will fail in the first 300,000 hours of operation," he says on his blog. "MTBF, therefore, says nothing about how long any particular drive will last." In other words, MTBF does a very poor job communicating what the actual failure profile looks like, he says.
It's like providing the average woman's height in the U.S. but without showing the numbers used to derive that average, Smith says. "MTBF became the standard because it was perceived as a simpler answer to the question of reliability than showing the data of how they arrived at it," Smith says. "It's an honest-to-God simplification."
Mean time between failure
Additional Resources



White Papers & Webcasts
Why Email Must Operate 24/7 and How to Make This Happen
Learn how to avoid an email outage by implementing a hosted email continuity solution.
Insight from an Auditor: Ensuring a Successful PCI Audit
Ensure a successful PCI audit. Watch this webcast now.
Preventing Data Loss When Migrating to Microsoft 2007
Download this new white paper today!
Beyond Basic Back-Up: Disaster Recovery
It's not always a flood or fire- 50% of "disasters" are caused by users. Learn more now!
Serving Up Faster Registration
Download this Case Study now!
Disaster Recovery 2008: Reduced Costs and Improved Performance
How long can your Enterprise afford to be without your data? With an accelerated disaster recovery program, you never have to answer this...
Technical Guide: Using a Comprehensive Virtualization Solution to Maintain Business Continuity
Learn how virtualization reduces operational risk.
HP StorageWorks EVA4400 & Microsoft
Download this video, free, compliments of HP.
Virtual Workforce: The Key to Expanding The Business While Cutting Costs
How to cut costs while growing your business. Learn more now!
Data Protection and Disaster Recovery with iSCSI and VMware
Get this on demand webcast now
Computerworld Reports
Business Continuity ZoneAn organization's business continuity plan helps keep critical functions running during an emergencythe power fails, a virus is unleashed on your network, a natural disaster has occurred. Even the slightest downtime or loss of data can cripple your operation. CDW can help you prevent disaster by implementing a well-planned recovery strategy. Click here to visit the Zone See All Zones
|




Forrester Analyst Report: X86 Server Virtualization For High Availability and Disaster Recovery
Yankee Group. "Disaster Strikes! Is Your Business Ready? Disaster Preparedness for Mid-Sized Firms"
VMware White Paper: Transforming Disaster Recovery - VMware Infrastructure for rapid, reliable and cost-effective Disaster Recovery