Facebook rolls out storage system to wrangle massive photo stores
Homegrown system, Facebook Haystack, built to handle multiplying files of user photos
Computerworld - Needing to better deal with 50 billion files worth of photos, engineers at Facebook are installing a new photo storage system they say is 50% faster than traditional systems.
The storage system, dubbed Haystack, has been under development in-house for the past couple of years, and Facebook has been rolling it out in limited test versions to parts of the network for the past few months. The company expects to use Haystack to store all Facebook photos by next week, according to Bobby Johnson, director of engineering at Facebook.
And Jonathan Heiliger, vice president of technical operations, told Computerworld today that based on tests, Haystack is more than 50% faster than traditional photo storage systems.
"In terms of cost, if it's twice as efficient, we can have 50% less hardware," said Johnson. "With 50 billion files on disk, the cost adds up. It's essentially giving us some [financial] headroom."
Johnson and Heiliger said they began building the new storage system to better handle the growing number of photos Facebook has to store. Many of their 175 million and 200 million users share photos of everything from their pets to vacations, weddings and days at the beach. That means users are posting and calling up their own photos, as well as their friends' and family members' photos. Keeping the system running efficiently was a growing challenge.
Johnson noted that Facebook deals with 15 billion photos - not including all of the replications. User data grows by 500GB per day. And Facebook has 50 million requests per second to its back-end servers.
A spokesman for Facebook said more specifics about the new system will be released in a few weeks.
Johnson, though, said the system is so much faster than the previous one because of changes made to its setup. Haystack is tailored for small files that don't change very often, instead of for a small number of large files that are changing all the time. Traditional file directories also need file names, and a lot of resource cost goes to just finding the files. The new system uses ID numbers instead of names; that mapping is very small and doesn't involve directory structures or file names.
Johnson said that so far, the rollout of the new system has gone very smoothly.
Five-year-old Facebook's user base passed one-time leader MySpace last year, according to a recent report.
Facebook, once regarded as the up-and-coming social network, had almost 222 million unique visitors last month, while MySpace came in at 125 million, according to online researcher comScore Inc. That's a dramatic change, since the Facebook-MySpace race for unique visitors was a near dead heat in April 2008.
The company is closing in on a big milestone -- 200 million users, executives said today.
Read more about Web 2.0 and Web Apps in Computerworld's Web 2.0 and Web Apps Topic Center.



- Excel 2010 Cheat Sheet
- Register for this Computerworld Insider Cheat Sheet and gain access to hundreds of premium content articles, guides, product reviews and more.
- Why Business Ethernet Services?
- Everybody's heard the cliché, "the network is your business." But that's not going to help you choose the best wide area networking service...
- Overcome Top 7 Admin Challenges of Active Directory
- As Active Directory's role in the enterprise has drastically increased, so has the need to secure the data. Gain insight on creating repeatable,...
- Insiders Can Ruin Your Company. Take Action.
- Did you know that 80 percent of threats to an organization come from the inside? The threat from insiders is often overlooked in...
- Top Solutions and Tools to Prevent Devastating Malware
- Custom malware frequently goes undetected. According to Forrester Research, the best way to reduce risk of breach is to deploy file integrity monitoring...
- Streamline Compliance and Increase ROI
- Streamline, simplify, and automate compliance related activities; especially those that impact multiple business units. This white paper from NetIQ, outlines solutions that will... All Web 2.0 and Web Apps White Papers
- Optimizing Networks for the Cloud
- Join guest speaker, Rohit Mehra, IDC Director of Enterprise Communications Infrastructure, to explore current trends, discuss best practices for optimizing Data Center and...
- Apps QuickStart Series Part 2: Designing and Deploying SQL Server on VMware vSphere
- Download this webcast to learn about the design considerations for virtualizing SQL workloads, performance and scalability information and high-availability options, as well as...
- Apps QuickStart Series Part 1: Designing and Deploying Exchange 2010 on VMware vSphere
- Download this webcast to learn the virtual hardware design considerations for Exchange 2010, deployment using the building block approach, options for high-availability and...
- Customer Spotlight: How IPC The Hospitalist Company Implemented Oracle on VMware
- Have you been looking to hear about customer's experiences with the new VMware vCenter Site Recovery Manager product? View this webcast to learn...
- Virtualize Business-Critical Applications with Confidence
- Virtualizing business-critical applications has become a key focus for organizations as they move along their virtualization journey. With the launch of VMware vSphere®... All Web 2.0 and Web Apps Webcasts