Office 365 email conks out twice within a week
Some customers in North and South America were affected
Network World - Microsoft's Office 365 service has suffered two email outages within a week of each other that affected some customers in North and South America that stemmed from different causes but ended in the same result: failed email delivery.
The first outage Nov. 8 stemmed from an overwhelmed antivirus engine and the subsequent backup that caused the service degradation. The second on Nov. 13 resulted from the failure of unspecified network elements, routine maintenance and increased load that combined to degrade service, according to the Office 365 blog posted by Rajesh Jha, the corporate vice president of Microsoft's Office division.
He didn't say how many customers were affected or where they were located other than somewhere on the two continents. Both outages affected just Office 365 Exchange Online mail services.
Affected customers are entitled to a service credit. Jha apologizes and promises a post mortem on the outages as well as an update on how the Office 365 service level agreement was affected.
The Nov. 8 incident started when an antivirus engine bogged down as it processed emails that the engine determined carried a particular virus. That delay processing emails led to retries that further bottlenecked email flow including legitimate emails, he says.
The issue was resolved by intercepting the tainted messages and quarantining them directly.
To head off similar problems down the line, the company has set a lower threshold for diverting problem emails and implementing faster remediation tools. It is also adding unspecified safeguards that automate remediation of this type of problem, Jha says.
The second incident Nov. 13 started with some scheduled maintenance that required shifting some of the load out of those data centers undergoing maintenance. During this work unspecified network elements failed but sent no alerts of their failure, he says. And finally the entire infrastructure was handling more traffic from new customers, all of which resulted in some customers being unable to access email services.
Traffic for affected users was shifted to healthy data centers while the issues were dealt with.
Jha says the company is in the midst of increasing capacity and is automating how equipment failures are handled to speed up recovery time.
In addition, the company is reviewing its processes to head off future outages.
"As I've said before," Jha blogs, "all of us in the Office 365 team and at Microsoft appreciate the serious responsibility we have as a service provider to you, and we know that any issue with the service is a disruption to your business - that's not acceptable. I want to assure you that we are investing the time and resources required to ensure we are living up to your - and our own - expectations for a quality service experience every day."
(Tim Greene covers Microsoft for Network World and writes the Mostly Microsoft blog. Reach him at tgreene@nww.com and follow him on Twitter https://twitter.com/#!/Tim_Greene.)
Read more about wide area network in Network World's Wide Area Network section.
- The 20 Best iPhone/iPad Games of 2013 So Far
- 9 Steps to Build Your Personal Brand (and Your Career)
- 7 Consumer Technologies Coming to an Enterprise Near You
- 11 Signs Your IT Project is Doomed
- A walking tour: 33 questions to ask about your company's security
- 15 social media scams
- The 7 elements of a successful security awareness program
- IT Certification Study Tips
- Register for this Computerworld Insider Study Tip guide and gain access to hundreds of premium content articles, cheat sheets, product reviews and more.
- File Archiving - The Next Big Thing or Just Big This white paper from Osterman Research discusses best practices for archiving file-based content and offers some recommendations about how organizations should manage the...
- 3 Steps to Unlock Savings from Legacy Applications Explore a three step process to free your business from unnecessary costs and to protect your business from unnecessary risks.
- Red Hat JBoss Fuse Compared with Oracle Service Bus Competitive Brief Read this paper to learn how to start more projects, deploy technology more pervasively within the enterprise, and apply more of your budget...
- Red Hat JBoss BRMS Best Practices Guide Learn the technical best practices for development with Red Hat JBoss Enterprise BRMS. Following the best practices outlined in these guides will result...
- Boost Performance & Profitability with Better Planning & Mobile Reporting This session will discuss how Ashurst, a top-tier legal service provider for private and public sector clients worldwide, was able to effectively manage...
- Apps and BlackBerry 10 - Tips for IT Learn how to easily create, deploy and manage both off-the-shelf and custom apps, improving productivity and efficiency for employees by mobilizing apps, processes... All Applications White Papers | Webcasts
Our weekly newsletter will cover a wide range of topics and trends related to consumerization. Stay up to date with news, reviews and in-depth coverage of BYOD, smartphones, tablets, MDM, cloud, social and how consumerization affects IT. Subscribe now!