As of today around 10 am (CET), Mobypicture has been down because Amazon Web Services is experiencing problems with connectivity in some of their Availability Zones, amongst the one hosting our servers.
Besides taking us down with them, also Foursquare, Reddit, Quora and many other services are down or experiencing heavy troubles.
Read more:
Mashable
Techcrunch
You can follow our tweets to stay informed or keep an eye on
http://mobypicture.com or our status page.
Sorry for the inconvenience. We’re aware of the implications of relying on cloud services and are designing a more redundant way to ensure our up time.
Hosting companies can stop emailing us with propositions. We love Amazon and their services and because of their scalable proposition we’ve been able to grow lean and mean. We trust in a prompt recovery.
We’ll keep you updated. Hold your Mobys
Update 21:21
After hours of hard work we managed to get back online with a series of workarounds, Amazon AWS is still having severe issues with an ETA of several hours.
We changed key elements of our infrastructure and are running on a degraded setup of our database servers at the moment, expect degraded performance for the next couple of hours. We will monitor everything closely until all issues are resolved. We are truly sorry for not being there today.
Update #2: Our workarounds
Because most of the troubles originated from EBS and all our database servers run on multiple EBS blocks in a RAID configuration, we were hit the hardest in that area. Luckily one of our slave database servers survived, which we then promoted to one of the masters. Starting new EBS backed instances was at that time still impossible. We started non-EBS backed instances which we must replace with the faster EBS backed when all troubles are over.
At that time we were running in a development environment on a platform with degraded performance, with one big issue: ELB, our loadbalancer, was still offline. We were luckily already using our own loadbalancing solution HAproxy, to handle API connections and we migrated our webserver architecture to that as well.
So at the moment we are not as scalable as we would like to be and have a slight degraded performance, mostly noticable on our MobyNow platforms, but are running smoothly since 21:00 hours CET yesterday. We expect to upgrade to full performance somewhere today and add our scalability when all AWS issues are resolved.
bron: http://mathys.vanabbe.com/thunder-in-the-cloud/