“The damage to our infrastructure spanned multiple Availability Zones and exceeded what our regional and multi-AZ services are designed to withstand,” the AWS update said.
There is an old AWS quote that has aged in a particularly interesting way.
In 2017, CBS asked AWS executive Matt Wood a rather pointed question:
CBS: “I don’t mean to give anyone ideas, but let’s say I figured out that one of these unmarked buildings was an AWS data center, and I blew it up. Are you saying that it’s so backed up and redundant that you probably wouldn’t notice?”
AWS: “Yeah, you wouldn’t notice. I mean, we might be a bit upset, but you wouldn’t notice!”
Nine years later, somebody effectively ran the experiment, although on a much larger scale than the loss of a single building, and people did, in fact, notice. //
I Guess Maybe Now They’ll Do 3-2-1 Backups
The gold standard for data protection is “3 copies, 2 different kinds of media, 1 of which is offsite”.
I’m all for the cloud, but if it’s important data and you’re limited to keeping all of your cloud data in one set of buildings…then for pity’s sake, have off-site backups. You don’t need to have another failover site ready to light up if Amazon goes down for you, but you shouldn’t lose the data. Replicate it off-site. I mean, forget even acts of war – what if hackers get into your systems? Into Amazon’s systems? What if an employee goes rogue? What if, what if…
A replicated off-site solution with ransomware protection and disk snapshots (or heaven forbid, tape!) doesn’t mean you can restore service quickly, but it means you won’t lose the data. In other words, go ahead and say “if we have to go to that dire situation, we’re not promising a recovery time objective (RTO). But we are promising a Recovery Point Objective (RPO).”
Losing data sucks.
The Trilogy Nobody Wanted¶
Let me be real about what this three-part series documents:
Part 1: A cloud provider deletes a decade of work because of a broken verification process and a Java parameter parsing quirk. Support gaslights the customer for 20 days.
Part 2: One human inside the machine fights the bureaucracy, escalates to the CEO, and restores the account. Hope restored. Faith in humanity renewed.
Part 3: That human gets fired. The systemic issues remain unfixed. The machine continues.
This is the arc of modern tech. The system breaks. A human fixes it despite the system. The system removes the human. Repeat.
The Real Lesson¶
In Part 2, I wrote: “My trust isn’t fully restored. What is restored is my faith that even in massive corporations, one person can make a difference.”
I still believe that. But I’ll add a corollary: the difference that person makes is often inversely proportional to how long the corporation keeps them around.
The people who challenge broken systems, who go off-script, who escalate when the template says “close the ticket”. Those people are threats to institutional inertia. They’re expensive. They’re inconvenient. They make leadership answer uncomfortable questions.
And eventually, they get optimized away. Just like a “low-activity” AWS account.
A Note to AWS¶
You don’t need me to tell you this, but I will anyway: Tarus Balog was worth more to your reputation than any GenAI keynote. Every developer who read my story and thought “maybe AWS isn’t so bad after all”? That was because of him. Not your PR team. Not your marketing budget. One human being who decided to do the right thing.
You’ll replace him with someone who hits KPIs and doesn’t ask uncomfortable questions. And you’ll wonder why developers keep building exit strategies from your platform. //
To AWS: You had a human circuit breaker. You removed it. Good luck with the next cascade failure.
To everyone else: Keep your backups distributed. Keep your exit strategies current. And if you find a Tarus inside your cloud provider, thank them before the system optimizes them away.
Remember my article about AWS deleting my 10-year account? The one where support gaslit me for 20 days while claiming my data was “terminated”?
Here’s the plot twist: My data is back. Not because of viral pressure. Not because of bad PR. But because one human being inside AWS decided to give a damn.
This is that story.
I’d done everything right. Vault encryption keys stored separately from my main infrastructure. Defense in depth. Zero trust architecture. The works.
My security posture was textbook—protect against compromise by ensuring no single failure could take down everything. What I hadn’t protected against? AWS itself being the single point of failure.
I built a hardened bunker with multiple escape routes, only to have AWS drop a nuke on the entire complex. //
You might be thinking, “What are the odds they target me?” But that’s the wrong question. I thought the same thing—with my level of exposure and contributions, surely they could just write my name down and not bother me with stupid verification requests about whether I exist.
But you’re not being targeted—you’re being algorithmically categorized. And if the algorithm decides you’re disposable, you’re gone. //
After 20 days of appeals, AWS support finally responded with this gem: “Because verification wasn’t completed by the due date, your resources were terminated.”
But here’s the dilemma they’ve created: What if you have petabytes of data? How do you backup a backup? What happens when that backup contains HIPAA-protected information or client data? The whole promise of cloud computing collapses into complexity.
This isn’t a system failure. The architecture and promises are sound. AWS doesn’t lose data—they have backups of backups of backups, stored in vaults that last far longer than the stated 90 days, where no rogue AI script can reach.
What’s happening here is simpler: teams in MENA are trying to cover up a massive fuck-up. Restoring data from those deep vaults would require explanations. Incident reports. Post-mortems. “Why did we have to open the vaults?”
Their entire communication strategy screams: “He’s nobody. He’ll give up soon. We won’t have to report this up the chain.” //
Lessons Learned¶
- Never trust a single provider—no matter how many regions you replicate across
- “Best practices” mean nothing when the provider goes rogue
- Document everything—screenshots, emails, correspondence timestamps
- The support theater is real—they literally cannot help you
- Have an exit strategy executable in hours, not days
AWS won’t admit their mistake. They won’t acknowledge the rogue proof of concept. They won’t explain why MENA operates differently. They won’t even answer whether your data exists.
But they will ask you to rate their support 5 stars.
The cloud isn’t your friend. It’s a business. And when their business needs conflict with your data’s existence, guess which one wins?
Plan accordingly.
“Affected devices include Kindle 1st and 2nd Generation, Kindle DX and DX Graphite, Kindle Keyboard, Kindle 4, Kindle Touch, Kindle 5, and Kindle Paperwhite 1st Generation,” reads the message from the Kindle team. Older 2011 and 2012-era Kindle Fire tablets will also lose access to the Kindle Store.
Amazon’s Kindle generational branding is occasionally confusing—that “Kindle Paperwhite 1st Generation” is also referred to as “Kindle Paperwhite (5th Generation)” on Amazon’s support pages because it’s part of the fifth generation of Kindle releases overall. But if you check your Kindle’s software version and see anything older than 5.12.2.2, it means your Kindle is losing access to Amazon’s store and your e-book library.
It’s been a while since any of these devices received active software support from Amazon; only 2024-and-later devices have received the latest 5.19.3.0.1 software update, though 2021 and 2022’s Kindles have been updated as recently as February. Historically, though, Amazon has been willing to allow older, un-updated Kindles to continue to buy and download more books, even if they’re no longer benefitting from new features.
Mar 02 4:22 PM PST We are providing an update on the ongoing service disruptions affecting the AWS Middle East (UAE) Region (ME-CENTRAL-1) and the AWS Middle East (Bahrain) Region (ME-SOUTH-1). Due to the ongoing conflict in the Middle East, both affected regions have experienced physical impacts to infrastructure as a result of drone strikes. In the UAE, two of our facilities were directly struck, while in Bahrain, a drone strike in close proximity to one of our facilities caused physical impacts to our infrastructure. These strikes have caused structural damage, disrupted power delivery to our infrastructure, and in some cases required fire suppression activities that resulted in additional water damage. We are working closely with local authorities and prioritizing the safety of our personnel throughout our recovery efforts.
In the ME-CENTRAL-1 (UAE) Region, two of our three Availability Zones (mec1-az2 and mec1-az3) remain significantly impaired. The third Availability Zone (mec1-az1) continues to operate normally, though some services have experienced indirect impact due to dependencies on the affected zones.
Mar 02 4:22 PM PST We are providing an update on the ongoing service disruptions affecting the AWS Middle East (UAE) Region (ME-CENTRAL-1) and the AWS Middle East (Bahrain) Region (ME-SOUTH-1). Due to the ongoing conflict in the Middle East, both affected regions have experienced physical impacts to infrastructure as a result of drone strikes. In the UAE, two of our facilities were directly struck, while in Bahrain, a drone strike in close proximity to one of our facilities caused physical impacts to our infrastructure. These strikes have caused structural damage, disrupted power delivery to our infrastructure, and in some cases required fire suppression activities that resulted in additional water damage. We are working closely with local authorities and prioritizing the safety of our personnel throughout our recovery efforts.
In the ME-CENTRAL-1 (UAE) Region, two of our three Availability Zones (mec1-az2 and mec1-az3) remain significantly impaired. The third Availability Zone (mec1-az1) continues to operate normally, though some services have experienced indirect impact due to dependencies on the affected zones.
I received an email / billing notification from AWS this week that may be the most diplomatically crafted communication in the history of cloud computing. Here it is, stripped of the usual boilerplate around it:
"AWS is waiving all usage-related charges in the ME-CENTRAL-1 Region for March 2026. This waiver applies automatically to your account(s), and no action is required from you."
No explanation. No mention of the Iranian drone strikes that physically destroyed two of three availability zones in the region on March 1st. No reference to the 109 services that went down, nor the customers who spent weeks unable to terminate EC2 instances via the console because the control plane was as dead as the hardware underneath it. No acknowledgment that an entire month of cloud infrastructure effectively ceased to exist. Not even a link to their remarkably short (presumably because it wasn't insulting the Financial Times' reporting) corporate blog post explaining that you probably shouldn't expect that region to be working reliably again any time soon.
Just: we're waiving the charges. You're welcome. Move along.
I want to be clear: I have no problem with this. It's a tough situation, and it's not AWS' fault, given that there is not yet an Amazon standing military force.
But here's the part that caught my attention. The email continues: "You will not see any March 2026 usage for the ME-CENTRAL-1 Region in your Cost and Usage Report or Cost Explorer once processing is complete."
They're not just waiving customer charges for a month; they're erasing the billing and inventory data! //
For most organizations, the AWS bill isn't just an invoice. It's the canonical record of what infrastructure exists, where it's running, and how long it's been there. The Cost and Usage Report (CUR) is the closest thing many companies have to a single source of truth that accurately describes their cloud footprint.
AWS's destiny isn't to lose to Azure or Google. It's to win the infrastructure war and lose the relevance war. To become the next Lumen — the backbone nobody knows they're using, while the companies on top capture the margins and the mindshare.
The cables matter. But nobody's writing blog posts about them. ®
It’s always DNS
Amazon said the root cause of the outage was a software bug in software running the DynamoDB DNS management system. The system monitors the stability of load balancers by, among other things, periodically creating new DNS configurations for endpoints within the AWS network. A race condition is an error that makes a process dependent on the timing or sequence events that are variable and outside the developers’ control. The result can be unexpected behavior and potentially harmful failures.
In this case, the race condition resided in the DNS Enactor, a DynamoDB component that constantly updates domain lookup tables in individual AWS endpoints to optimize load balancing as conditions change. As the enactor operated, it “experienced unusually high delays needing to retry its update on several of the DNS endpoints.” While the enactor was playing catch-up, a second DynamoDB component, the DNS Planner, continued to generate new plans. Then, a separate DNS Enactor began to implement them.
The timing of these two enactors triggered the race condition, which ended up taking out the entire DynamoDB.
All those words, and yet there is no mention of where the product is made. The title hints it may be Italy. But under the product information, the country of origin line is missing. The manufacturer is Superbuy. It sounds like an American name, but no, it is a company that buys and ships Chinese items. The seller, GoPlusUS, has a Chinese address.
How do Chinese manufacturers manage to produce things so much cheaper?
Prisoners.
China has detained Uyghurs, Falun Gong practitioners, and members of other ethnic and religious minority groups in roughly 1,200 state-run internment camps.
“Detention in these camps is intended to erase ethnic and religious identities under the pretext of ‘vocational training.’ Forced labor is a central tactic used for this repression,” a U.S. State Department statement said in January.
“In Xinjiang, the government is the trafficker. Authorities use threats of physical violence, forcible drug intake, physical and sexual abuse, and torture to force detainees to work in adjacent or off-site factories or worksites producing garments, footwear, carpets, yarn, food products, holiday decorations, building materials, extractives, materials for solar power equipment and other renewable energy components, consumer electronics, bedding, hair products, cleaning supplies, personal protective equipment, face masks, chemicals, pharmaceuticals, and other goods — and these goods are finding their way into businesses and homes around the world.”
If you care about liberty, if you hate slavery, if you want fair trade, then you give a cock-a-doodle-damn where your rooster was painted.
Amazon should make it simple to find the country of origin for every product.
In accordance with certain free and open source software licenses, Amazon is pleased to make available to you for download an archive file of machine readable source code ("Source Code").
Review distinguishing features to help determine which device you have.
Haven't used or updated your Kindle device for a while? Depending on your device and current software update version, you may have to install a previous software update before installing the latest version.
Software updates automatically download and install on your Kindle when connected wirelessly. These free software updates include general improvements and performance enhancements.
Before You Start
Be sure to download software updates onto a desktop computer in order to install them onto your device. For more information, visit Manually Update Your Kindle E-Reader Software.
Note: Determine what Kindle E-Reader model you're using before downloading any software updates. Refer to Identify Your Kindle E-Reader.
Starting Feb. 26, 2025 (via ZDNet) you will no longer be able to download copies of your Kindle books and use those files as a backup. After that date, you will only be able to download books via Wi-Fi or through the Amazon platforms. //
All in all it is a reminder that you don't actually own many or most of your digital purchases, as what you are typically actually "buying" are licenses to use content that can be revoked at any time.
A couple of years ago, Amazon removed the ability to purchase e-books with their flagship Kindle for Android from the Google Play Store. Google implemented a policy that all in-app purchases had to be made using their billing system. Instead of paying Google 30% of each e-book sold, Amazon removed the ability to buy e-books or audiobooks. //
Amazon operates their own Android App store. You can install the Amazon App Store and download the Kindle app if you have a smartphone or tablet. All of the audiobooks and e-books that are available in the store can be purchased. This is because Amazon uses their billing system, the same one that their Kindle e-readers use.
Google, Amazon, Microsoft dive into costly deals that aren't generating anything yet. //
Nuclear power contracts signed by hyperscalers show they're desperate for reliable "clean and green" energy sources to feed their ever-expanding datacenter footprints, however, investment bank Jefferies warns that these tech giants are likely to end up paying over the odds to get it.
Under the terms of the sale, Amazon not only acquired Cumulus' datacenter facilities and associated power infrastructure, but has direct — behind the meter — access to a sizable chunk of the energy generated by the nuclear plant's two reactors.
Over the course of its contract with Talen, Amazon expects to unlock upwards of 960 MW of power supply. However, we'll note that the cloud titan has the option to cap this at 480 MW if it doesn't actually need all of it.
We've now learned at least 15 new datacenters will be built adjacent and connected to the fission plant over the next ten years. As we've previously reported, an AWS campus with five buildings may take up 600,000 square feet, or around 13 acres, and capacity of between 50 and 60 megawatts. //
As generative AI has taken off, it's not uncommon to see clusters of 20,000 or more GPUs capable of consuming in excess of 25 megawatts of power, deployed.