AWS confirms unrecoverable data loss after Iran strike
AWS cannot restore some data from Iran-struck Mideast facilities. A recovery copy that shares a failure domain with the primary is not recovery.
AWS states it cannot restore some data from Middle East facilities struck by Iran. That is the fact, and it is the entire fact. This is not a breach, not an intrusion, not an access event. It is confirmed data loss with no recovery path for a defined subset of data. When a provider says restoration is not possible, the recovery objective for that data is zero. There is nothing to negotiate down to. The data is gone, and the provider has said so directly.
A physical strike on a facility is a destruction event, not a security event in the traditional sense. The question is not whether the facility was defended, hardened, or monitored. None of that is confirmed and none of it changes the outcome. The only question that matters is whether the data existed independently of the facility that held it. For some data, AWS has answered that question. It did not.
Treat the statement as a control test that already ran. The event happened. The recovery function was invoked. For a portion of the data, it returned nothing. Everything downstream of that outcome is either a stated fact, a logically necessary consequence of it, or not confirmed. I will hold that line for the rest of this briefing.
What failed is the recovery function itself, for that subset of data. Backup exists for one reason: to make loss recoverable. Here, loss is confirmed and recovery is confirmed impossible. That is the observable failure. The restore operation for the affected data does not return the data, and the provider has stated that it cannot. This is externally observable system behaviour. I am not describing what happened inside the storage layer, because that is not stated and I will not invent it.
The scope of the loss is not confirmed. Which data, how much, which customers, and which services are affected are not stated. The word “some” tells us only that a subset failed to restore. It does not tell us the size of that subset, and it does not describe the state of the data that was not named. Whether other data was successfully restored is not confirmed. I will not scale the impact beyond the single fact provided: some data cannot be restored.
The internal design of the affected systems is also not confirmed. Replication topology, backup frequency, storage tiering, and region configuration are not stated. Whether any backup control existed at all is not confirmed. Absence of restoration is what is observable. Absence of a described control is not evidence that a control existed and failed quietly. It is simply not confirmed. The strike is a stated fact. The inability to restore is a stated fact. The connection between them is asserted by the source. Beyond those three points, there is no confirmed information about how the system behaved.
The reason for the failure follows as a logically necessary implication of the stated outcome. If the affected data could be restored from any surviving copy within reach, it would be restored. AWS states it cannot. Therefore, for that data, no independent recoverable copy was available outside the loss boundary of the struck facility. The recoverability of the data was bound to the facility that was destroyed. That is the mechanism, and it is the only mechanism the facts support.
A copy that shares the fate of the primary is not a backup. If the ability to restore depends on hardware in one location, then destruction of that location is destruction of the data. The strike did not defeat an enforcement point or bypass an access boundary. The observable outcome indicates the recovery path and the primary were exposed to the same single event. Recoverability was not separated from the asset it was meant to protect.
Why the recoverable copies shared that exposure is not confirmed. Whether backups existed elsewhere and also failed, whether they were configured to a single geography, or whether they never existed, is not confirmed. Whether this was a customer configuration or a provider default is not confirmed. What is confirmed is the outcome, and the outcome constrains the cause: the destruction of one facility removed the ability to restore some data. That is only possible if recovery for that data was not independent of the facility. Everything past that boundary is not confirmed, and I will not present it as if it were.
The mechanism is single-failure-domain collapse. A failure domain is any boundary within which one event can destroy everything inside it. When the primary data and its only recoverable copy sit inside the same domain, that domain is a single point of failure for both the asset and its recovery. The strike operated on that domain. Everything inside it, including whatever recovery capability was present for the affected data, was inside the reach of one event. The outcome AWS stated is only possible if that condition held.
This is not a control bypass. No identity boundary was crossed, no trust relationship was abused, no enforcement point was disabled. The recovery function did not fail because it was attacked. It failed because, for that subset of data, it did not exist outside the destroyed domain. A recovery path that lives inside the thing it is meant to recover from is not a recovery path. It is a second instance of the same exposure. The strike did not have to defeat two controls. There was one domain, and the destruction of that domain removed both the data and the means to restore it.
The distinction that matters here is redundancy versus independence. Redundancy inside one domain survives component failure. It does not survive domain failure. Independence means the copy is reachable after the domain is gone. For the affected data, independence was absent, by logical necessity of the stated outcome. Whether it was never configured, configured to the same geography, or configured elsewhere and also lost is not confirmed. The mechanism holds under every one of those cases: the recovery capability shared a fate with the primary, and the primary’s fate was destruction.
The pattern derived from that mechanism is shared-fate recovery. Any recovery copy that can be destroyed by the same event as the primary provides no recovery against that event. The scope of the shared domain can be a rack, a power feed, a building, a region, or an administrative boundary, any unit that a single event can take out at once. The size of the domain changes. The mechanism does not. If the primary and its recovery are inside the same domain, that domain is the single point of failure for the data, and its recovery objective under a domain-level event is zero.
This exposes a measurement error, not a technology error. Backup existence is measured by whether copies exist. Recovery is measured by whether a copy survives the loss event and is reachable after it. Those are different tests. Counting copies answers the first. Only mapping copies to failure domains answers the second. A recovery plan validated by whether backups exist, and not by whether backups exist outside the loss boundary, will report healthy until the loss boundary is actually tested. The strike tested it. For some data, the measured result was zero recovery, and it was measured after the event rather than before it.
The pattern is a property of topology, not of vendor. It applies to any provider and any customer because it depends on where copies sit relative to failure domains, not on whose name is on the facility. The question is never whether a backup was taken. The question is what single event can destroy both the data and every copy capable of restoring it. If that event is possible and plausible, the recovery objective for that data under that event is zero whether or not anyone has counted it. AWS has now confirmed one such result for a defined subset of data. The confirmation is the outcome of the test. It is not the cause. The cause was the topology that existed before the strike.
What must now be true is that recovery capability is independent of the failure domain of the asset it protects. If a single event can destroy the data and its recovery together, recovery does not exist for that event, regardless of what a backup inventory reports. Independence is the control. It is enforced by geographic and administrative separation that no single physical or logical event can span. A copy that cannot be shown to survive the primary’s destruction does not count as recovery, and it must not be recorded as recovery in any plan that claims to have one.
This is an accountability line, not a vendor line. Whether the affected recovery configuration was a provider default or a customer choice is not confirmed. The obligation does not move with that distinction. The owner of the data owns the recovery objective for that data. Delegating storage to a provider does not delegate the requirement that recovery be independent of the loss boundary. The party that needs the data restored is the party that had to have verified, before the event, that a copy existed outside the domain that could be destroyed. Verification after the event is not verification. It is notification of the result.
For the affected data, the test is over and the result is permanent. There is no configuration change that restores data with no surviving copy. Every available action is forward. Identify, for each data set, the single events that can destroy the primary and its recovery together. Where such an event exists, treat the recovery objective as zero until an independent copy exists outside that boundary. Assume the event will occur, because if a system permits an outcome, that outcome is available. AWS has now shown one such outcome in full and stated it plainly. Treat the statement as the control test it is, and do not wait for the next facility to run the same test again.
See also: NordVPN for tunneled traffic when operating outside controlled networks.
#ad Contains an affiliate link.
Keep Reading
cloud resilienceRubble where the records used to be
A cloud provider states some Middle East data cannot be restored after a strike - what permanent loss means for board-level accountability and recoverability.
AI safetyMay 2024's Bend is no proof assistant
Bend is a parallel language, not a proof assistant. What proof-based programming actually does for AI safety, and the errors it can't touch.
technical debtPermanent by default
A temporary PHP fix reached 20M installs because a label is not a control. The mechanism, the pattern, and what must now be true in production.
Latest on the Wire
Full wire →- Bend: a proof-checked language that aims to make AI coding bugs unmergeableHacker News
- Bonsai 2 27B: a 27B model in 5.9GB that keeps 98% of its benchmarksHacker News
- CrowdSec Confirms Private Source Code Leak Traced to Tanstack Supply-Chain BackdoorHacker News
- Unverifiable: no retrievable content for "Astra for Law"Hacker News
New signal daily · RSS
Stay in the loop
New writing delivered when it's ready. No schedule, no spam.