The decision was made in a room I was in, and I did not object at the time. A single provider felt like a single point of failure, the commercial team wanted leverage in the next contract negotiation, and somebody senior had read an article about concentration risk. So we committed to running across two clouds. Two years later we had spent a great deal of money and effort to be worse at operating infrastructure than we had been with one, and the outage we had been insuring against had still not happened.
The cost was not the bill, although the bill was not small. It was that every capability had to exist twice, in two dialects, and the two dialects were never quite equivalent. Identity worked differently. Networking worked differently. Managed databases had different failure modes, different backup semantics, and different upgrade cadences. To keep the workloads genuinely portable we had to avoid the good managed services on both platforms and rebuild the boring middle layer ourselves, which meant we were paying cloud prices for something close to a data centre. The abstraction that was supposed to free us became the largest piece of undifferentiated software we owned.
The human cost was worse than the technical one. Expertise thinned out. Instead of a team who knew one platform deeply, we had a team who knew two platforms adequately, and adequate is exactly the wrong depth of knowledge at three in the morning. On-call got harder because the runbooks forked. Hiring got harder because the job description was twice as long. Every architectural discussion acquired an extra half hour of caveats about which environment we were talking about.
What I would defend today is something narrower and more honest. Know precisely which risk you are buying down, and check whether it is the one that actually threatens you. Provider-wide failure is rare; a bad change of your own is not. Keep your data exportable, keep your infrastructure described as code, and avoid gratuitous dependence on the one service you could never replace. That is most of the real insurance, at a fraction of the premium.
Redundancy you cannot operate confidently is not resilience. It is just a second system to be surprised by.
– Serguey Shinder