Your dependency map should be written down before the outage, not discovered during it.
Tuesday night one of our vendors went down for about seventy minutes. Spanish generation, PSD export, text edits, and video all failed. Core English generation kept working, because it does not touch that vendor. That single fact is the only reason the outage was a bad hour instead of a bad night.
Here is the uncomfortable part. Our dependency map only became visible during the outage. We learned which features depended on which vendor by watching what broke. That is the wrong time to learn it. A customer who had bought that same afternoon spent ninety minutes failing generations before we caught it. We auto-refunded credits, topped up the account, sent an apology, and I proved the fix end to end with a real generation rather than assuming it worked.
For a PE-backed platform, this is not a hobby-project problem. It is the exact risk that surfaces the week after an acquisition closes. Every brand you roll up brings its own quiet stack of third-party dependencies: the dispatch tool, the payment processor, the call routing, the review platform. Nobody has written down what dies if any one of them goes dark, so nobody can tell you your true blast radius until the day it detonates.
The reliability move is boring and cheap. Write the dependency map before you need it. For every customer-facing capability, name the external services it touches and what the failure looks like to the customer. Then decide, on purpose, which capabilities must survive a vendor outage and wall those off from anything external.
Resilience is not something you discover at 11pm. It is something you decided months earlier and simply confirmed.
AI Value-Creation Diagnostic | All insights