Going entirely on gut feel, I suspect that outages like this will be 1) less common and 2) addressed more quickly with AI-supported DevOps in the not-too-distant future.
The actor pattern exists for this reason - it's not like AI is a singular thing checking itself - you could easily have 1000 instances with supervisors automatically rotating out unhealthy instances so that the system is self-repairing even if instances can fail.