An on-call agent gets one action, then a human
The failure mode of an autonomous system is a chain of reasonable actions that compound. Restart did not help, so scale up, so roll back, and now three things changed. Let the agent do one thing per incident, announced, from a fixed list. If it did not resolve the alert, a person decides the next step.
ai-agentssre
Longer version: the post this came from.