Your alerts are noise and your noise is an outage
I have never audited a startup whose alerting was under-configured. I have audited dozens whose alerting was so loud that the team had stopped looking at it. The alert that pages you at 3am for a CPU spike that resolves itself is not a safety net. It is the reason you miss the alert that matters. Here is how I strip alerting down to what actually predicts incidents, with the Prometheus rules and routing configs I install for every team.