Alertmanager
Routing, grouping and silencing for your alerts
What is Alertmanager?
Alertmanager receives alerts from Prometheus and decides what to do with them: group related ones, route by severity, silence during maintenance windows, and deliver to the right channel. It is what keeps an on-call channel readable.
Best for
Making alerting bearable when you run many checks
Why choose Alertmanager
Alertmanager is the missing half of Prometheus. Prometheus decides that a condition is true; Alertmanager decides who should be told and how often, and that separation is what keeps an on-call channel readable. It groups related alerts into a single notification, so a rack losing power does not generate eighty separate messages. It routes by label, so database alerts go to the database team and disk alerts go elsewhere. It silences alerts during planned maintenance so real problems are not buried in expected noise, and it inhibits alerts — suppressing downstream symptoms while a root cause is already firing. Without it, most monitoring setups eventually collapse under their own volume.
Replaces
- PagerDuty
- Opsgenie
- VictorOps
Key features
- Grouping to collapse alert storms into one message
- Route trees by label, team or severity
- Silences and inhibition rules
- Email, Slack, webhook and many other receivers
What to watch out for
The configuration file is the hard part and the place where most mistakes live. Routing trees are evaluated in order with continue flags that behave subtly, and it is entirely possible to write a config where alerts vanish because no route matched and no default receiver existed. Every notification channel has its own quirks — grouping keys, rate limits, retry behavior — and a misconfigured receiver can fail silently. Silence and inhibition rules are powerful enough to hide real incidents if written carelessly. Alertmanager holds state in memory and on disk but is not a long-term record; it will not tell you what fired last month.
How to deploy
- Docker or binary
- Configure route tree and receivers in YAML
- Point Prometheus at it via alerting rules
Getting started
Write a default receiver that catches everything before you write any routing rules, so nothing can silently disappear. Start with one notification channel and add a second of a different kind only after the first is proven, because two failing channels hide each other. Group by the labels that identify the actual incident — service and cluster — rather than by instance, or a single outage becomes a flood. Test the whole path by firing a throwaway alert rather than trusting the config on inspection. Review every silence with an expiry date, and keep the config in a repository so a bad edit can be reverted.
Typical setup
One small service beside Prometheus, with a persistent volume because silences and notification state survive restarts. A default receiver catches anything unmatched, so no alert can silently disappear, and at least two notification channels of different kinds are configured so one provider outage cannot mask an incident. Routing groups by the labels that identify real incidents — service and cluster — rather than by instance. Silences carry expiry dates and get reviewed. The configuration lives in a repository and is validated before reload, and the service is monitored by something independent of the stack it serves.
Who should look elsewhere
If you have one alert and one recipient, Alertmanager is process for its own sake — send that notification directly. It is also the wrong tool if you need alert history, dashboards or an incident timeline, because it deliberately forgets. And if nobody will own the routing configuration, a complex config drifts into a state where alarms are quietly discarded, which is worse than having none.
Project health
- GitHub stars: 8,636
- Last code push: 2026-10-01
- Open issues: 426
- Status: actively developed
Figures pulled from the GitHub API and refreshed periodically.