Prometheus

Metrics collection and alerting built for dynamic infrastructure

Monitoring & Analytics Apache-2.0 Advanced ★ 66,346 stars

What is Prometheus?

Prometheus scrapes metrics over HTTP on a schedule, stores them as time series, and evaluates alert rules against them. It is the reference implementation of the pull-based monitoring model and the default target for most exporters.

Best for

Collecting and alerting on infrastructure metrics

Why choose Prometheus

Prometheus is the reference implementation of pull-based monitoring, and that single design decision explains most of its adoption. Instead of agents pushing data at a central server, Prometheus scrapes a /metrics endpoint on a schedule, which means the target's health is observable from the outside and a service that stops responding is detected rather than going silent. Its query language, PromQL, is powerful enough to express real questions — error rates over a sliding window, the 95th percentile of latency by route — and its alert rules are plain files you can review in a pull request. Because it is the default target for hundreds of exporters, adopting it means every other tool you install already speaks to it.

Replaces

  • Datadog
  • New Relic
  • AWS CloudWatch

Key features

  • Pull-based scraping with service discovery
  • PromQL query language for arbitrary aggregation
  • Built-in alert rules and Alertmanager integration
  • Client libraries and exporters for almost everything

What to watch out for

Prometheus is not built for long-term storage. Its local time-series database is fast but keeps data in a single directory with no replication, and the standard advice is to retain weeks rather than years — anything longer needs Thanos, Mimir or VictoriaMetrics bolted on. It is also single-node by design, so high availability requires running two independent instances and deduplicating. Cardinality is the sharpest edge: a label with unbounded values, like a user ID or a full URL path, will multiply your time series until memory runs out. The query language is genuinely learnable but is not SQL, and newcomers routinely write queries that look correct and return nonsense.

How to deploy

  • Docker Compose with a persistent data volume
  • Configure scrape targets via file or discovery
  • Pair with Alertmanager for routing

Getting started

Run it with the official container or binary and give it a persistent volume, because losing the data directory loses all history. Start by scraping yourself and one real service, then add node_exporter and any application-specific exporters you need. Write your first recording rules for anything expensive you query twice, and treat alert rules as code — keep them in a repository and reload them on change rather than editing in place. Set an explicit retention window on day one rather than letting the disk fill and finding out. Finally, decide early where long-term data will live, because retrofitting remote-write later means rewriting your scrape configs.

Typical setup

Typically a single container or binary on a modest always-on machine, with a persistent volume holding the time-series database and a scrape configuration listing every target and exporter. Node exporter runs on each host, application exporters alongside the services they describe, and recording rules pre-compute whatever dashboards query repeatedly. Retention is set deliberately to a few weeks, with anything longer shipped to a remote store. Alert rules live in a repository and are reloaded on change, and Alertmanager sits beside it handling routing and silencing. The whole stack is usually monitored by a second, dumber instance so a failure of the clever one is still visible.

Who should look elsewhere

Do not choose Prometheus if you want a monitoring system that works without learning a query language, or if you need multi-year retention in the same install — that is VictoriaMetrics or Thanos territory. It is also the wrong tool if your metrics have unbounded label values by nature, because cardinality will kill it before traffic does. And if you only run one small server and want a screen of numbers with no configuration at all, Netdata answers that question faster.

Project health

  • GitHub stars: 66,346
  • Last code push: 2026-10-01
  • Open issues: 931
  • Status: actively developed

Figures pulled from the GitHub API and refreshed periodically.

Prometheus as an alternative

More in Monitoring & Analytics