Grafana Loki
Log aggregation that indexes labels instead of content
What is Grafana Loki?
Loki collects logs from every service and stores them cheaply by indexing only labels rather than the full text. Queries use LogQL, and it plugs directly into Grafana so logs and metrics sit side by side.
Best for
Centralising logs without an Elasticsearch-sized budget
Why choose Grafana Loki
Loki's central bet is that you rarely need to full-text search old logs, so it indexes only the labels and stores the log lines compressed and unindexed. The result is an order of magnitude less storage and memory than an Elasticsearch cluster for the same volume, which is what makes centralised logging affordable on hardware you own. It uses the same label model as Prometheus, so the service, environment and instance labels you already attach to metrics work here unchanged. Because it plugs into Grafana, a dashboard panel showing a latency spike and the log lines from that same minute sit side by side, which is usually the fastest route from symptom to cause.
Replaces
- Elasticsearch
- Splunk
- Datadog Logs
Key features
- Indexes metadata labels, not message content
- LogQL query language with metric extraction
- Object storage backend keeps costs low
- Native Grafana data source
What to watch out for
Label discipline matters more than in most systems. Putting high-cardinality values like request IDs or user IDs into labels will explode the index and degrade the whole cluster — those belong in the log line, not the label. Queries over large time ranges without a label filter are slow by design, because the system has to scan. Loki deliberately does not do full-text indexing, so if your workflow is 'search every log for this obscure string across six months', a real search engine will be better. Running it in the recommended microservices mode is a genuine distributed system to operate, and the single-binary mode that is easy to start is not what you want under heavy load.
How to deploy
- Docker Compose with object or filesystem storage
- Run Promtail or Alloy on each host to ship logs
- Add Loki as a Grafana data source
Getting started
Start with the single-binary deployment and a filesystem or object-storage backend, and only split into separate components when a specific bottleneck forces it. Decide your label schema before shipping logs: typically service, environment, level and maybe cluster — never identifiers. Ship logs with Promtail or Alloy from the machines that generate them, and verify a query returns results before wiring up more sources. Set a retention period and a compactor configuration explicitly, because the default will keep accepting data until something fills. Once it works, add a Grafana panel that links a metrics spike to the logs for the same service and time window.
Typical setup
The common shape is a single binary or small container set with object storage or a persistent filesystem for chunks, fed by an agent on every machine that produces logs. The label schema is fixed before anything ships — typically service, environment and level — and it is enforced, because retrofitting labels means re-indexing everything. LogQL queries in Grafana sit next to the metrics for the same service, so a spike and its log lines open together. Retention is set explicitly with a compactor, the storage backend is backed up or versioned, and the deployment itself is watched so that losing logs is something you learn about from a monitor rather than from an incident.
Who should look elsewhere
Do not pick Loki if your primary need is interactive full-text search over long historical windows — it is designed to make that cheap to store, not fast to grep, and Elasticsearch or OpenSearch is the honest answer there. It is also the wrong choice for a single application that already writes structured logs to a file you rotate, where the operational cost of a log cluster exceeds any benefit. And if you cannot control what your applications put into labels, Loki will punish you for it.
Project health
- GitHub stars: 28,983
- Last code push: 2026-10-02
- Open issues: 1,196
- Status: actively developed
Figures pulled from the GitHub API and refreshed periodically.