ClickHouse

Columnar database for analytical queries

Databases & Storage Apache-2.0 Advanced ★ 50,191 stars

What is ClickHouse?

ClickHouse is built for aggregating enormous numbers of rows quickly. It is what sits behind open-source observability platforms and any analytics workload where a normal relational database would choke on a GROUP BY.

Best for

Analytical queries over hundreds of millions of rows

Why choose ClickHouse

ClickHouse is what you reach for when a query has to aggregate hundreds of millions of rows and come back quickly. Its columnar storage means a query touching three columns does not read the other forty, its compression routinely reduces analytical data to a fraction of its original size, and it parallelises across cores as a matter of course. Benchmarks aside, the practical effect is that questions which required pre-computed rollup tables in a relational database become ad-hoc queries. It is the engine behind several open-source observability platforms, which is a strong signal about where it fits: high-volume append-only data that gets sliced and summarised rather than updated row by row.

Replaces

  • Google BigQuery
  • Amazon Redshift
  • Snowflake

Key features

  • Column-oriented storage with aggressive compression
  • Vectorised query execution
  • Materialised views for pre-aggregation
  • SQL with a large function library

What to watch out for

It is not a general-purpose database and will punish you for treating it as one. Updates and deletes are expensive and asynchronous, so anything needing frequent row-level modification belongs elsewhere. Small, frequent inserts are the wrong pattern — the engine expects batched writes, and a per-row insert loop will perform catastrophically. Joins are supported but are not its strength, and a query joining huge tables can exhaust memory rather than simply run slowly. Running a cluster is a real distributed system to operate, and the memory requirements are not modest. Version churn is rapid and configuration compatibility across releases is something to verify rather than assume.

How to deploy

  • Docker or a single-node install
  • Size memory generously, it is the main constraint
  • Set a TTL policy for retention

Getting started

Start with a single node and load real data before adding any cluster machinery, because most analytical needs fit on one adequately sized machine. Batch your writes — collect events and insert in blocks of thousands rather than one at a time — and design your table ordering around the columns you filter by most. Enable compression explicitly for the columns that benefit. Set a memory limit per query so a runaway aggregation fails cleanly instead of taking the host down. Do not expose the HTTP interface publicly, and understand that backups should use the native backup mechanism or a replica rather than a filesystem copy.

Typical setup

A single node sized with real memory headroom, receiving batched inserts from a collector rather than per-row writes. Table ordering is designed around the most common filters, compression is enabled explicitly, and a per-query memory limit is set so a runaway aggregation fails cleanly. Dashboards query it directly for ranges that would be impractical elsewhere. Backups use the native mechanism or a replica rather than a filesystem copy, and the HTTP interface is bound internally only. Cluster mode is added only once one node is demonstrably the limit.

Who should look elsewhere

Do not use ClickHouse for transactional workloads, for data that is frequently updated in place, or as a drop-in replacement for PostgreSQL. It is also a poor fit if your data volumes are modest — the operational cost is not repaid until queries actually hurt on a conventional database. And if you need strong consistency for individual records with frequent small writes, this is precisely the wrong engine.

Project health

  • GitHub stars: 50,191
  • Last code push: 2026-10-02
  • Open issues: 8,085
  • Status: actively developed

Figures pulled from the GitHub API and refreshed periodically.

More in Databases & Storage