MongoDB

Document database for flexible schemas

Databases & Storage SSPL Intermediate ★ 28,617 stars

What is MongoDB?

MongoDB stores JSON-like documents with a flexible schema, which suits applications whose data shape changes often. Several popular self-hosted apps use it as their primary store.

Best for

Applications whose documents do not fit a fixed table

Why choose MongoDB

MongoDB is the database for applications whose data refuses to be tabular. Documents nest naturally, new fields can appear without a migration, and an entire user profile or product record is a single read rather than a join across five tables. Several genuinely popular self-hosted applications — including some of the biggest in the media and automation categories — use it as their primary store and expect it to be running, so in practice it is often a requirement rather than a choice. Its aggregation pipeline is expressive enough to do real reporting work in the database, and replica sets give you failover without a separate clustering layer.

Replaces

  • Amazon DocumentDB
  • Firebase Firestore
  • CouchDB

Key features

  • Document model with flexible fields
  • Aggregation pipeline for transformations
  • Indexing including text and geospatial
  • Replica sets for high availability

What to watch out for

The flexibility that makes it approachable is also the classic source of trouble: without schema discipline enforced at the application layer, one document shape drifts from another and queries silently miss records. Indexing is not optional the way it can be for small relational tables; an unindexed query on a large collection performs a full scan and will look like a bug for weeks. Memory use is aggressive, because the storage engine wants to keep working data in RAM — a small instance will page heavily. The licensing has changed over releases and the open-source edition is not the same product as the commercial one, so features described in the documentation may not exist in what you installed. Self-hosted replica sets and sharding are real operational work.

How to deploy

  • Docker with a persistent data volume
  • Enable authentication immediately
  • Back up with mongodump or snapshots

Getting started

Install a version that matches what your applications document, and enable authentication before the port is reachable by anything but localhost — an unauthenticated MongoDB on a public address is compromised within hours. Set a data directory that is backed up with mongodump rather than by copying files at random, since a filesystem copy of a running instance is not a valid backup. Create the indexes your queries need before the collection grows, and watch for collections the application creates that you did not anticipate. Monitor disk and memory rather than CPU, because it degrades in ways that look like application slowness. For anything you care about, run a replica set rather than a standalone node.

Typical setup

A mongod instance with authentication enabled and bound to an internal interface, never a publicly reachable port. Backups use mongodump on a schedule rather than copying files from a running server, and a restore has been verified. The indexes the application's queries need are created before collections grow, and the ones the application creates at runtime are reviewed rather than ignored. Disk and memory are monitored because degradation shows there first. For anything beyond a lab, a replica set rather than a standalone node, with the connection string updated to match.

Who should look elsewhere

Do not choose it for data that is genuinely relational, because you will end up implementing joins in application code and losing transactional guarantees that a relational database would give you for free. Avoid it on very small instances where its memory appetite will cause swapping. And if you cannot enforce document schema discipline anywhere in your stack, the flexibility will turn into data quality problems that are expensive to unwind.

Project health

  • GitHub stars: 28,617
  • Last code push: 2026-10-01
  • Open issues: 36
  • Status: actively developed

Figures pulled from the GitHub API and refreshed periodically.

More in Databases & Storage