Paperless-ngx

Scan, index and archive your paper documents

Productivity & Notes GPL-3.0 intermediate ★ 46,199 stars

What is Paperless-ngx?

Paperless-ngx turns your physical documents into a searchable online archive. It performs OCR on scanned files, automatically tags and classifies them, and lets you find any document by content. A community-maintained fork of the original Paperless project.

Best for

Households drowning in paper bills, contracts and receipts

Why choose Paperless-ngx

Paperless-ngx is the rare self-hosted project that changes a physical habit. You scan a document — a bill, a contract, a medical letter — and from then on it is text-searchable, tagged, filed and retrievable by any word it contains. The archive is on your own disk, encrypted or not as you choose, and it never gets sent to a document service. For anyone with years of paperwork and no desire to keep filing cabinets, the payoff starts on day one.

Replaces

  • Evernote
  • Dropbox paper workflows

Key features

  • OCR with multiple languages
  • Automatic tagging and classification
  • Full-text search
  • Email and scan ingestion

What to watch out for

Getting documents in is the awkward part. Quality depends on your scanner and on OCR accuracy, and handwritten or low-resolution scans will produce text you cannot search. The machine-learning classification and document matching work well once trained but need a decent volume of examples first, so early results are underwhelming. Consumption directories and email ingestion have surprising edge cases, and a poorly named file can create a duplicate document that you then have to reconcile by hand.

How to deploy

  • Docker Compose

Getting started

Deploy the Docker Compose stack and decide your folder structure for originals, consume and media before importing anything — reorganising later means moving files the application tracks by path. Set the correct OCR language and the right date-parsing rules for your locale early, because both affect every document you add. Point your scanner at the consume directory over a network share and scan a batch of twenty boring documents to learn how names and tags come through. Take a proper backup of the database and media folder together; a restore needs both.

Typical setup

The standard build is Docker Compose on a home server with a scanner sharing the consume directory over the network, a database and a media directory on persistent storage, and OCR configured for the user's language. Documents arrive by scan or by email and are tagged by rules, and the archive is reached from a browser on the local network. The database and media folder are backed up together, since neither alone can restore the archive.

Who should look elsewhere

Not for anyone who will not keep scanning. A half-populated archive is worse than none, because you will search it and conclude the document is missing. It is also a poor fit if your paperwork is primarily handwritten or low-quality scans where OCR cannot produce searchable text, and if losing a single document would be intolerable without a separate archive.

Project health

  • GitHub stars: 46,199
  • Last code push: 2026-10-01
  • Open issues: 4
  • Status: actively developed

Figures pulled from the GitHub API and refreshed periodically.

Paperless-ngx as an alternative

More in Productivity & Notes