Paperless-ngx
Scan, index and archive your paper documents
What is Paperless-ngx?
Paperless-ngx turns your physical documents into a searchable online archive. It performs OCR on scanned files, automatically tags and classifies them, and lets you find any document by content. A community-maintained fork of the original Paperless project.
Best for
Households drowning in paper bills, contracts and receipts
Why choose Paperless-ngx
Paperless-ngx is the rare self-hosted project that changes a physical habit. You scan a document — a bill, a contract, a medical letter — and from then on it is text-searchable, tagged, filed and retrievable by any word it contains. The archive is on your own disk, encrypted or not as you choose, and it never gets sent to a document service. For anyone with years of paperwork and no desire to keep filing cabinets, the payoff starts on day one.
Replaces
- Evernote
- Dropbox paper workflows
Key features
- OCR with multiple languages
- Automatic tagging and classification
- Full-text search
- Email and scan ingestion
What to watch out for
Getting documents in is the awkward part. Quality depends on your scanner and on OCR accuracy, and handwritten or low-resolution scans will produce text you cannot search. The machine-learning classification and document matching work well once trained but need a decent volume of examples first, so early results are underwhelming. Consumption directories and email ingestion have surprising edge cases, and a poorly named file can create a duplicate document that you then have to reconcile by hand.
How to deploy
- Docker Compose
Getting started
Deploy the Docker Compose stack and decide your folder structure for originals, consume and media before importing anything — reorganising later means moving files the application tracks by path. Set the correct OCR language and the right date-parsing rules for your locale early, because both affect every document you add. Point your scanner at the consume directory over a network share and scan a batch of twenty boring documents to learn how names and tags come through. Take a proper backup of the database and media folder together; a restore needs both.
Typical setup
The standard build is Docker Compose on a home server with a scanner sharing the consume directory over the network, a database and a media directory on persistent storage, and OCR configured for the user's language. Documents arrive by scan or by email and are tagged by rules, and the archive is reached from a browser on the local network. The database and media folder are backed up together, since neither alone can restore the archive.
Who should look elsewhere
Not for anyone who will not keep scanning. A half-populated archive is worse than none, because you will search it and conclude the document is missing. It is also a poor fit if your paperwork is primarily handwritten or low-quality scans where OCR cannot produce searchable text, and if losing a single document would be intolerable without a separate archive.
Project health
- GitHub stars: 46,199
- Last code push: 2026-10-01
- Open issues: 4
- Status: actively developed
Figures pulled from the GitHub API and refreshed periodically.