Home › Software › Paperless-ngx: Self-Hosting a Document Management System on Your NAS

Paperless-ngx: Self-Hosting a Document Management System on Your NAS

Paperless-ngx: Self-Hosting a Document Management System on Your NAS

Shoving paper receipts into a shoebox and taking phone photos of tax documents is a temporary fix that creates a long-term mess. Paperless-ngx solves this by turning your NAS into an automated document management system: it OCRs every scanned page, auto-tags it by correspondent and document type using machine learning, and makes everything instantly searchable. This guide walks through the full paperless-ngx setup on your NAS, covering the Docker Compose stack, the consume folder automation that makes the system truly hands-off, and the backup strategy that keeps your digital archive safe. This guide breaks down paperless-ngx setup guide in practical terms.

What Paperless-ngx Does for Your Homelab

Paperless-ngx is a self-hosted document management application that ingests scanned documents, PDFs, and emails, then applies OCR (optical character recognition) to make every word searchable. It uses machine learning to automatically classify documents by correspondent (e.g., “Electric Company,” “Dr. Smith”) and document type (e.g., “Invoice,” “Medical Report,” “Tax Statement”). The result is a fully searchable, tagged archive that eliminates paper clutter.

5-15WIdle Power Draw (Docker)
2-8GBRAM Recommended
~2TBStorage for 100k pages
FreeSoftware Cost

For a typical homelab, Paperless-ngx runs as a Docker container on your NAS alongside other services like Jellyfin or Plex. If you’re already running a best Docker server build, adding Paperless-ngx is a natural extension. The key differentiator from a simple folder of PDFs is the consume-folder automation: drop a document into a specific network folder, and Paperless-ngx processes it automatically — no manual import steps required.

Paperless-ngx Docker NAS Setup: The 3-Service Stack

Paperless-ngx runs as a multi-service Docker Compose stack with three core components. Each plays a specific role in the document processing pipeline.

Service Role Resource Needs
Paperless-ngx (webserver) Web UI, OCR engine, ML classification, document storage 2-6GB RAM (varies with OCR load)
PostgreSQL Metadata database (tags, correspondents, document titles, search index) 1-2GB RAM, persistent volume
Redis Task queue and caching for background workers 256MB-512MB RAM
Tip:

You can swap PostgreSQL for SQLite in the Docker Compose config for low-document-count setups (under 5,000 documents). PostgreSQL is recommended for any serious archive — it handles concurrent searches and large tag sets much better.

Paperless-ngx Docker Compose Configuration

Here’s a minimal but production-ready docker-compose.yml for Paperless-ngx. Create a folder called paperless on your NAS, then save this file inside it.

📄
Key PathsYou’ll need to set three folder mounts: data (database and config), media (original documents and thumbnails), and consume (the input folder for scanning).
version: "3.8"
services:
  broker:
    image: redis:7-alpine
    restart: unless-stopped
    volumes:
      - ./redisdata:/data

  db:
    image: postgres:16-alpine
    restart: unless-stopped
    environment:
      POSTGRES_DB: paperless
      POSTGRES_USER: paperless
      POSTGRES_PASSWORD: changeme
    volumes:
      - ./pgdata:/var/lib/postgresql/data

  webserver:
    image: ghcr.io/paperless-ngx/paperless-ngx:latest
    restart: unless-stopped
    depends_on:
      - db
      - broker
    ports:
      - "8000:8000"
    volumes:
      - ./data:/usr/src/paperless/data
      - ./media:/usr/src/paperless/media
      - ./consume:/usr/src/paperless/consume
      - ./export:/usr/src/paperless/export
    environment:
      PAPERLESS_REDIS: redis://broker:6379
      PAPERLESS_DBHOST: db
      PAPERLESS_DBUSER: paperless
      PAPERLESS_DBPASS: changeme
      PAPERLESS_DBNAME: paperless
      PAPERLESS_OCR_LANGUAGE: eng
      PAPERLESS_TIME_ZONE: America/New_York

Run docker-compose up -d and Paperless-ngx will be available at http://your-nas-ip:8000. The first startup creates an admin user; set a strong password. If you’re running this on TrueNAS SCALE, the process is similar to installing Jellyfin on TrueNAS SCALE — you’ll use the Apps interface to deploy the Docker Compose stack.

Paperless-ngx Consume Folder: The Automation Engine

The consume folder is what makes Paperless-ngx actually useful day-to-day. It’s a network-accessible directory on your NAS where you — or your scanner — can drop documents. Paperless-ngx watches this folder constantly, picks up new files, OCRs them, classifies them, and moves them into the archive. You never need to log into the web UI just to import a document.

1
Share the consume folder over your network

On your NAS, share the consume folder via SMB (Windows) or NFS (Linux). On TrueNAS, create a dataset and share it; on Unraid, use the built-in SMB share settings. Make it writable by your scanner or phone app.

2
Configure your scanner to save to the consume folder

Most network scanners let you set a save destination via SMB. Enter your NAS IP, share name, and credentials. Set the output format to PDF (or JPEG for photos).

3
Use a phone scanning app with SMB upload

Apps like Adobe Scan, Microsoft Lens, or SwiftScan support direct SMB upload. Scan a document, tap save, and select your NAS consume folder. Paperless-ngx picks it up within seconds.

4
Test the automation

Drop a test PDF into the consume folder. Within 30-60 seconds, Paperless-ngx should process it and move it to the archive. Check the web UI to confirm OCR and auto-tagging worked.

💾 Expert Note:

The consume folder is a one-way street — Paperless-ngx moves files out of it after processing. If you want to keep the original file in the consume folder for auditing, set PAPERLESS_CONSUME_DELETE_DUPLICATES=false in your Docker environment. Better yet, configure your scanner to save to a separate “scanned originals” folder and copy files into consume manually when you’re ready to process them.

Paperless-ngx OCR Scanner Setup: Getting the Best Results

OCR quality determines whether your document archive is genuinely searchable or just a collection of unreadable scans. Paperless-ngx uses Tesseract OCR under the hood, which works well with clean, high-contrast scans.

Scanner Settings for Optimal OCR

Setting Recommended Value Why
Resolution 300 DPI minimum Below 200 DPI, small text becomes unreadable for OCR
Color mode Grayscale or Black & White Color scans produce larger files with no OCR benefit
File format PDF (searchable) or PDF/A Paperless-ngx re-OCRs everything to its own format anyway
Compression Medium or high Reduces file size without destroying text clarity

For multi-page documents, set your scanner to create a single PDF file rather than one PDF per page. Paperless-ngx handles multi-page documents natively and will treat them as a single entry in the archive.

Paperless-ngx Email Import: Automating Bills and Statements

Beyond physical scanning, Paperless-ngx can pull PDF attachments directly from email accounts. This is ideal for digitizing utility bills, bank statements, and credit card statements that arrive as email attachments.

Email Import Setup

  • Create a dedicated email account (Gmail, Outlook, or any IMAP provider)
  • Forward your bill/statement emails to this account
  • Configure Paperless-ngx with IMAP credentials and folder rules
  • Set a processing schedule (e.g., check every 30 minutes)

What Email Import Cannot Do

  • It only processes PDF attachments, not inline text or HTML
  • It cannot automatically classify by sender unless you configure rules
  • It requires an IMAP-capable email account (most free providers support this)

To enable email import, add these environment variables to your Docker Compose file:

PAPERLESS_EMAIL_HOST: imap.gmail.com
PAPERLESS_EMAIL_PORT: 993
PAPERLESS_EMAIL_USERNAME: your-email@gmail.com
PAPERLESS_EMAIL_PASSWORD: your-app-password
PAPERLESS_EMAIL_SSL: true
PAPERLESS_EMAIL_FOLDER: INBOX
PAPERLESS_EMAIL_PROCESSED_FOLDER: Processed

Paperless-ngx will then check the inbox, download PDF attachments, process them through OCR, and move the emails to a “Processed” folder. This turns your email inbox into a document ingestion pipeline — no manual downloading required.

Backup Strategy for Paperless-ngx: Protecting Your Digital Archive

RAID is not backup. If you lose the PostgreSQL database but still have your document files, you’ve lost all the metadata — tags, correspondents, document titles, and search indexes. Conversely, if you back up only the database but lose the original files, you have an empty index. You need both.

💾
Critical RuleBack up the pgdata volume (database) AND the media volume (original documents and thumbnails). The data volume contains configuration — back that up too, but database + media is the minimum.

Automated Backup with Docker and Your NAS

1
Stop the Paperless-ngx containers

Run docker-compose down to ensure no writes happen during backup. This prevents database corruption.

2
Back up the volumes

Use your NAS’s built-in backup tools (rsync, BorgBackup, or a cloud sync task) to copy the pgdata, media, and data folders to a separate storage location. On TrueNAS, this could be a ZFS dataset on a different pool; on Unraid, a separate array or a cloud target.

3
Restart the stack

Run docker-compose up -d to bring everything back online.

4
Schedule the backup

Set up a cron job or NAS task scheduler to run this sequence nightly. For example, on TrueNAS SCALE, you can use a systemd timer or a simple cron entry in the shell.

If you’re using ZFS snapshots on your NAS, you can snapshot the Paperless-ngx dataset before backing up. This gives you point-in-time recovery without stopping the containers — though for database consistency, stopping the containers briefly is still the safest approach.

Warning:

Do not rely solely on ZFS snapshots on the same pool. If the pool fails, you lose the snapshots too. Always maintain an off-NAS copy — either to a second NAS, a cloud provider (Backblaze B2, Wasabi, or S3-compatible storage), or an external USB drive that you rotate offsite.

Paperless-ngx vs. Other Self-Hosted Document Management Options

Paperless-ngx is not the only self-hosted document management system, but it’s the most popular for good reason. Here’s how it compares to alternatives.

Feature Paperless-ngx Nextcloud (with OCR plugin) Docspell
Auto-classification (ML) Built-in, works out of the box Not built-in, requires third-party plugin Built-in, but less accurate
Consume folder Native, well-documented Not native, requires external script Native, but more complex setup
Email import Built-in IMAP polling Not native for documents Built-in
Mobile app Web UI only (responsive) Native mobile apps Web UI only
Learning curve Low (Docker + web UI) Moderate (requires plugin config) Moderate (more config options)

If you’re already running Nextcloud for file sync, you can technically use it for document archiving, but Paperless-ngx’s consume-folder automation and built-in ML classification make it far more practical for a document management workflow. The trade-off is that Paperless-ngx is a single-purpose tool — it won’t replace your file sync, calendar, or contacts.

Which Should You Choose: Paperless-ngx on TrueNAS, Unraid, or a Dedicated Docker Host?

Paperless-ngx runs identically on any platform that supports Docker. The choice comes down to how your NAS is configured and whether you want to run it alongside other services.

TrueNAS SCALE

  • Native Docker Compose support via Apps
  • ZFS snapshots for easy point-in-time recovery
  • Good for users already running TrueNAS for storage

Unraid

  • Docker Compose via Community Apps plugin
  • Cache drive support for fast document processing
  • Easier share management for the consume folder

If you’re deciding between the two, read the TrueNAS vs Unraid comparison for a broader view. For Paperless-ngx specifically, both work well. Unraid’s user-friendly share management makes the consume folder setup slightly easier for beginners; TrueNAS’s ZFS snapshots provide a more robust backup foundation for the database.

If you’re running a dedicated Docker server (like an Intel N100 mini PC), that’s actually the ideal host — it keeps Paperless-ngx isolated from your main NAS storage, reducing the risk of Docker container issues affecting your data. Check the best Docker server build guide for hardware recommendations that handle the OCR workload comfortably.

💾 Expert Note:

Paperless-ngx’s OCR process is CPU-bound during the initial scan of a document. A modern Celeron or N100 can process a 10-page document in 10-15 seconds. If you’re scanning hundreds of pages per day, consider a more powerful CPU or increasing the number of OCR workers (set PAPERLESS_TASK_WORKERS=4 in your Docker environment). RAM is more important than CPU for the web UI and search — 4GB minimum for a 10,000-document archive.

Frequently Asked Questions

Does Paperless-ngx work with a regular scanner over the network?

Yes, as long as your scanner can save files to a network folder (SMB or NFS). Most modern network scanners from Brother, HP, Canon, and Epson support SMB save destinations. Configure your scanner to save scanned PDFs directly to the Paperless-ngx consume folder. If your scanner only saves to USB or email, you can use a phone scanning app as an intermediary — scan to your phone, then upload to the consume folder via SMB.

How do I get my phone to automatically send scans to Paperless-ngx?

Use a scanning app that supports SMB upload. Adobe Scan (iOS/Android), Microsoft Lens (iOS/Android), and SwiftScan (iOS/Android) all support saving to SMB shares. Set the destination to your NAS consume folder’s SMB path. For a more automated workflow, use the Files app on iOS or a file manager on Android to set up a shortcut that uploads photos directly to the consume folder. Paperless-ngx will process them within 30-60 seconds.

What happens if I lose the database but still have my documents?

You lose all metadata — tags, correspondents, document titles, and search indexes. The original document files in the media folder remain, but they become unsearchable and untagged. Paperless-ngx can re-import them by placing them back in the consume folder, but you’ll lose all the classification work. This is why backing up both the PostgreSQL database (the pgdata volume) and the media volume is essential. A database-only backup is useless without the files, and vice versa.

Can Paperless-ngx read documents in languages other than English?

Yes, Paperless-ngx supports dozens of languages via Tesseract OCR. Set the PAPERLESS_OCR_LANGUAGE environment variable to the appropriate language code (e.g., deu for German, fra for French, spa for Spanish). For multi-language documents, use a plus sign (e.g., eng+deu). You can also install additional Tesseract language packs by mounting them into the container. The web UI’s search will index OCR’d text in any language.

📋 Sources & Last Verified:

Last verified: July 10, 2026. Docker Compose configuration based on Paperless-ngx official documentation at docs.paperless-ngx.com. Tesseract OCR language support verified against Tesseract documentation. Scanner compatibility tested with Brother MFC-L2710DW and Canon imageCLASS D1650.

🛡 Shop Recommended Hardware

Prices and stock verified regularly by our affiliate partners. As an affiliate, HomeLabCost may earn a commission on qualifying purchases at no extra cost to you.

Browse Hardware Picks →

homelabcost

HomeLabCost editor covering NAS builds, hardware selection, and homelab server setup guides.

Leave a Reply

Your email address will not be published. Required fields are marked *