r/Paperlessngx 1h ago

Asking for AI Configuration Help - Paperless 3.0 (openai - local endpoint)

Upvotes

The error, when I click SUGGEST on a document in PaperlessNGX3.0 -

URL

http://192.168.1.105:8000/api/documents/317/ai_suggestions/

Status

400 Bad Request

Error

{"ai":["Invalid AI configuration."]}

____

I'm using a local open-ai endpoint, and I get this when clicking on suggestions. When I read my endpoint (oMLX's) log, it looks like Paperless is hitting it successfully:

2026-07-23 08:57:25,988 - omlx.server - INFO - [-] - Embedding: 1 inputs, 2560 dims, 1058 tokens, max_length=40960, truncation=True in 1.271s

2026-07-23 08:57:29,900 - omlx.scheduler - INFO - [-] - Cache phase timings: cleanup_finished_sync=11.7ms/5, store_cache_main_boundary=0.2ms/5

2026-07-23 08:57:29,901 - omlx.server - INFO - [-] - Chat completion: 148 tokens in 3.74s (39.5 tok/s), prompt: 2854, finish_reason=stop, max_tokens=80000, request_max_tokens=None

2026-07-23 08:57:34,071 - omlx.scheduler - INFO - [-] - Cache phase timings: cleanup_finished_sync=14.0ms/6, store_cache_main_boundary=0.2ms/6

2026-07-23 08:57:34,072 - omlx.server - INFO - [-] - Chat completion: 199 tokens in 4.16s (47.9 tok/s), prompt: 546, finish_reason=stop, max_tokens=80000, request_max_tokens=None

My AI setup:

What might I be doing wrong here?


r/Paperlessngx 6h ago

One-liner Paperless V3 installer

26 Upvotes

Six months ago I posted about the document management setup I built for my family around Paperless-NGX V2, and later a one-command installer that stands the whole stack up on a fresh Ubuntu box.

Paperless-NGX v3.0.0 shipped, and one of the changes was folding LLM classification (title, tags, correspondent, document type, storage path suggestions) into the main app. So paperless-gpt, which was carrying that job for me, could finally go.

Did the upgrade end to end yesterday on my 1,805-doc library, ripped out the sidecar, wired v3's native AI up to Gemini, pushed the changes to the installer, and wrote up what I found.

The writeup (with the rollback plan and the one gotcha):
https://turalali.com/from-paperless-gpt-to-paperless-ngx-v3-dropping-a-container-cutting-complexity/

The installer (v3-native, one command on a clean Ubuntu box):
https://github.com/tural-ali/paperless-overconfigured

tl;dr for people looking at this exact migration:

  • v3's AI covers everything paperless-gpt did for me. Suggestions drawer on every document details page: title, correspondent, doc type, tags, storage path, dates. Same underlying LLM, one fewer container to keep updated.
  • Four config changes are forced by v3. Image tag pin, PAPERLESS_DBENGINE: postgresqlPAPERLESS_OCR_MODE: skip_noarchive (renamed from skip), drop PAPERLESS_OCR_SKIP_ARCHIVE_FILE. That's it. Everything else in my compose carried through unchanged.
  • The gotcha, if you want Gemini via the OpenAI-compat endpoint: every doc and tutorial names the embedding model text-embedding-004. That value returns 404 through Google's OpenAI-compat gateway. The name that actually works is gemini-embedding-001. Blog explains why (different API surface routing under the hood).
  • Migration timing: 11 minutes on my box (1,805 docs). Django migrations + SHA-256 recompute + Whoosh → Tantivy index rebuild. No data loss, all workflows + mail rules + storage paths intact.
  • Rollback: the migration doc explicitly has no downgrade path. You have to build your own with pg_dump and cp -al snapshots before you flip the image tag. I've got the exact commands in the writeup.
  • Cost on Gemini paid tier: ~$0.54 one-time for the initial embedding rebuild, roughly $2/month steady-state at ~20 new docs/day. Google AI Studio free tier absorbed the whole rebuild without a rate-limit hit.

The installer wizard picks between Gemini (OpenAI-compat), OpenAI (native), or Ollama (local) during setup, and pins the Paperless image to 3.0.0 so nobody accidentally does a major-version upgrade unattended via docker compose pull.

Happy to answer questions on the migration or the installer.


r/Paperlessngx 19h ago

Releases · paperless-ngx/paperless-ngx V3.0.0

Thumbnail
github.com
145 Upvotes

r/Paperlessngx 1d ago

We built Docpose.cloud — file conversion, OCR, PDF, archive, and email file tools with API access

Thumbnail
0 Upvotes

r/Paperlessngx 2d ago

Self-provisioning SMB inbox shares for multi-user `paperless-ngx`.

5 Upvotes

Hi there,

this is another part of my personal paperless-ngx setup: I’m using a SMB capable network scanner to push new documents right away to the inbox folder of my paperless instance.

This project is a lightweight Docker container meant to be ran side-by-side with paperless-ngx. It contains a samba/SMB service and automatically takes care of (de-)provisioning users from the paperless instance as well as creation of workflows for owner assignment in a multi user paperless environment:

https://github.com/flobernd/paperless-smb-sync

This project is a follow up to my earlier blog post in which I described my multi-user paperless setup:

https://blog.flobernd.de/2026/02/paperless-ngx-document-management/

The project is not enterprise ready (no SSO, Active Directory, etc. support), but might still be useful in smaller homelab setups.


r/Paperlessngx 4d ago

Any way of optimising this simple workflow?

2 Upvotes

Hi all,

I've recently mostly migrated from Dropbox to PaperlessNGX for the storage and sharing of documents (mostly receipts) scanned from my phone.

The process of scanning a receipt with my phone and having available a shared link to that file is a matter of 5 steps using Dropbox, taking about 5-10 seconds:

  1. Click Add photo
  2. Take photo
  3. Click checkbox
  4. Click Upload
  5. Click 'Link'. (Share link is now copied to clipboard automatically).

Using the Paperless app (for iPhone) it's 14 steps, and around 60 seconds per receipt:

  1. Click '+'
  2. Scan Document
  3. Take photo
  4. Click green checkmark
  5. Click Save
  6. Wait for upload and processing to complete (takes around 30 seconds)
  7. Click document
  8. Click Share icon
  9. Choose 'Share link'
  10. Click '+' icon
  11. Choose Expiration: never
  12. Click green checkmark
  13. Click new share link
  14. Choose 'Copy'.

Anyone got any clues or hints as to how to optimise this? Happy to look at alternative apps - some of which I've tried already - or include the use of a browser on my PC. Thanks!


r/Paperlessngx 5d ago

How do you handle ASN labels on documents that you send to other parties?

5 Upvotes

I have paperless ASNs set up with small QR stickers (that also contain the ASN in plain text). However, I am not completely satisfied with the use of those stickers on documents that I have to submit to other outside parties (e.g., course certificates).

If I place the sticker directly on the document (which should be the correct workflow for ASNs), the recipient might be wondering, what this weird QR code and ASNXXX is, that they have never seen before on such a certificate. So I would rather like to keep the actual document "pure" but still assign an ASN.

My current workaround is the use of a blank page with the QR code that is scanned before or after the actual document. In the beginning, I then transferred the sticker from the blank page to the actual document, but now I just use the blank QR page as a divider in the physical folder.

However, I am not satisfied with this solution, as it produces unnecessary, mostly blank pages.

How do you handle this case? Do you care that the QR code is on the page when you send it off to someone else?


r/Paperlessngx 5d ago

Chandra OCR for paperless-ngx (v3)

34 Upvotes

Hi there,

with paperless-ngx v3 (currently in beta), the team added a plugin system for ingest providers.

I created an example implementation of a provider that uses Chandra OCR (instead of Tesseract) for greatly improved text recognition:

https://github.com/flobernd/paperless-chandra

The provider uses ocr-my-pdf to create transparent text overlays for your PDFs (exactly like the stock Tesseract provider) and also honors all of the built-in paperless OCR settings / environment variables.

The project has been bootstrapped with the help of LLMs, but I carefully reviewed + tested the code during the whole process.

Let me know what you think!


r/Paperlessngx 6d ago

My family wouldn't use the Paperless-ngx web UI, so I built them a Telegram bot — now open source (heads up: the AI part uses Gemini API, not local)

36 Upvotes

I set up Paperless-ngx in my homelab a while ago for our family documents. It's great, but I hit a problem: my family. The web UI takes time to learn, and they found it inconvenient and basically never opened it. Meanwhile, we all live in Telegram anyway. So the idea was simple: make the day-to-day interface a chat.

Fair warning before the feature list: the AI part runs on Google's Gemini API, so this is not a fully local setup. More on that below.

I first experimented with the community Paperless MCP server (@baruchiro/paperless-mcp) plus Google's Antigravity agent SDK and was surprised how much that combination can do with documents. The next step was a small hobby bot the whole family can talk to, each member with their own Paperless API token, so permissions are enforced by Paperless itself, not by the bot. Unknown Telegram users are simply ignored.

What we actually do with it:

  • "List all contracts with Acme signed before 2020" — works in any language, answers in yours, with download buttons for the original PDFs
  • "When does my passport expire?", "What's the notice period in my rental agreement?" (it actually reads the document text)
  • "How much did we spend on utilities in 2025?" and then just "And compared to 2024?" — it remembers the conversation per user, so follow-ups work (/clear wipes it). This is something I couldn't do in the web UI anyway.
  • Send a PDF or a photo into the chat: it uploads to Paperless, waits for OCR, sets title/date/correspondent/document type/tags and writes a short note. If Paperless flags a duplicate, the bot links the existing document instead of re-uploading.
The GIF is a mock-up with sample data, not our real archive.

The honest part. Document text and your queries are sent to Google. A free-tier key from AI Studio works, but check the free tier's data-use terms (a paid key has different ones). Telegram bot chats aren't end-to-end encrypted either. For us this trade-off is fine, since we already share these documents in our family chat anyway, but if it's a dealbreaker, this project isn't for you (yet). Local model support could be a future direction if people want it.

One more caveat: the totals above are computed by the LLM from retrieved documents. Treat them as an assistant's summary, not accounting-grade numbers. That's what the download buttons are for.

The bot keeps no database of its own; conversation history lives in memory only.

Setup: a bot token from BotFather, your Paperless URL plus per-user API tokens, and a Gemini key. Docker compose (image on GHCR) or systemd. AGPL-3.0, just tagged 1.0.

Repo: https://github.com/rabestro/paperless-genie

Docs: https://jc.id.lv/paperless-genie/

My family actually uses it now, which was the whole point. The project is young and there's a lot of room for improvement. Would something like this be useful to anyone besides us? Any ideas, UX or otherwise, are very welcome.


r/Paperlessngx 6d ago

Built a one-way sync from Notion to Paperless-ngx (Rust, MIT licensed)

5 Upvotes

If you keep notes in Notion but archive everything else in Paperless, your notes are usually the one thing that's not searchable alongside the rest. I built notionless (https://github.com/Script-hpp/notionless) to close that gap: it watches a Notion database and mirrors pages into Paperless as Markdown, on a schedule, diffed by content hash rather than Notion's minute-rounded timestamps.

A few things that took some iterating to get right:

- Custom fields (notion_id, notion_last_edited, notion_content_hash) are auto-discovered by name at startup and created if missing, no manual Paperless setup needed.

- If a matching document already exists in Paperless (e.g. from before you started using this), it gets linked instead of endlessly re-uploaded and rejected as a duplicate.

- Config loads from ~/.config/notionless/.env or .env in the working directory, so it plays nicely with systemd services, not just cargo run.

Current limitations: sync is one-way (Notion → Paperless only), and only paragraph/heading blocks are exported so far, no lists/code/tables yet. Both are on the roadmap.

Rust, Dockerfile included, MIT licensed. Built with Claude (Anthropic's coding assistant); I drove requirements and did all verification against my own live instance. Repo: https://github.com/Script-hpp/notionless

Happy to answer questions or hear if the custom-field/duplicate-handling approach could be done better.


r/Paperlessngx 8d ago

Looking for an iPhone/iPad scanning workflow that gives crisp, noise-free text and true-to-original colors for highlighted notes

2 Upvotes

I've been using a few AI-powered scanner apps on iPhone (and sometimes iPad) to digitize my special notes (which have color highlighting) into PDFs, but I can never get the result I'm after: crisp, noise-free text and colors that actually match the original highlights. Everything comes out slightly soft, a bit grainy, and the highlight colors look duller or shifted compared to the real page.

A couple of questions for anyone who's solved this:

  1. Is there a scanner app for iOS/iPadOS that's genuinely optimized for AI-enhanced resolution and quality, especially one that captures color highlights accurately (true color reproduction, not washed out or over-saturated) and keeps text edges sharp and free of scan noise/grain?

  2. Is there a web app, website, or macOS tool that can take an already-scanned PDF and enhance/upscale it with AI afterward, sharpening text, boosting resolution, and cleaning up noise, without shifting the colors or messing up the layout?

Ideally I want a repeatable workflow (iPhone or iPad, whichever gives better results) where I scan and consistently end up with a clean, high-res PDF where the text is sharp and noise-free and the highlight colors are accurate to the original page. Any app recommendations, settings tips, or full workflows that have actually worked well for you would be really appreciated.


r/Paperlessngx 11d ago

What dont you like about SaaS document managing platforms?

0 Upvotes

I'm seeing paperless come up whenever I google "document management", which is weird because most people prefer to just use an existing solution than to self host. Do most web apps suck? Why did you choose selfhosting over a SaaS?


r/Paperlessngx 12d ago

Status of Paperless 3.0

65 Upvotes

The first beta/rc1 was released 2 months ago on GitHub (Paperless-ngx v3.0.0-beta.rc1). Does anyone here have insight into the latest developments? Are many fixes still pending?


r/Paperlessngx 12d ago

I built a self-hosted, bring-your-own-keys AI companion for paperless-ngx. Better OCR + automatic titles, tags, correspondents and dates

0 Upvotes

I've been using paperless-ngx for a while and the one thing that always nagged me was the metadata. OCR on scanned stuff was hit-or-miss, and I still ended up manually setting titles, tags, correspondents and dates on everything. So I built a thing to hand that off to an LLM and I'm posting it here in case it's useful to anyone else.

It's called Paperless Starfruit. The idea:

  • You tag a document in paperless with psf-process.
  • It gets picked up, (optionally) re-OCR'd, and sent to an LLM which suggests a title, tags, correspondent and date.
  • Depending on your setting, the result is either written straight back to paperless, or dropped into a Review queue where you approve / edit / reject each suggestion before anything touches your documents.

Things I cared about while building it:

  • Bring your own keys. Anthropic, OpenAI, Google, Mistral, or anything OpenAI-compatible including fully local Ollama / LM Studio / vLLM if you don't want documents leaving your network at all.
  • Self-hosted, single container. SQLite on a mounted volume, no Postgres/Redis to babysit. Pre-built images on GHCR, so nothing to clone or build.
  • Everything configured in a web UI paperless connection, provider keys, models, prompts, page limits, polling interval. The only thing outside the UI is a handful of env vars.
  • Credentials encrypted at rest (AES-GCM), never returned to the browser unmasked.

It's meant for homelab scale, simplicity over throughput.

It's open source and free. Docs and install walkthrough are on the site.

Happy to answer questions and genuinely interested in feedback, especially on the prompt/tagging side.


r/Paperlessngx 13d ago

Apple Mail Attachments and paperless-ngx

2 Upvotes

Hi,

I try to setup paperless-ngx and to consume attachments sent by apple mail. Unfortunately this does not work, sending the same (empty) mail with Proton-Mail paperless-ngx consumes the attachment, deletes the mail as expected.

Is there a special method to handle mails forwarded by apple mail ?


r/Paperlessngx 16d ago

Archi 2.0 — on-device scan + AI metadata straight into Paperless-NGX (iOS/iPad/Mac)

66 Upvotes

Quick update for the Paperless crowd: Archi is a native Apple app that scans a

document, runs OCR + AI metadata extraction fully on-device (Apple Vision +

Gemma), and files it into your Paperless-NGX with suggested title / correspondent

/ type / tags — which you review before upload. No cloud in the loop; the only

endpoint is your own server.

New in 2.0: native iPad & Mac, an offline archive with full-text search, and an

AI-transparency page. Uses your existing correspondents/tags as hints so it

matches your setup.

App Store, $3.99. I'm the dev — feedback very welcome, especially on the

Paperless-specific workflow.


r/Paperlessngx 20d ago

I built a Paperless-ngx companion for AI metadata and owner assignment — looking for workflow feedback

5 Upvotes

I have been working on Archivista AI, a self-hosted companion that reads Paperless OCR text and writes back a title, tags, correspondent, document type, date, language, custom fields, and optionally an owner.

The part I most wanted to improve was setup: connect an existing Paperless instance in a browser, choose Ollama or a hosted/OpenAI-compatible provider, then inspect history and manually re-run documents from the UI. It supports OpenAI Flex and OpenAI/Anthropic batch processing for lower-cost, asynchronous workflows.

It runs as one Docker container with SQLite for processing history and retries. The published image is `ghcr.io/arturict/archivista-ai:1.1.0`.

Repo: https://github.com/arturict/archivista-ai (MIT)

Privacy boundary: local Ollama/OpenAI-compatible endpoints keep classification on your network; choosing a hosted provider sends the OCR content needed for classification to that provider.

I am looking for Paperless-specific feedback rather than stars: should generated values be limited to existing tags/correspondents/types, and what would make optional owner assignment feel safe enough for a household installation?

Disclosure: I am the author. AI coding tools assisted with parts of implementation, review, documentation, and testing, and the app itself uses the configured model for classification.


r/Paperlessngx 21d ago

Has anyone gotten QuickScan to work with HTTP only?

1 Upvotes

I’m trying to get QuickScan to connect by HTTP over Tailscale, but it doesn’t seem to work. I get the NSURLError for App Transport Security. From what I’m seeing it’s to do with HTTP vs HTTPS, but the app does say it supports HTTP for lan connections.

There’s a thread about it on the GitHub but the dev didn’t really answer the question. Would love to use the app but can’t get it to connect to my instance.

Thanks!


r/Paperlessngx 22d ago

how can i use paperless-ngx to build a personal rag systerm

3 Upvotes

title says it all. I have already set up paperless-ngx, I was planning on setting up paperless-ai do ocr and rag but when I read the readme of the project, it said that it was no longer maintained, what should I do? is it worth installing or should i wait for the official implementation?


r/Paperlessngx 23d ago

consolidate tags

3 Upvotes

Does anyone have any ideas on how to use AI to consolidate similar existing tags in an automated fashion?


r/Paperlessngx 26d ago

Waiting for the 3.0 release to setup llm-based OCR?

23 Upvotes

I have used base ngx for about a year, and recently start to get interested in a better OCR, plus potentially chat/tagging bot with Paperless-GPT. Then I realized that v3.0 is about to be released, should I just wait for that?

Another question for people tried v3.0 beta: I don't have powerful hardware to run reasoning models, but enough for a lightweight OCR model, (like qwen3.5-0.8b or minicpm). So can I use ocr-model locally, but use cloud AI providers (like GPT5 api) for tag/chat bot?


r/Paperlessngx 27d ago

Not everyone can afford a local LLM

Enable HLS to view with audio, or disable this notification

0 Upvotes

I am living paperless since many years. With the rise of AI it would be nice to have a VLM for OCR, some good LLM for tagging, a strong RAG system for my whole library and an AI assistant which does all my paperwork, like replying to Mails and attaching all necessary documents.

I cannot afford running models locally and doing all the administrative tasks around it. I also don't want my document library hosted on any Big Tech company datacenters or any LLM context send to an US American AI company API Endpoint.

The logical next step for me is hosting my paperless-ngx on a EU sovereign cloud, three times redundant, with a backup location and always the possibilty to download a local backup.
There I have the protection by strong EU laws and it is not reachable by the CLOUD Act.

I have recorded a demo, how the assistant works. Additional demos can be found on paless.eu , like filling out forms automatically by using the documents context, tagging, OCR, etc.


r/Paperlessngx 27d ago

AI without redoing OCR in paperless with paperless-GPT or paperless-IQ ?

14 Upvotes

Hi,

I have been using Paperless-ngx for a long time, actually I started with paperless-ng before commiting to paperless-ngx.

I really love it, but wouldn't mind adding a little AI to auto-tags some of my documents, so it's easier for me. I want to stay local with Ollama.

I don't want/can't use my GPU for AI on this purpose, so I wanted to use my large CPU for this. The cpu can handle the text part with small models like qwen3 but even if capable of doing vision models, it struggle and can impact the server.

All (or 95%+) of my pdfs already have the OCR processed correctly and the content in Paperless-ngx is usually quite good for this. So I don't see a reason why I need to reprocess it in the LLM.

Is there a way to only process the auto-tags with the text from OCR pdf without the vision LLM in Paperless-GPT ?

Otherwise my solution is to wait or use the beta v3 ? Cause I also found a post from a month ago for : https://github.com/knows-cloud/paperless-iq
That's seems interresting but before going to start and test a new container, I wanted to know if there was a parameter or option I missed in paperless-gpt.

thank you !


r/Paperlessngx 28d ago

Slow Performance on m4 Mac Mini

3 Upvotes

I have the great task to digitize all documents at my workplace – and there are a lot of documents! Around 230 Leitz folders (if that's a measurement outside of Germany), some are full to the brim, some are almost empty...

My issue is: the performance of my paperless instance is quite slow. When looking at a 10-page PDF scanned with 200 dpi, it takes about 5 seconds to load the high-res version of the document. I'm running it on a Mac mini M4 10/10 with 16gb RAM, 1TB ssd with Postgres via Docker.

I have (so far) scanned 289 documents (1.6M characters) with ASN-Labels – that's about 5% of the whole thing, and the performance is already concerning to some point, and I don’t know what will happen when I start scanning more and more documents... Any ideas what the issue could be, if I'll be fine or what to do about it? I've heard that some people migrate to a different database cause it scales better? hm...

I'm also running Paperless-GPT and n8n for OCR and some handy tag-scripts on a different device, so there should be no compromise in performance... 


r/Paperlessngx 29d ago

How does batch scanning affect auto-filling in Paperless-ngx?

2 Upvotes

Does Paperless-ngx apply its matching algorithms to documents only at the moment of ingestion, or can I scan all my documents first and benefit from the auto-filling features later? Any insights or tips would be greatly appreciated!