Paperless-ngx – Open source document management system
nerdyarticles.com
nerdyarticles.com
With DT, say you’ve scanned or saved 20 docs to your inbox and you want to sort them to their long-term homes. DT will suggest folders based on how closely the new file matches the contents of those folders. It has the UI equivalent of “this looks like 2023 state taxes. Is it? This looks like kid #2’s school stuff. Is it? This looks like the older dog’s veterinarian records. Is it?”
That’s so, so nice.
Lately, as an experiment, I’ve been playing with organizing my docs with Johnny Decimal, then using the Hazel app to sort known docs with fixed structures (think bank statements and the like) into the right folders. My ScanSnap scanner’s software does OCR, so by the time docs land in the inbox folder, they’re ready for automated processing. It’s working pretty well so far, and I may stick with it.
But if I were to go back to an app, it would be DEVONthink or something with most of its features. That classifier is too darn nice, plus its smart rules, plus its scriptability, plus multi-device sync, plus Markdown notes with wiki links to stored docs, plus a thousand other niceties.
I've never used DT, so it's possible that their system is substantially better in some way.
I’d still miss DT’s zillion other things I’ve used over the years, but that one would have been a dealbreaker.
But I’m still experimenting with just plain files in iCloud. 99% of my mobile file access is to pull up scans of vaccination cards and things like that. It’s basically read-only from my iPhone (and iPad while I was using it). My phone’s Files app is full capable of handling that use case for me.
I scan into a Samba share that paperless-ngx picks up automatically, OCRs, tags, and deletes.
A web application is pretty cross platform too, at this point.
Plus I can get to them on my phones with less trouble than a share.
How does it handle when you have digital documents you want to store (a la google drive or similar)?
I have a Fujitsu ScanSnap which is one of those feed-through scanners. I have it hooked up to a Raspberry Pi which listens for the button press on the scanner. You press the button, the paper feeds through the scanner and once it has finished the scan a script runs to collate everything into a PDF and drops the result onto a Samba share that's running on the box where paperless-ngx is.
It's pretty neat and feels seamless. The worst part was dealing with SANE and finding linux drivers for my scanner.
https://chrisschuld.com/2020/01/network-scanner-with-scansna...
I've switched my hand-crafted scripts recently to use scanpdf[1] which seems to give better results (once I tweaked it to be a little less eager to downconvert to B+W). I experimented with using OpenCV models for cropping and straightening (based on examples in a stackoverflow thread at [2]) but I found results were worse than scanpdf so far.
1. http://badge.fury.io/py/scanpdf 2. https://stackoverflow.com/questions/28935983/preprocessing-i...
What does listen for button press mean? and how?
This works fine on say, a Mac, with the official Fujitsu ScanSnap software, and I'm guessing _that_ supports saving to a samba share, but I wanted a solution that's
1. completely headless, i.e. no desktop machine required and experience needs to be friction free as the headless part means the only way to interact with the scanning function is to press 1 button
2. linux compatible, as I wanted to connect it to a Pi. I had to dig for the drivers, Fujitsu didn't have the right ones for my model on their website!
I couldn't find any official software from Fujitsu, but I found the drivers eventually, so ended up coming up with connecting the scanner to the Pi over USB and glueing the bits together to drop the PDFs onto the samba share
The button is located on the scanner, and I run "scanbd" [1] to listen for the button press, this is what coordinates the scan function (feeding the paper through) and then post-scan -> running a script to collate + create PDFs
I run this locally on my workstation and send PDFs many times a week from Brother ADS2800w scanner via SFTP. Paperless NGX has reduced my home office paper piles to almost zero. It is a fantastic open source project and I am very thankful it exists.
They will also automatically launch if you have docker running at boot. Is it just because you prefer redhat/IBM's docker equivalent stack to the much more common and cross platform docker install?
I've been using docker compose in production for a couple of years now and it adds another layer on top of systemd that is a continuous source of headache, especially during updates.
Podman gets it right: no central daemon, can automatically generate systemd services for a whole pod. Updates are seamless.
This by itself is enough of a reason to me.
That is a lot of dependency. How stable is Paperless with all those applications making uncoordinated changes on their own schedules?
Tika and Gotenburg are additional features for scanning and converting MS Office documents to PDF. Not necessary and I don't use them in my setup at all. Same with sftpgo. I'm not sure for its usecase. But paperless doesn't directly depend on it in anyway.
Entirely possible the instability was my fault but the error messages didn't make it obvious what was wrong.
I'd probably still use it since there's awesome tools like this to set it all up, but it seems like a confusing architecture.
Just recently started working on an iOS/macOS app for it. Hope you like it!
One question though, my paperless-ngx is behind an SSO login (I use Authelia) with 2FA. Would it be possible to make your app work with that?
If you'd like some help testing this out let me know. Email is on my profile.
I want to add custom metadata to documents by categories/tags/folders, for example like this:
Invoice {issued: date, invoiceNumber: string, amount: number, due: date}
Contract { validFrom: date, renewsAt: date, autoRenew: boolean}
When adding a tag like this, it should either automatically fetch this information from the content document (probably very hard) or give you a manual workflow to type it into a form, while showing the document next to it. Maybe just by selecting the text from the PDF.In the folder list and in the search you would be able to add those meta data information as columns, sort them by value or do queries (tag:invoice AND invoice.amount > 1000)
Edit: this feature seems to be one of most upvoted feature requests for paperless https://github.com/paperless-ngx/paperless-ngx/discussions/1...
That said, open source is absolutely table stakes in this, to me. From the documents I have in the system one could trivially impersonate me. Perhaps even as good as clone me. So sending all that off to random internet corporations, no can’t do.
But yeah, let's connect. Take a look at my project as well! https://turtledev.net/projects/refind-ai
The documentation itself is so full of implementation details that, as someone who is interested in the concept of this, I'm scared off even trying to setup and use this
The project would be much more approachable if there was a simple native installer. My parents could also benefit from this but there's no way they would ever even understand how to install this, much less troubleshoot docker things.
It looks to sit in the self hosted space that has an admin manage all the sysadmin tasks. They’ve provided docker which is a pretty good step.
There are desktop apps designed at the single user/less experienced user, which might be more suitable
I've been using it for over a year and am very happy with it, though I intend on moving it from my home Pi docker swarm onto a free Oracle cloud instance to improve the performance and uptime (I've got my Pis auto updating and rebooting, so services get shunted around fairly often).
Still an overly complex FOSS user interface for a tech-unsavvy target with lots of digging around to configure it (OCR setup, for instance[2]), but at least you don't need to know what Docker is to install it.
1: https://www.lesbonscomptes.com/recoll/
2: https://www.lesbonscomptes.com/recoll/usermanual/webhelp/doc...
Actually the very first example on https://docs.paperless-ngx.com/setup lists an interactive installer which asks the user some question and eventually arrives at a working docker-compose setup.
$ bash -c "$(curl -L https://raw.githubusercontent.com/paperless-ngx/paperless-ngx/main/install-paperless-ngx.sh)"
If you ask me, this is already pretty user friendly. Although I agree that if your needs are more involved, there is some reading you'll have to do.I am currently in the process of migrating from mayan-edms to paperless-ngx and it feels pretty approachable to me if you know your way around docker (compose).
Like a sibling commenter is already doing, for example: https://news.ycombinator.com/item?id=37802337
Consider pitching in time/money there (if welcome) instead of complaining when not everything is served to you on a silver platter.
A sqlite backend would be another thing that could reduce the complexity for a minimal OoB setup, I guess.
Truth the told, I wouldn't find it unfair to call the architecture/setup "insane". This is not one. If you've done any meaningful self-hosting in the past decade, it's as straightforward as it can be.
https://github.com/geimist/synOCR/wiki
English translation: https://github-com.translate.goog/geimist/synOCR?_x_tr_sl=au...
I've been using this for several years and it works great.
Right now I am looking at OpenProDoc [1] and bitfarm-archiv [2] as document management possibilities.
[1] http://jhierrot.github.io/openprodoc/Spec_EN.html
[2] https://www.bitfarm-archiv.com/document-management/features....
It's a bit "scary" since even documents I delete in paperless-ngx are thus preserved forever, but it may come in handy someday.
What I find most useful is grouping together holiday documents such as travel insurance, holiday booking details, passports etc. and assigning a suitable tag, so I can easily find the relevant info. You could easily replicate that by copying those documents into a separate folder for easy access, but with Paperless-NGX it does most of the organisation for you and the search is more flexible as you can specify what kind of document you're looking for and who it came from.
I’ve killed my instance twice now and had to restore from backup, which is also surprisingly pleasant to do. Their document exporter makes that possible. Having everything in a single JSON and otherwise just the raw PDFs makes a ton of sense and has me confident my documents are “just there” and moving to a different system would be feasible.
>I’ve killed my instance twice now and had to restore from backup, which is also surprisingly pleasant to do
Stable, but murderable?
As one of the least digitized countries in Europe, and the digitalization budget recently cut 99%, it seems like they still need to use paper in their lives, and it's not gonna improve soon.
This feels so incredibly archaic to me as a Norwegian, I would have to print out documents to have anything to fill paperless-ngx with.
But yeah, then there are for example the bank account contract updates which come by physical mail only.
> expense notes (...) I throw everything away once it's handled.
Don't know about your location, but I need to keep the tax related documents for 5 years in case of an audit.
Getting all my invoices from last year to prepare taxes is now just a simple query in the paperless UI, the result would be about 95% digital and 5% physical documents, probably. Of course I could do all that old-school using filesystem folders, but having all my documents indexed and searchable in a single place was definitely worth the (small) effort of setting it all up and keep it running.
I just add all purchases/sales right when they happen in my accounting app and attach the invoice PDF. Then when I have to file taxes, I export the correct numbers.
Are you doing your bookkeeping in Excel or something?
Also it was just meant as an example, paperless is generally useful (to me) in situations where I need to access somehow related documents, like traveling and such, or searching my documents for some information. As I said, there are other systems and ways to do this, but for me this is the one that stuck.
I'm sorry to tell you that is a an oversimplification and especially for documenting expenses as a company/freelancer it's kind of worse.
Last time I checked if you want to follow the tax law to the word you're not allowed to change the medium:
If an invoice came as a paper copy (e.g. by snail mail), this paper copy is the original. If you scan it the digital version isn't.
If an invoice came as a digital document (e.g. a PDF by email), this digital document is the original - a printed version of that digital document isn't.
So if a tax inspector asks for "originals" it's technically almost impossible to provide them in the sense of the law. If even a tax inspector would care is another question.
And yes, they care about those rules and that you provide "originals" according to that definition - in particular that you didn't modify digital documents in any way. You can (and should) comply with that and there are service providers to help if you are to small to set that up yourself.
I admit that my last paragraph was kind of hyperbole, but I never heard (at least from other freelancers) of a tax inspector which wasn't happy with either everything printed or everything digital. I guess they really start to care if they suspect something fishy.
I think more than a few of these projects are started and/or maintained by Germans due to the astonishing number of documents received - e.g., paperless-ng appears to have been done by a German, although neither the original Paperless nor Paperless NGX immediately appear to be.
Of course, I still quickly download my year-end bank/salary/mortgage statements and cross-verify the tax departments numbers. The whole process takes at most a few hours.
IME Germany has significantly more hard-copy requirements.
But technically you're expected to keep the documents at least until you receive the "definitieve aanslag" and if you're nitpicking I think there is a 7 year term for the tax services to come back on your filed taxes and change things or demand proof.
Practically that doesn't happen if you accepted their pre-filled numbers and they match your employers. But if you're a freelancer or other non-standard case I would keep digital copies for a few years just to be sure.
Ah, interesting. I just assumed my situation never triggered it.
But when I was a freelancer I used a document scanning system provided by my bookkeeper. It worked similar to this open source thing, scan to PDF, automatic OCR and classification. Needed it because many invoices still arrived on paper, and receipts for restaurants etc I usually took a picture to upload.
I recently purchased a house. As part of the process, I needed to apply for a mortgage. The bank wanted a statement from my employer about my income from them, along with my last 2 complete years tax documents.
The bank had an inquiry. My employer had said my salary + bonus was X, but in the first of these two years, my tax documents said my income from my employer that year was 2.5X. The extra 1.5X was due to the employer being bought out and some change of control terms in the RSUs causing immediate payout of what would normally have been paid out over 4 years. Since I kept the documents of the RSU terms and the payslips, I could provide these to the bank to clear the matter up.
Notably, had I not kept my own copy of these documents, I could not have gone back to my employer for new copies. Due to the change of control, they had changed payroll vendors, and had eventually terminated the contract with the old vendor, so I could not have gotten a payslip from 1.5 years ago. Similarly, in the move to the new owner's HR system, the company had lost many of their records of agreements with employee's, including contracts etc., so it's not clear they would still have the terms of the RSUs, especially since the change of control payout rendered this a "completed" transaction. And later events made it clear that they did not have, e.g. a copy of my employment contract.
Similarly, if I ever had had a dispute over the terms of those contracts - if I hadn't kept a copy of the contract, and the company definitely hadn't kept theirs, any dispute would have been my word against theirs.
In a real contract dispute your copy of a contact from your documents isn't notably different in the eyes of the court than one from your employer. They're both notarized and if there's a dispute between them there is established processes. Aside from some titles or etc., historical filing ownership is typically relegated to the document originators.
Seriously, a couple cases of “sorry, I don’t have proof to back up that tax deduction” or “hey, here’s the receipt proving that our TV is still covered by warranty!” make it all worthwhile.
services.paperless.enable = true;
In contrast to Docker, it composes well with other things on the same system (e.g. the Samba server already set up to receive documents from my scanner), and I get automatic security updates for it.We had this ridiculous money-making scheme from the Federal Government a few years back called "Robodebt", which basically amounted to sending demands to pay back welfare money that they claim was incorrectly claimed ... long after the retention period for the paperwork. Conveniently they declined to show their working for the so-called incorrect claims.
Most people didn't keep their timesheets and records, and then couldn't prove they weren't overpaid, and were supposedly liable for the "debt".
Some people comitted suicide over it.
If you had your records, you probably wouldn't have lost any sleep over it, much less your life.
I've only added documents I've received this year (plus a couple of dozen documents going further back), and I've got ~250 in there, with a total of ~2.5m words (although I think word-count is a fuzzy concept in German).
I've posted a top level comment in more detail, but yeah, it's helpful to me.
Some correspondences take years and only add a mailing every few months. You would like to have a thread-like view -- as in an electronic mail. That is the strength of document management systems.
The other end of the spectrum in USA is filling with the 1040-EZ which is like a 3-4 document process.
That's the Marie Kondo approach. I feel the same way. I'm not sure that digitizing everything really removes clutter. It removes the physical paper, yes. But not the mental overhead of knowing that you have all those documents.
Some people obviously have a need to retain documents for things like business expenses that will be deducted from income. I don't have any of that. I get my W2s and 1099s and do my taxes. I throw all that in one folder and put it in a box in the closet. That's good enough for my purposes; I see no need to expend the time and mental energy necessary to scan and tag (even automatically) every receipt, utility bill, and other statements I receive. Why? I'll never look at them again.
And there are some important documents you want to keep longer than that -- birth certificates, wedding certificates, death certificates, documentation around the house you buy, documentation around your car, any personal health documents you keep, etc....
So, you'll want to have electronic copies of all that, and ideally backed up to an offsite location, in case you have a fire in your house. If you at least have electronic copies of those documents, even if you lose the original document in a fire, there's a good chance you can use that information to get new certified copies sent to you by the appropriate city/state government in question.
So, yeah, this can become a thing that needs to be carefully monitored and managed over the decades.
If I want something done by a third party i.e. my bank ill send them a mail or setup a meeting with a representative. This creates documents, that i want to add to paperless. Over the course of the next weeks additional documents will be created that outline the "progress" of the topic. I want all of the documents for this single issue im trying to solve to be kinda linked together, so that i always know the current progress and where it all started.
Does anyone here found a good way to solve this issue or is there even a built in function from paperless that i did not see ?
A bit more contrived than copying the volume but you don't need to shut down the server. There's probably some scripts out there for doing this in a structured way but I usually do it more or less manually/use a bash script.
[0]: https://www.postgresql.org/docs/current/app-pgdump.html
docker compose down && docker compose run backup && docker compose up -d
The restore procedure is the same, you restore the composefile through restic on the host and then `docker compose run backup restic restore latest --exclude "/data/self/*" --target /`
I find it's fast enough because restic is incremental, but if you can set this up on a filesystem with snapshots that would be a great option too.
Restic takes a bit of fiddling around too. I mount a prepared ssh config, a known hosts file and a private key.
What I have to do currently is use one of those document scanner apps on my phone to take a picture. It will give me a reasonably good pdf doc from that picture, but then I got to press multiple buttons to eventually email the doc to myself. Then I have to go into my mailbox and sort the doc into my google drive. Ultimately, it works, but also very annoying.
Source: happy user for 3+ years. No affiliation.
I'd love it if I could also use my mobile devices to bring up paper docs instantly (mobile phone, tablet, kindle).
On my mobile I use Firefox, and for me it works great.
I traveled for a month in Europe this summer and never had any issues.
I’ve found more good information about Paperless right here in the comments than anywhere else so far.
While the folder organization criticism in the article is on point (although you could also use tags that many file systems support, but that's not a reliable system to invest time in, or maybe if it's backed up by some app that can restore all the tagging it could be), the range of native tools for viewing/editing various document formats as well as your ability to customize your workflows in unparalleled
I really really appreciate the work that went into paperless, but for us the business risk of self-hosting this is far too high because if we lose our docs we lose our tax proof.
Paperless-ngx is a great example of how open source can work around the issue of lack of maintenance (quite a big problem with commercial services if they don't get much profit from them) as IIRC Paperless-NG was a fork from Paperless and now Paperless-NGX is a fork from Paperless-NG. Even if it becomes completely obsolete, I can choose to keep an instance running for as many decades as I wish.
Angular is an under-appreciated, solid, no-gimmicks framework. Been using it for years rather than React and it seems the the pendulum is swinging back toward "this side" now.
My only complaint is I've had the odd letter get scanned upside down and there's no way to rotate pages in Paperless-ngx.
>This app isn't available for your device because it was made for an older version of Android.
“Stack is only available on Android in the U.S. You can install it through the Google Play store.”
[1] https://techcrunch.com/2023/01/25/google-spares-three-area-1...
It's a self-hosted application, so it depends on your setup. I suppose it's arguably using local storage on the server you run it on which is often going to be a cloud hosted machine.
I've purchased a new-build home in Germany, and I'm currently in the stage between "purchased" and "ready for move-in," and if you've ever purchased a Neubau in Germany you know how much paperwork is involved - I get so many documents over email, many of which are scanned (to preserve the wet signature and stamps), and some of which I need to copy into a translator, that this is incredibly helpful. It checks my email, grabs PDFs, straightens them, OCRs them, adds a correspondent, tags them, and makes them available through a web UI.
I also appreciate the full-text search (for all that it might struggle if I had tens of thousands of documents) as I've had to go and try to find particular documents where the name of the document I've received might be a synonym for what the other person is asking for, but the word they're asking for is at least used in the text.
I'll also set it up to pull documents from my NAS as well, where the scanner writes to, as I also receive a number of documents via mail (that I also occasionally need to translate or copy/paste from).
There are also some limitations that annoy me:
* I really wish the email filters were more flexible - right now, I have to have three filters, one of PDFs, one for JPEGs, and one for PNGs, so I wish I could just set a regex for the attachment name. This one annoys me enough that if I ever have time I'd look at doing a PR for it (assuming the filtering is done locally and not on the IMAP server). * I'd also like to be able to setup rules to tag documents based on the email domain (e.g., house-builders get tagged as "house-builder, house") without having to manage a gigantic explosion of rules. In theory the ML should handle that, but... I'm mistrustful of ML. We'll see in a few months if I was too hasty in my judgement or not. * I'd like to retain slightly more information about the correspondent, like both name and email address (there's no consistency about who has their From line as "Name <email>" and who's just "email", even within the same company), both for de-duplication of correspondents and domain-based searching. * I wish I could share documents more easily than downloading it and re-uploading it to my email client (or mounting the folders and trying to find the right document, but that has its own set of problems). This one of those problems that's really easy to state, but potentially quite difficult to actually implement - could a web application add a PDF to the clipboard in such a way that GMail, say, would understand what was happening and add it as an attachment when pasted?
Overall though, I'm pretty happy with it, and finding it useful so quickly was somewhat surprising.
https://clis-everywhere.k8s.best/16
My approach is scanning the documents with airscan1, indexing them with a custom OCR Server (using the MLKit by Google on an Android phone which does completely offline OCR scanning) and indexing everything in OpenSearch. I've then created a backend + frontend to see the documents and di full text search with that.
Everything is (going to be) open source with a permissive license.
Funny situation because it's two extremes, we in Switzerland have way too much paper and need to find a solution, especially for tax filling. Opposite people are complaining they'd have nothing to feed it with.
I'm still on my way to have a proper document management solution for my needs, so that I can retrieve any correspondence of the last X years.
Finding an apartment contract or some other important mail is worth more than just being able to pay the bills "automatically" - since you sort of can already do that with eBill or the QR Bills.
OCRing everything is the first step towards this goal :)