Academic papers, standards documents, white papers etc... get downloaded as PDF and renamed according to a (human-readable) standard scheme that I use (year, title, [author(s)], [paper-type]), and stored in a 'meaningful' filesystem hierarchy. Podcasts and some videos are also stored in the same hierarchy.
Documentation for the software libraries and APIs that I use is downloaded from readthedocs (where possible) and stored in a parallel system that takes account of versioning. (So I can concurrently store differently versioned copies of the documentation for a single library).
I have a simple python script that iterates through my directory hierarchy and produces a sqlite database and a couple of xlsx files with a row-per document (one spreadsheet for reporting document-management metadata and another that allows me to assign labels and write precis notes). The script also extracts the content of the PDFs as plain-text and feeds some NLP tools that I'm playing with.
I use the spreadsheets to keep notes on the files, and the act of manually renaming and sorting the PDFs into the 'right' place in the hierarchy helps me to understand what's in them and remember what I've got. (I'm constantly reorganising the hierarchy as my understanding develops and evolves. My python script keeps everything -- notes and documents and other metadata - in sync).
So far this has scaled OK to around 26,000 PDFs.