Glad it helps! There are 4 key steps that I took:
- Upscaling (using Upscayl[0])
- OCR (using tesseract[1])
- Indexing (using Algolia[2])
- Scaling the processing and running on AWS (Klotho[3] - our startup)
I wrote a more in-depth blog post about it[4]
[0] https://github.com/upscayl/upscayl [1] https://github.com/tesseract-ocr/tesseract [2] https://www.algolia.com/ [3] https://github.com/KlothoPlatform/klotho [4] https://www.alashiban.com/search-the-deck/