Handwritten Text Recognition for Xournal++ Using Deep Learning
github.com
github.com
What you're talking about would be "online" handwriting recognition, where timing information about each stroke is available.
As such, my stroke-order-aware attempt over at https://github.com/PellelNitram/OnlineHTR/ uses a dataset from 2000 with around 12,000 samples. Contrary, the internal Google dataset is reported to feature around 16,000,000 samples :-D.
I have developed another model however (based on a somewhat recent Google paper by Carbune et al. 2020), that operates on pen dynamics and thereby implements online HTR, see here:
https://github.com/PellelNitram/OnlineHTR
This model is open-source as well and will be part of the HTR system for Xournal++ in the future. Feel free to give it a try yourself locally.
One question that has been bothering me a long time and prevented online HTR so far for me is how to find text on a page in temporal domain (i.e. in online domain and not offline domain). If you have any ideas on that, please do let me know as I would greatly appreciate that! One possible way is a transformer model - but again that feels a bit overkill and introduces a context length.
Currently, the machine learning model only supports offline HTR (i.e. using images) but online HTR (i.e. using pen time series data) is in the making, see here:
Xournal++ is a great project that features a bunch of really great developers who dedicate a lot of time to it.
I am not much involved in the Xournal++ development itself but then try to utilise my machine learning skills to build an HTR system for Xournal++ in the form of a plugin.
Cheers! :-)
Great to hear that you found out about Xournal++! It's really the power of open-source that you own your own handwritten notes as you are not locked in.
BTW personally I use Xournal++ to add text/images to pdfs, typically where I have some crappy low importance pdf-form, not a real one, and I do not want to invest time in a nice LaTeX + cart.el (Emacs artist mode wrapper to get coordinate of any form clicking with the mouse on them [1]). I still have to do with some scanned documents but originally printed from a computer not handwritten.
Handwritten text recognition might be very welcome to scan and index old public archives, witch is damn complex since there are countless of style of cursive, but it's still a very needed thing to merge the old paper world to the digital one not to loose history.
The HTR feature in Xournal++ does not require you to scan anything though. You just write handwritten notes in Xournal++ as always and upon saving with the plugin, the resulting PDF is searchable (the handwritten texts at least). So there's no scanning or retyping involved :-).
These are the decomposed prime time metrics:
1. The plugin itself is at 10/10; e.g. I use it in production myself.
2. The prediction quality is at, say, 6/10. I am actively working on this.
3. The installation process is 3/10. That's because I have only tested it using my machine setup, for which it actually works fine. To improve here, I'd love to get a bit of user feedback to improve the installation process.
Cheers!