Maybe you could use this for when there are "enough" features, and keep using the manual method for the remaining situations. I'd have to take a look at the images to have a better idea of what you are dealing with.
A bigger issue is that the project is running on a Pi, so just getting the image alignment running along with our face detection wouldve also been a pretty tall task if we want to stay at around ~5fps.