Also, for low frequencies with lots of overtones look into autocorrelation, that works quite well and can be efficient if you program it right.
Then I switched to very simple autocorrelation and then I switched to a fancier autocorrelation-based method called "Yin".
Also, There's also a comment "The logarithmically binned spectrum has a resolution of one cent". Does that mean that there is a bucket for each 1 cent? When I was using WebAudio I think the the buckets were something like 14 units of frequency apart (hz?), and I didn't discover a way to get finer resolution.
So you need to run multiple methods in parallel and decide based on the very rough distribution of the energy in the spectrum which method has the biggest chance of success, or, alternatively, to use the output of both methods to drive some logic that will assign a weight to the output of each.
It's a tricky problem, to put it mildly. Also, this is the simplest form of the problem, doing this accurately for multiple pitches at once is much harder.
Another source of inspiration is the 'onsets and frames' software that powers some automated transcription software:
https://github.com/magenta/magenta/tree/master/magenta/model...
I think if this code is over your head that maybe a good introduction course on signal processing would be a nice thing to have under your belt.
Best of luck!