Ancient secrets of computer vision
pjreddie.com
pjreddie.com
Let's take the circle Hough transform as it's one of the most enlightening ones!
Say you are looking for a circle of a given diameter. After a binarization to make the edge stand out, make all the potential points "vote" for a circle center.
The method is simple: using a matrix, you +1 all the points that are as far from this point as the radius of the circle will allow.
Do this for every point, and take the max: https://en.wikipedia.org/wiki/Circle_Hough_Transform
Simple, and works in guaranteed time.
Extension 1: if you don't know the radius, apply iteratively for a range of values, then again, take the max: if you imagine how it works (or code it as an example then animate the result), it's like doing a "mathematical" focus.
Extension 2: if it's too costly to do a dense exploration of the space of values for the radius, while you know there's only one circle, do a gradient descent on the increase.
Extension 3: If there are more that one circle, other techniques exist - the easiest to picture are based on the maximization of variance of the distribution of values in the matrix resulting from the binarization, but you can also use 2d lattices and other fun tricks.
It's even simpler to make artificial neurons vote for a circle center.
You don't need the binarization step, and you can apply the method to other shapes as well.
Is it?
It's not conceptually simpler: people can more easily imagine circles around points converging to a center, so they can also put that idea into code more easily.
> you can apply the method to other shapes as well.
Yes you can. Read about Hough.
I just presented the one that is the most enlightening.
I may be biased against neural network approaches and their likes, because I see them as black boxes with failure modes that are hard to predict or work around: I prefer what I can understand and explain, and unfortunately, it seems at odd with the current demographics of ML (cf https://news.ycombinator.com/item?id=27361812 ) who has no clue about what makes these black boxes tick, sometimes even after they get a PhD in the dark art of tweaking black boxes.
* https://towardsdatascience.com/lines-detection-with-hough-tr...
* https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.2....
It comes with a wonderful demonstration tool that allows you to apply the various included algorithms to images and tweak the parameters in real-time – including the Hough transform. A great tool for helping to understand how these kinds of algorithms work!
A "faint" circle will often score worse than 2 high-contrast parallel lines that happen to be the right distance apart, since the lines manage to trigger pixels along 20% of a circle's arc and their contrast massively inflates their score compared to the faint circle (higher edge pixel density).
It seems like there should be a simple way to weight the results by how dispersed within the circle's arc the pixels are, but I've never dug any further, after hitting this problem I had to move on.
1. The most valuable take from computer vision
2. The simplest take from computer vision
Not to mention this is rarely useful unless you're in a specific context where you're looking for circles in an image.
(personally I was impressed by mixture of gaussians background removal))
Edge binarization is dependent upon edge detection algorithm choice, threshold algorithm choice, and both of their respective parameters. It's often very difficult to find a set of parameters that aren't brittle due to occlusions, poor contrast, camera noise, etc.
Hough works great if you can do this part confidently. But in my experience, robust edge binarization for Hough is often not very feasible in the wild.
Most CV tasks are borderline impossible if your input is acquired under uncontrollable lighting. Whereas the right illumination setup can often let you get away with nothing but a threshold binarization.
It used to be the one taught by Justin Johnson at UMich [0].
But the publicly available videos have been last updated in 2019.
[0]: https://web.eecs.umich.edu/~justincj/teaching/eecs498/FA2020...
[0] https://uni-tuebingen.de/fakultaeten/mathematisch-naturwisse...
Previous authors didn’t start their axes at 0, so he kept their axes and just put the timing for YOLO outside the original chart area.
A few years ago he said he’d thought about quitting research and opening a vegan cafe or something. Not sure what he’s planning to do now though.
https://ecommons.cornell.edu/bitstream/handle/1813/6165/92-1...
In a different life I toyed with it a bit: http://pugoob.blogspot.com/2008/01/pugoob-image-search-tool....