RetinaFace: Single-stage Dense Face Localisation, implemented in TensorFlow 2.0
github.com
github.com
Mxnet | 96.5 | 95.6 | 90.4
Ours | 95.6 | 94.6 | 88.5
My professors would be so mad if I submitted data like this. I already can hear "What are the units? Seconds? So yours are by one second better? Error margin? Percent? So yours is worse?"
I get that this is a very specific information for a specific audience. People who stumble on this repo should know what is that.
However we can all be better at presenting our data.
https://medium.com/@jonathan_hui/map-mean-average-precision-...
"So this is precision? 0 to 1? 1 being totally accurate? And you got 96.5? I assume it is percentage?"
If your professor is working in on object detection they know what mAP is - all the major datasets use it as their standard evaluation criteria.
Figures like precision and recall are often expressed either as 66% or 0.66. Confusion really isn't that big a problem.
If you would write mAP: 0.9 or mAP (%): 90 or similar my "professor" would be OK with that.
I guess general rule in science (or maybe engineering) is to define units. I have no doubt that e.g. when you are building rockets (or maybe model airplanes) you would know that e.g. velocity is in m/s. Until somebody thinks in some other way and things blow up.
I just wanted to make some fun comment about missing units and that is it. I have no doubt that anybody browsing that repo would not be confused about mAP being 95.6.
Professors will rail against imprecision to undergrads and young graduate students but they use it all the time in real life.
I can already hear my professors saying "stop commenting on style, what do you have to say about substance?"
Note however that Retina does not support real-time performance on the CPU especially on IoT devices and web browsers (WebAssembly). That's why we opted for a standard cascade approach for our WebAssembly port: https://sod.pixlab.io/articles/porting-c-face-detector-webas...
Would you consider using an open source license like for example the ISC license?
For model file size I have a few ideas of things that could be tried easily.
For a faster model, this model is based on resnet50 architecture. The original authors also published a lighter mobileNet architecture. I havent converted it to tensorflow but I could try to take a look
Generally it means something roughly like "the best known approach for this specific problem". Often it means "the best known approach for this specific dataset" (eg "SOTA on ImageNet").
> why is lower accuracy acceptable?
Lower accuracy is worse, but these numbers look close enough that it's probably acceptable for most people.
There are plenty of environments where TF is preferable to MXNet (eg, you have TPUs/want to use TFLite on mobile/want to slice the model weights up and use it for your own custom TF model).
It's probably lower accuracy because it wasn't trained as long. Those extra couple of points could take days (or more) of training.