Show HN: Using deep learning to detect objects from your webcam
modeldepot.github.io
modeldepot.github.io
After seeing the announcement of Tensorflow.js, and being a lover of both JS and ML, I couldn't help myself from creating a demo with a non-trivial ML model. YOLO is definitely one of the most interesting/useful ML models out there (as we saw on ModelDepot) and thought that'd be cool to get working.
As far as I know, there hasn't been any prior work with getting YOLO object detection working in JS, I assume mainly because the post-processing code to transform the model's outputs to bounding boxes is a whole heap of math. Luckily tf.js makes it a lot easier to translate from Python and so here it is!
You can drop this into your project with just a couple lines of code (and no ML required) using NPM! (via npm i tfjs-yolo-tiny)
If you want to help with perf/bug fixes, PR's are more than welcome @ https://github.com/ModelDepot/tfjs-yolo-tiny !
Depending on how accurate you need your model to be, (and what platform you're looking to build on), you could use tfjs-yolo-tiny if it's a browser-based environment. Otherwise I'd recommend doing processing in the cloud using the full YOLO model which is more accurate than Tiny YOLO but slower (https://modeldepot.io/mikeshi/yolov2/overview) or FSSD which is more accurate than YOLO but slower (https://modeldepot.io/lzx1413/fssd/overview)
Oh and if you might be doing something self-driving car related, a segmentation network might be useful as well (ex. https://modeldepot.io/hellochick/pspnet/overview)
Thanks for all this information, it seems I have a lot to learn! :)
Here is what it detected for me https://imgur.com/a/DSIZQ
As for that detection, this model is using Tiny YOLO, it's a simpler ML model made to run on more constrained hardware (end user devices) and therefore does not perform as well as the full YOLO or SSD. The other issue is that usually these models take context heavily into account, a fun experiment is to raise anything (ex. wallet) up to your ear and it'll interpret it as a cell phone. I've read an interesting paper recently that forces models to focus more on the object itself than the context to classify things.
Does HN use the rate of exclamation points in their ranking algo?