OpenCV in the browser using WebAssembly and web workers
aralroca.com
aralroca.com
> Cloud Vision does not track, record or send any images, videos, or information provided by the user to any server. All image processing is performed within the application.
You can install the free Chrome/Firefox extension to test it. In general, I continue to be amazed how powerful the web assembly concept is.
WebGPU holds a lot of promise for fast image processing on the client. 10x boost is not uncommon for RTX 2000 devices ;)
Ask since the intent of the code is not about embedding OpenCV in a browser, but offloading the computation workload from +1 users from a server to the endpoint user’s computer.
Might be wrong, but for a single user, this setup would likely not be optimal.
While it'll likely be a lot slower than native implementation, the benefits of an image not leaving your computer could unearth some interesting applications (for example, I am very hesitant to use online OCR services and use them only for data which is public anyway).
If anyone is interested you may check it out on https://xroom.app or even contribute with ideas and commits: https://github.com/punarinta/xroom-plugins/tree/master/nisdo.... For the masks I'd still recommend a desktop browser though.
If you're curious but don't want to bother clicking too much here's how it all looks like: https://imgur.com/93YR87e
Another good resource is https://learnopencv.com/ . Again, good quality stuff, if you know your ways it tends to be enough[0] but it's also a big funnel to get you to but one of their courses.
[0] though Satya Mallick does hide a lot of complexity from the readers, and that bites you if you try to implement things on your own
On my iPhone 11, it requested access to the camera and showed the image, says it’s running at 60fps and is using the camera, but it only captured a single frame.
https://opencv.org/opencv-ported-to-google-chrome-nacl-and-p...
It also has direct script functionality but I think the SDK is more powerful.
[0]: http://camtwiststudio.com/developers-objective-c/ [1]: https://pyobjc.readthedocs.io/en/latest/
OBS: https://obsproject.com/es
Now I know it can be done, but I'm not sure how things have changed (this has been some 3 or 4 years ago), so I can't really give you too many details other than it's possible.
I have been putting AR in-browser when Java applets with JOGL was a thing! I've been nominated twice this year for the Webby awards on AI and AR in browser (1). Small innovative team who have been utilising Emscripten and likeminded technologies for a few years from when Emscripten and WebRTC was starting to be a thing.
I wanted to share some pain points taking this tech to production.
- Bandwidth
This is huge with OpenCV, ~4.5Mb+ to take a picture is quite a difficult bandwidth cost to accept. Especially the clients I worked with have millions of views per day. The total binary for Max Factor VMUA (2) is the same size which includes a large data set needed for a neural network for skin tone analysis and face feature detection.
Learning: Do not include all of OpenCV. You don't need it all, but if you do cherry pick the parts you need. I do recommend writing the simplistic parts (this is for you who just use cv::mat!).
- Speed
If you want a 60 FPS AR effect / AI algo on an Android device OpenCV isn't always the fastest approach. Do not rely on a framework, you will need to get your hands dirty and optimise/rewrite the slow areas. WebAssembly is fast, but not as fast as the desktops and native environments you normally create this code on.
- Market
Not everyone has an iPhone in London. Bandwidth means seconds, JS and WebAssembly execution adds to this. In a world where m-commerce is king this does matter. Think Poland, middle of nowhere in Ohio, Brazil, etc. If it takes 60 seconds for a web app to run on 3g and then another 20 for the executable to start, and then the experience is then sluggish it wont be commercially successful.
- UX
When you put this into a large site most traffic will come via instagram and facebook. On iOS this is typically within a WKWebView which does not support getUserMedia. Make sure you have some nice hints on how to open within iOS Safari (or Android Chrome if the parent app has not enabled permissions).
Nevertheless I wish this blog post existed when I started out. I regret in not writing something similar. In this post I especially love the simplicity of the Emcripten pipeline which is great. It is a fantastic demo post. I do hope it inspires many to play with this innovative stack.
1) https://twitter.com/Holition/status/1258068773623431177 2) https://www.maxfactor.com/vmua/
https://github.com/vinissimus/opencv-js-webworker/blob/maste...
7.75MB.