Tensorflow.js: Machine Learning in JavaScript
tensorflow.org
tensorflow.org
https://github.com/paruby/mnist
The second is the classic game of snake, controlled by pointing your head in the direction you want to steer the snake.
https://github.com/paruby/snake-face/
This uses the MediaPipe Facemesh model with the device camera to work out which direction your head is pointing in.
Looking at the page this links to, it gives pretrained models for "common use cases", among other things. Can you actually do anything interesting or novel with those? I feel like any meaningful new use of ML will not be possible with such "beginner building blocks" and you would just wind up with a weaker version of existing offerings. Could someone give me any examples of meaningful ways this can be used?
You can reuse that, and fine tune to the model to your specific data set, which could save you days worth of computation, and also decreases the data set size requirement on your end.
https://github.com/paruby/snake-face/
I'm sure that many people would not class this as interesting or novel, but it demonstrates that at least some of the building blocks can actually be used to make actual things.
There are many boring but meaningful tasks in which this can be used. For example, I'm sure many industries could be benefit from image classification for very specific cases (e.g., fruit categorization). In those cases, you are not interested in the classification of "general" objects such as car, person, bike or horse (as provided by MobiletNet pretrained model), but you can use that model as base to classify different categories of fruit.
Therefore, you are right that they might not be very useful to build new ML algorithms or network architectures. But they are useful to build specific (and novel) uses of current neural networks.
They ship a network to convert landscape photos to landscape paintings, for example. You can retrain that with limited effort and knowledge by using their Collab notebook. That way, you can make an AI to convert photos of cats to photos of dogs. But you wouldn't be able to create an AI to convert between cats and rabbits, for example, because those are too different for that AI architecture to tolerate.
Similarly, they ship a classifier to detect objects of some categories based on surface texture. You can retrain that to detect different categories of things. But if someone paints a chair to look like elephant skin, the AI will be confused and short of inventing a new architecture yourself, there's nothing you can do to fix it.
Plus in any case, you will need to have large high quality datasets for your retraining. The biggest cost in building a bird classification app isn't the one ML employee that creates the final AI, it's hiring biologists to correctly tag thousands of photos to train the AI on.
That's why many AI research nowadays tries to offload the data generation to random people on the internet. For example, recaptcha, or crawling Flickr for that faces dataset.
Edit: since it was brought up in the comments, the retraining that I mentioned can be done only on parts of the model, which is then called transfer learning and can save you some computing time.
But, you'll need to determine if the first feature layers of the original model are suitable for your new classification task. Even experts sometimes guess wrongly here, so as a beginner, I'd say you can only pray and hope.
I reckon you probably actually could with some of the more recent results.
I also believe that I haven't read a paper with a suitable architecture yet.
Which one are you referring to?
Try this: https://github.com/NVLabs/FUNIT (2019)
Or this: https://github.com/clovaai/stargan-v2 (2020)
Either of those should probably work for cats to rabbits.
Both of them massively reduce the complexity by working on position-aligned images of purely the head. I would estimate that even a picture of a cat from the side will already be too much for this network to handle.
These kind of networks learn their feature space mapping by treating the input data set as continuous. So for them to learn a good mapping from cat to dog, it would need to also see photos of an animal that is half cat half dog. If I had to train this case for work, I guess I'd try to go through baby pictures. Dog -> baby dog -> baby cat -> cat. That might work if baby cats and baby dogs look similar enough.
They have pose detection: https://github.com/tensorflow/tfjs-models/tree/master/posene.... One day during a talk with someone who works in aged care I worked up a demo of fall detection using this.
I think small AI models that run completely in the browser and provide personalization by learning from how the user interacts with a given website are the future. This empowers the user and puts them in charge of how their data is used. The example I mentioned previously to demonstrate this was about ranking new submissions on HN: https://news.ycombinator.com/item?id=23407549. I'll quote the relevant part
> TensorFlow.js is a pretty nifty piece of software and it's underutilized. If the model parameters can be stored in IndexedDB then users could train TensorFlow.js based site augmentation to suit their own needs. For example, what if HN had a TensorFlow.js model for ranking new submissions based on the user's preferences? This model could be trained like a spam filter and would eventually learn the types of articles that someone likes to see but they would be in charge of the model's evolution and so would be empowered to use it however they saw fit. Maybe I don't care about politics then my model parameters would eventually converge on downgrading all political posts and the more technical submissions would rise to the top based on how I upvoted new and front page submissions.
One can even imagine a decentralized sharing mechanism where users can create ensembles of models by combining models trained by different users.
All the building blocks are there. Just requires mindshare and a few killer applications.
Also if you have any novel model it will be trivial to reverse engineer it. You gotta send the weights over to the client and they can just run tensorflow.js themselves right?
Re: bandwidth. Do you mean the weights? If so then yes, you'd have to think about distilling large models to smaller ones. There are techniques for doing this. Here is an explanation of distillation from Floyd Hub's engineering blog: https://blog.floydhub.com/knowledge-distillation/.
Novel models will probably need to have some custom layers, which are quite painful to write. The weights for the web version will probably be a low quality of what you achieve on a more powerful machine. And you don't need to provide the code for the training, nor the dataset used, so you will still have an edge over people copying you.
The main problem I have with tensorflow.js is that it's not production ready yet. It's probably 10 years ahead of it's time. Most people don't have recent GPU or recent video drivers. As it's doing some tricks with the GPU to have some acceleration, a fraction of your clients will encounter random bugs and crashes. I even got some bad review for causing reboots on Firefox :). For some people it will be so slow it will feel unresponsive, and others will find that it's unimpressive because you would have to use the minimal model. In a day and age where you need 99% satisfaction to exist, picking tensorflow.js is a mistake.
[1] chrome : https://chrome.google.com/webstore/detail/colorify/pipnfjhpb...
[2] firefox : https://addons.mozilla.org/en-US/firefox/addon/colorify/
[3] Wisteria : https://gistnoesis.github.io/
Internet speeds are still increasing exponentially, so hopefully the model size problems will become less of an issue - perhaps aided by some CDN-served (and thus cached across domains) "base" models that are fine-tuned with some parameters downloaded from the server. I think I could start playing with it seriously if I could get the model sizes under 30mb or so. In a few years (with increasing internet speeds) that might be 50mb. I think huggingface's distilled GPT-2 model is a couple of hundred megabytes, for reference, so we're certainly not going to be doing anything revolutionary in the browser, but I have a bunch of neat little ideas that I think would be useful.
For bringing the "big" stuff to the web we're probably stuck with APIs, like OpenAI's new GPT-3 offering. I think access to SOTA models on hard problems is going to be almost exclusively via APIs for the foreseeable future.
I might have misunderstood your comment, but if you've just started learning to code, I'd probably start with simpler tutorials and projects - just to get a handle on the basics. Depending on your background it could take a good 6 months of daily practice to get comfortable with the basics. It took me and my friends at least that long - similar time frame to learning a spoken language like Japanese, say, to a basic level. Learning to code is absolutely worth it though - 6 months is a very cheap price in my opinion. I'd probably start with a tutorial like this one: https://www.youtube.com/watch?v=yPWkPOfnGsw Or the Khan Academy computer science course.
I'm trying to tackle the topic of ML / AI for some time but I don't want to just use any existing tool and go with the flow - what I would like to do, is to understand how everything works, including underlying math.
Just to give you an example, I would like to build my own image recogniton tool from scratch and understand every part of it.
Where can I start? Any book recomendations would be greatly appreciated.
If you want to train your own models then you'll need to use Python (or in theory Swift or something..) and you want an NVidia video card.
https://github.com/paruby/mnist
See model/mnist_js.ipynb to see how the model is trained and exported to a js readable format, and you can see in lines 13-15 in index.html how the model is loaded.
Note: I have a decent amount of ML experience but almost no javascript experience, so YMMV
https://github.com/tensorflow/tfjs/tree/master/tfjs-converte...