479 karma · joined January 23, 2022
Newsletter coming, and will definitely add filtering by tags/topics + search!
Comments sounds like a good idea as well!
I'm working on a new project called Uminal, which enables anyone to extend large language models (like GPT-3) with custom apps that auto-compose w/ each other.
In short, imagine if anyone could extend LLMs with new skills that can interact w/ the real world. That's what Uminal enables.
I would love to get some feedback on it, so thought I'd share a demo video of it on HN!
Demo: https://www.loom.com/share/211327fd8e854513b909b0f69eadd2f8
I'm currently working to get the prototype online very soon.
Thanks!
So, one possible way to train a model on ImageNet using this tool could be to:
1. Store the dataset on S3
2. Have the training script download the necessary data from S3 at the beginning (for example, each process could only download the shard of data that it needs, etc.)
3. Use MLbot to run your local training script on AWS as a distributed training job with X number of nodes in your EKS cluster
In this case, when you execute "mlbot run ..." from the command-line, MLbot will package your code into a docker image and then create X number of pods within EKS (where X = the number of training nodes).
Once all of the pods are running, etcd + PyTorch Elastic begin handling all of the communication / synchronization between the various pods & distributed training happens.
So in this example, your code would be transferred to your EKS cluster (where the training would happen) and the training pods would then fetch the ImageNet data from S3.
And most importantly, all of this happens inside of your cloud infrastructure.
Now that I think about it, I think it might be helpful to add an ImageNet-based example to the repo.
I just released the initial version of MLbot: a new open-source tool that I’ve been working on for running ML training jobs in your cloud, with a single command.
In short, it allows you to run your training script in the cloud by simply swapping “python” for “mlbot run”.
For example, if “python train.py …” can run your training script locally, then “mlbot run —instance-type p3dn.24xlarge —num-nodes 2 train.py …” should be able to run your code in the cloud across 2 GPU machines.
Since this tool runs entirely inside your cloud environment, you don’t have to transfer your training data to a 3rd party, while having full observability into the underlying infrastructure.
Why I built this:
In a recent ML project I was working on (with a 10+ TB training dataset), I found that the tools to run distributed training jobs in the cloud were fairly painful & complex to use for a single developer like myself.
Current solutions like Kubeflow felt way too bloated, while typical 3rd-party hosted solutions didn’t work for my use-case either, since it often meant transforming & duplicating my large training dataset into a service-specific format (which would be both very expensive & impractical due to data security).
After unsuccessfully trying numerous solutions (including months w/ AWS Sagemaker), I finally ended up using AWS Elastic Kubernetes Service combined with PyTorch Elastic + PyTorch Lightning, which turned out to work really well for scalable, distributed training.
At the time, I wrote some scripts that automated the process of packaging my raw training code and launching it as a distributed training job on EKS, and I found that this significantly improved my speed of iterating & launching new distributed training jobs.
So, I thought it would be nice to package my scripts and open-source it as a single tool, in case others find it useful.
The current code definitely needs to be refactored into a much better structure, but I wanted to first get the current version out there, so that I can learn if this solves a real pain-point that others also face.
Finally, I’d love to hear your feedback, and any thoughts/experience/pain-points you might have related to running distributed ML jobs in the cloud!
Thanks! :)
Re: the off-by-one bug, it appears that it's actually due to the underlying indexed catalog having some metadata issues, where some songs can have the wrong audio associated with them.
And since the search algorithm is only audio-based, this can cause it to return results where the audio is similar but the associated song/artist name for that audio is incorrect.
In your example, for instance, the underlying catalog has the "Uprising (Muse)" duplicated under an incorrect, different name "Sojourner (Phillip Keveren & London Symphony Orchestra)". This tricks the search algorithm into thinking that this is another very similar track, when in reality it's just a duplicate version under an incorrect name.
This is a tricky bug that I still need to figure out how to solve.
Nevertheless, thanks a lot for your feedback! :)
Yeah, currently the indexed catalog has some metadata issues, where some songs can have the wrong audio associated with them.
And since the search algorithm is only audio-based, this can cause it to return results where the audio is similar but the associated song/artist name for that audio is incorrect. This can also cause issues such as having duplicates of the same song under different names, etc.
I definitely agree that this can be confusing, and I’m still trying to figure out a way to best address this issue.
And the search ranking algorithm can definitely be improved too! :)
So I’ve been working on this project where you can search for similar-sounding songs.
Basically, you enter the name of a song in a search bar, and it will use an audio-based search algorithm to return a list of other songs that sound similar.
Since the similarity search is audio-based, it attempts to prioritize songs with similar melodies, beats, and feel, irrespective of their popularity.
Here are a few example searches for similar songs:
- Fetish (feat. Gucci Mane) - Selena Gomez: https://maroofy.com/songs/1563859943
- Faded - Alan Walker: https://maroofy.com/songs/1087732055
- Bandit - Juice Wrld: https://maroofy.com/songs/1577864923
- L.I.F.E - Remady & Manu-L: https://maroofy.com/songs/1170014710
- A chill instrumental track: https://maroofy.com/songs/1592077074
It’s still a very rough prototype, but thought it would be nice to share it here to get some early feedback! :)
Some current limitations:
- Sometimes, the results may contain similar songs that are duplicates under a different name