75 karma · joined January 20, 2023
Qwen-Image-Edit being Apache 2.0 allows you to fine-tune and apply a few tricks like compilation, lightning lora and quantization to cut costs.
The base model takes ~15s to generate an image which would mean we would need 1,200,000*15/60/60=5,000 compute hours.
Compilation of the PyTorch graph + applying a lightning LoRA cut inference down to ~4s per image which resulted in ~1,333 compute hours.
I'm a big fan of open source models, so wanted to share the details in case it inspires you to own your own weights in the future.
We've had a few requests to integrate with music production workflows, but haven't taken it on yet. If anyone wants to collaborate to integrate Oxen with their DAW or workflow let us know! Here's the project:
We have an open source CLI and server that mirrors git, but handles large files and mono repos with millions of files in a much more performant manner. Would love feedback if you want to check it out!
Grab a cup of coffee and enjoy!
Feel free to check it out here: https://github.com/Oxen-AI/oxen-release
Or a hub you can host data on (we have public and private repos, or private VPC deployments): https://oxen.ai
The CLI mirrors git so it's easy to learn. It has some interesting build in tooling for diff-ing datasets and working on them remotely without downloading a full copy of the data as well.
Happy to answer any other questions!
If you haven't seen the Oxen project yet, we have been building an open source unstructured data version control tool.
We were inspired by the idea of making large machine learning datasets living & breathing assets that people can collaborate on, rather than the static ones of the past. Lately we have been working hard on optimizing the underlying Merkle Trees and data structures with in Oxen.ai and just released v0.19.4 which provides a bunch of performance upgrades and stability to the internal APIs.
To put it all to the test, we decided to benchmark the tool on the 1 million+ images in the classic ImageNet dataset.
The TLDR is Oxen.ai is faster than raw uploads to S3, 13x faster than git-lfs, and 5x faster than DVC. The full breakdown can be found here.
https://docs.oxen.ai/features/performance
If you are in the ML/AI community, or rust aficionados, would love to get your feedback on both the tool and the codebase. We would love some community contribution when it comes to different storage backends and integrations into other data tools.
https://blog.oxen.ai/arxiv-dives-mixture-of-experts-moe-with...
Feel free to join here: https://lu.ma/oxenbookclub
https://blog.oxen.ai/mamba-linear-time-sequence-modeling-wit...
I'm still not convinced on Mamba's performance on Natural Language tasks, but maybe it's just because they haven't trained a large enough model on enough data yet.
https://blog.oxen.ai/reading-list-for-andrej-karpathys-intro...
https://blog.oxen.ai/mamba-linear-time-sequence-modeling-wit...
I hadn't found a very satisfying explanation of the paper yet, and still had some questions at the end, but hopefully this can give people a good jumping off point for their understanding!
https://blog.oxen.ai/practical-ml-dive-how-to-customize-a-vi...
~ TLDR ~ ViT works the best in this small experiment, with minimal code. The experiment was classifying 7 different facial emotions such as "happy", "sad", "angry", etc...
Model Accuracy
* ViT - 69% * ResNet50 64% * Zero-Shot CLIP - 53%
Was honestly most impressed with CLIP's ability for zero-shot transfer, even though it had the worst accuracy. The ability to give it a freeform list of prompts or labels and it will automatically classify into the subset without training feels like the future of prototyping products and models, then once you define your use case go with something more performant like a ViT.
Anyways, I had fun writing the code and running the experiments, so thought I would share!
Many of these datasets have many many images, videos, audio files, text as well as structured tabular datasets that git or git-lfs just falls flat on.
Would love anyone to kick the tires on it and let us know what you think:
https://github.com/Oxen-AI/oxen-release
The commands are mirrored after git so it is easy to learn, but optimized under the hood for larger datasets.
Every Friday we've been going over the fundamentals of a lot of the state of the art techniques used in Machine Learning today. Hoping to learn a little each week, and spot patterns we can apply to our own work. I feel like there's always a little nugget of information I didn't fully understand before reading the paper, so have been finding it helpful.
Though it is not groundbreaking research as of this week, I think it's nice to take a step back and review the fundamentals as well as keeping up with the latest and greatest.
Posted the notes and video recap are here if anyone finds it helpful:
https://blog.oxen.ai/arxiv-dives-zero-shot-image-classificat...
Also would love to have anyone join us live on Fridays or suggest papers! We've got a pretty consistent and fun group of 400+ engineers and researchers popping in and out.
https://blog.oxen.ai/suds-a-guide-to-structuring-unstructure...
Though it is not groundbreaking research as of this week, I think with the pace of AI it is important to dive deep into past work and what others have tried! It's nice to take a step back and learn the fundamentals as well as keeping up with the latest and greatest.
Posted the notes and recap here if anyone finds it helpful:
https://blog.oxen.ai/arxiv-dives-vision-transformers-vit/
Also would love to have anyone join us live on Fridays! We've got a pretty consistent and fun group of 300+ engineers and researchers showing up.
The full talk can be found here: https://youtu.be/ zjkBMFhNj_g?si=fPvPyOVmV-FCTFEx
Here's the reading list: https://blog.oxen.ai/reading-list-for-andrej-karpathys-intro... video/
Let me know if you have any other papers you would add!
Website: https://oxen.ai
Dev Docs: https://docs.oxen.ai
GitHub: https://github.com/Oxen-AI/oxen-release
Feel free to reach out on the repo issues if you run into anything!
https://github.com/Oxen-AI/oxen-release#-oxen
Going down your list of requirements, Oxen has:
* Data versioning, similar paradigm to git, but built from the ground up for large ML datasets
* Inexpensive storage, comparable pricing to s3
* Branching/Merging for maintaining production training data sets
* Metadata storage and query capabilities, works with many structured data types. Have APIs for querying.
* User interface for less tech savy people, building out a hub at https://www.oxen.ai to enable this.
* Being able to define datasets that are a subset of the whole collected data (is this a similar requirement to querying?)
* Data ingestion pipeline - engineers would have to hook into APIs or CLI tools right now.
Feel free to check it out and leave any feedback on the GitHub repo!
Fundamentally even adding and committing data locally is slower, even before the push. But I agree the remote matters too.