It's open source: https://github.com/oxen-AI/oxen
Most users are building their own interface on top, but we also have a nice experience in our web ui: https://oxen.ai
Would love people to try it out and give us feedback!
28 karma · joined July 16, 2019
It's open source: https://github.com/oxen-AI/oxen
Most users are building their own interface on top, but we also have a nice experience in our web ui: https://oxen.ai
Would love people to try it out and give us feedback!
Would love any feedback on it or contributions if people are interested :)
Started with the machine learning use case for datasets and model weights but seeing a lot of traction in gaming as well.
Always open for feedback and ideas to improve if you want to take it for a spin!
Grab a cup of coffee and enjoy!
If you haven't seen the Oxen project yet, we have been building an open source unstructured data version control tool.
We were inspired by the idea of making large machine learning datasets living & breathing assets that people can collaborate on, rather than the static ones of the past. Lately we have been working hard on optimizing the underlying Merkle Trees and data structures with in Oxen.ai and just released v0.19.4 which provides a bunch of performance upgrades and stability to the internal APIs.
To put it all to the test, we decided to benchmark the tool on the 1 million+ images in the classic ImageNet dataset.
The TLDR is Oxen.ai is faster than raw uploads to S3, 13x faster than git-lfs, and 5x faster than DVC. The full breakdown can be found here.
https://docs.oxen.ai/features/performance
If you are in the ML/AI community, or rust aficionados, would love to get your feedback on both the tool and the codebase. We would love some community contribution when it comes to different storage backends and integrations into other data tools.
Happy to answer any thoughts or questions!
With Black Forest Labs’ Flux.1 variants being the current state of the art for image gen, we’re doing a technical dive into a few paper that inspired the work, starting with: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (also known as the Stable Diffusion 3 paper).
If you’d like to join the community tomorrow 10 AM PST we’d love to have you. We do it live over zoom and anyone is welcome to join.
Scaling Rectified Flow Transformers for High-Resolution Image Synthesis https://arxiv.org/abs/2403.03206
Join the paper club: https://lu.ma/arxivdive-27
Full process was documented here:
https://www.oxen.ai/blog/create-your-own-synthetic-data-with...
Next up we are going to train a classifier on the outputs, as well as do some classical NLP (named entities, keywords, sentiment, etc) on it to see what we find.
Mainly a fun side project, but could have some interesting implications assuming candidates are using LLMs in the upcoming elections.
We were messing around with PixArt as a way to fine tune DiT's for image generation. I was pretty impressed with the results and thought I'd share.
https://www.oxen.ai/ox/PixArtTutorial
In this example I downloaded a video from YouTube (the trailer of Wes Anderson's Asteroid City) chopped up the frames, captioned them with LLaVA, and then trained the model to generate in the style of the video. It's only about 340 frames of data so pretty quick to generate and train.
I also compare against pure prompting, which the model did not have encoded in it's base parameters.
Using PEFT and LoRA, it took less than 3 hours on an A10 GPU on Lambda Labs. So cost about $3 in total. Pretty wild that it worked right out of the gate for that cheap.
Hopefully it inspires others for what they could build!
I thought the paper “Discrete Diffusion Modeling by Estimating the ratios of the Data Distribution” was a pretty cool idea, so decided to dive deep into the code, strip it down so I could understand it, then train some models from scratch. My findings are linked here:
https://www.oxen.ai/blog/how-to-train-diffusion-for-text-fro...
I find the diffusion papers a bit difficult to read and looking at the inputs and outputs of code really help me grok what’s going on.
Main takeaways are:
1) It is yet to be seen if these techniques will scale in both data and model size 2) Is an interesting technique in general, kind of wild that the Monte Carlo sampling and denoising works at all 3) The infilling isn’t a super big selling point as is because the context length is fixed during diffusion. You’d have to layer in some hacks to make it work well for code completion or other use cases.
Curious what you guys think about diffusion for text, and hopefully this gives people a jumping off point for understanding and implementing your own!
Props to @louaaron and his team at Stanford and Pika Labs for the initial paper and implementation.
Feel free to join us as we build: https://oxen.ai/community
After reading the Self-Rewarding Language Models paper by the team at Meta, it felt very approachable and reproducible, so we spent some time implementing it.
The scripts provided take any base model and put it in a loop of:
1) Supervised fine-tuning on an initial dataset
2) Generating new prompts using the SFT
3) Generating N responses per prompt
4) Scoring the generated responses 1-5
5) Running DPO on the rewards from the model itself.
We've run it through one loop starting with a Mistral-7b base model and the results are pretty encouraging so far.
Feel free to check it out or run it for yourself and let us know what you think:
Have been diving into the Sora technical report for our paper club on Friday, and decided it would be nice to have a reading list of the background papers need to fully grok everything that is going on in that technical report - each with a little description of the part of the pipeline it would be used for (or a previous state of the art technique that was referenced in the review).
We are going to pick a few of the top papers and go over them as a group in the coming Fridays, so join us if you'd like! It's at 10am PST on Fridays over Zoom.
Paper Reading List:
https://www.oxen.ai/blog/road-to-sora-reading-list
Technical Report:
https://openai.com/research/video-generation-models-as-world...
Join the paper club:
https://blog.oxen.ai/mamba-linear-time-sequence-modeling-wit...
Took me awhile to wrap my head around some of the terminology, so hopefully this helps anyone else trying to grok it.