HNHacker News
TopNewBestAskShowJobs

gregschoeninger

28 karma · joined July 16, 2019

submissionscomments
gregschoeninger··on Perforce charges $500 for training training videos.. and it's AI narrated
We've been working for a few years on a VCS that is starting to get adopted by creative studios - called "Oxen"

It's open source: https://github.com/oxen-AI/oxen

Most users are building their own interface on top, but we also have a nice experience in our web ui: https://oxen.ai

Would love people to try it out and give us feedback!

gregschoeninger··on Lore – Open source version control system designed for scalability
We're also working on an open source large asset versioning tool called "oxen" - https://github.com/Oxen-AI/Oxen

Would love any feedback on it or contributions if people are interested :)

gregschoeninger··on The future of version control
We're working on this project to help with the non-text file and large file problem: https://github.com/Oxen-AI/Oxen

Started with the machine learning use case for datasets and model weights but seeing a lot of traction in gaming as well.

Always open for feedback and ideas to improve if you want to take it for a spin!

gregschoeninger··on [dead]
Over the past ~1.5 years I've been running a research paper club where we dive into interesting/foundational papers in AI/ML. So we naturally have come across a lot of the papers that lead up to DeepSeek-R1. While diving into the DeepSeek papers this week, I decided to compile a list of papers that we've already gone over or I think would be good background reading to get a bigger picture of what's going on under the hood of DeepSeek.

Grab a cup of coffee and enjoy!

https://www.oxen.ai/blog/no-hype-deepseek-r1-reading-list

gregschoeninger··on [dead]
Hey all,

If you haven't seen the Oxen project yet, we have been building an open source unstructured data version control tool.

We were inspired by the idea of making large machine learning datasets living & breathing assets that people can collaborate on, rather than the static ones of the past. Lately we have been working hard on optimizing the underlying Merkle Trees and data structures with in Oxen.ai and just released v0.19.4 which provides a bunch of performance upgrades and stability to the internal APIs.

To put it all to the test, we decided to benchmark the tool on the 1 million+ images in the classic ImageNet dataset.

The TLDR is Oxen.ai is faster than raw uploads to S3, 13x faster than git-lfs, and 5x faster than DVC. The full breakdown can be found here.

https://docs.oxen.ai/features/performance

If you are in the ML/AI community, or rust aficionados, would love to get your feedback on both the tool and the codebase. We would love some community contribution when it comes to different storage backends and integrations into other data tools.

gregschoeninger··on Data Version Control
Maintainer of Oxen here, we initially built Oxen because DVC was pretty painfully slow to work with, and had a lot of extra bells and whistles that we didn’t need. Under the hood we optimized the merkle tree structure, hashing algorithms, network protocols, etc to make it speedy when it came to large datasets. We have a pretty nice front end at https://oxen.ai for viewing and querying the data as well.

Happy to answer any thoughts or questions!

gregschoeninger··on Paper Club: How Flux.1 models work under the hood
Hey all,

With Black Forest Labs’ Flux.1 variants being the current state of the art for image gen, we’re doing a technical dive into a few paper that inspired the work, starting with: Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (also known as the Stable Diffusion 3 paper).

If you’d like to join the community tomorrow 10 AM PST we’d love to have you. We do it live over zoom and anyone is welcome to join.

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis https://arxiv.org/abs/2403.03206

Join the paper club: https://lu.ma/arxivdive-27

gregschoeninger··on Using Llama3.1 405B to generate political synthetic data
We thought it'd be interesting to see what political biases Llama 3.1 405B has by generating a bunch of "spam" or "ham" messages with it. We started with 5 hand crafted messages and let the LLM take it from there ending up with over 1k.

Full process was documented here:

https://www.oxen.ai/blog/create-your-own-synthetic-data-with...

Next up we are going to train a classifier on the outputs, as well as do some classical NLP (named entities, keywords, sentiment, etc) on it to see what we find.

Mainly a fun side project, but could have some interesting implications assuming candidates are using LLMs in the upcoming elections.

gregschoeninger··on Fine Tuning a Diffusion Transformer (DiT) from a Single YouTube Video
Hey all,

We were messing around with PixArt as a way to fine tune DiT's for image generation. I was pretty impressed with the results and thought I'd share.

https://www.oxen.ai/ox/PixArtTutorial

In this example I downloaded a video from YouTube (the trailer of Wes Anderson's Asteroid City) chopped up the frames, captioned them with LLaVA, and then trained the model to generate in the style of the video. It's only about 340 frames of data so pretty quick to generate and train.

I also compare against pure prompting, which the model did not have encoded in it's base parameters.

Using PEFT and LoRA, it took less than 3 hours on an A10 GPU on Lambda Labs. So cost about $3 in total. Pretty wild that it worked right out of the gate for that cheap.

Hopefully it inspires others for what they could build!

gregschoeninger··on How to train diffusion for text from scratch
Hey all,

I thought the paper “Discrete Diffusion Modeling by Estimating the ratios of the Data Distribution” was a pretty cool idea, so decided to dive deep into the code, strip it down so I could understand it, then train some models from scratch. My findings are linked here:

https://www.oxen.ai/blog/how-to-train-diffusion-for-text-fro...

I find the diffusion papers a bit difficult to read and looking at the inputs and outputs of code really help me grok what’s going on.

Main takeaways are:

1) It is yet to be seen if these techniques will scale in both data and model size 2) Is an interesting technique in general, kind of wild that the Monte Carlo sampling and denoising works at all 3) The infilling isn’t a super big selling point as is because the context length is fixed during diffusion. You’d have to layer in some hacks to make it work well for code completion or other use cases.

Curious what you guys think about diffusion for text, and hopefully this gives people a jumping off point for understanding and implementing your own!

Props to @louaaron and his team at Stanford and Pika Labs for the initial paper and implementation.

gregschoeninger··on Instruct-Tuning BitNet 1.58
This is work done for our arxiv dive paper club where we dive into research papers and implement code to see how the models work in practice. We have some internal use cases for BitNets so thought we'd share the work as we go along. Enjoy!

Feel free to join us as we build: https://oxen.ai/community

gregschoeninger··on Show HN: Implementation of the "Self-Rewarding Language Models" Paper by MetaAI
We used an A10 with 24GB of VRAM, this was enough for PEFT on Mistral-7B
gregschoeninger··on Show HN: Implementation of the "Self-Rewarding Language Models" Paper by MetaAI
The goal is to iteratively create training data and add it to its own training set. The LLM acts as its own judge and scores its own responses to decide if it should add the data. It’s expensive to have a human in the loop labeling preferences, so the folks at Meta showed you can have a clever prompt and fine tune the model to judge its own responses.
gregschoeninger··on Show HN: Implementation of the "Self-Rewarding Language Models" Paper by MetaAI
Hey all,

After reading the Self-Rewarding Language Models paper by the team at Meta, it felt very approachable and reproducible, so we spent some time implementing it.

The scripts provided take any base model and put it in a loop of:

1) Supervised fine-tuning on an initial dataset

2) Generating new prompts using the SFT

3) Generating N responses per prompt

4) Scoring the generated responses 1-5

5) Running DPO on the rewards from the model itself.

We've run it through one loop starting with a Mistral-7b base model and the results are pretty encouraging so far.

Feel free to check it out or run it for yourself and let us know what you think:

https://github.com/Oxen-AI/Self-Rewarding-Language-Models

gregschoeninger··on "Road to Sora" Paper Reading List
Hey all,

Have been diving into the Sora technical report for our paper club on Friday, and decided it would be nice to have a reading list of the background papers need to fully grok everything that is going on in that technical report - each with a little description of the part of the pipeline it would be used for (or a previous state of the art technique that was referenced in the review).

We are going to pick a few of the top papers and go over them as a group in the coming Fridays, so join us if you'd like! It's at 10am PST on Fridays over Zoom.

Paper Reading List:

https://www.oxen.ai/blog/road-to-sora-reading-list

Technical Report:

https://openai.com/research/video-generation-models-as-world...

Join the paper club:

https://lu.ma/oxenbookclub

gregschoeninger··on Guide to the Mamba architecture that claims to be a replacement for Transformers
Been diving deep into the Mamba paper and put together my notes here:

https://blog.oxen.ai/mamba-linear-time-sequence-modeling-wit...

Took me awhile to wrap my head around some of the terminology, so hopefully this helps anyone else trying to grok it.