871 karma · joined December 18, 2014
step 1, identify high value users by net worth, citation count, or number of followers
step 2, select all prompts by high value users
step 3, invest 10 billion thinking tokens in modeling an objective for each user
step 4, build an RL environment for each user
step 5, rollout 10 billion tokens per environment
step 6, train on resulting traces
Training a model this small on them is distillation
When models of this size were last studied seriously such corpora did not exist
X discussion: https://x.com/GregoryDiamos/status/2096873745420075020?s=20
I added some of the main points to the thread so they are easier to read.
Not everyone gets super voting shares, but everyone gets a life and has to live on the same planet.
It was always burning since the world's been turning
I needed a cheap model that runs at over 10k token/sec on a single CPU core for some data processing. So I gave Anthropic claude code a pile of tokens to build one.
It made three discoveries that I thought were interesting:
1) One Intel AMX core can train a 3M active parameter MoE foundation model at 6,616 tok/s on 4.91B NVIDIA Nemotron tokens in a few days.
2) That model shows emergent in-context copying, positional analogies, and basic arithmetic after about 250M tokens.
3) The foundation model gives large gains in downstream SFT, and the training & eval loss keep going down all the way through 4.91B (and likely beyond).
Claude is not as good as a great MLE at debugging MoE. It made a bunch of bone headed mistakes, but it got there in the end.
I asked it to write a paper about it's work, and it produced this.
Claude Co-Authored Paper: https://huggingface.co/gdiamos/amx-reasoning-v1-instruct/blo...
I read through it and it sounds a bit LLMy, but the main points and experiment results are correct.
Some of the models are published on HF: https://huggingface.co/gdiamos/amx-reasoning-v1-instruct
However, I know of no theoretical limits on scaling laws other than compute and data.
I’ve been using diffusion Gemma and it is very fast on GPUs in output token/sec.
In the diffusion Gemma whitepaper, they say they could have done better with more time and compute.
Even with those caveats, it is very uses-able as a local model.
Shouldn't that be 0 innovation tokens though?
Instead I like “only work on impossible problems”
Most of them turn out to be impossible, but some of them turn out to be possible.
I’ve never met anyone who could pick 3 and be confident in getting even one right. Tokens are a terrible analogy for innovation or research.
In hindsight I’ve had to sift through hundreds or more to fine one that worked.
I thought this post was helpful when I first started thinking about startups.
After more time, I think boring tech isn’t worth thinking about.
I wonder how much it would cost to vibe code the whole thing from scatch?
I wonder how much better models need to get before such a thing wouldn't look like code vomit?
That means that 10 years from now I don’t want think about making sure the hosting server is up.
I also want it to have a standard format for bibliography, DOI, and authors.
I agree it isn’t much, but it’s more than I get from a regular web hosting service and it is a standard format for papers so I don’t think it makes sense for every author to roll their own.
I wouldn't want my google drive to start telling me my paper was too sloppy. I just want a link.
There is a lot of follow on work that explains what happens as you change them, e.g. Scaling Laws for Transfer - https://arxiv.org/pdf/2102.01293
I think it’s fortunate that transfer works in a similar way.
Common crawl (and Reddit, stack overflow, etc but not 4chan) was much easier to get access to at the time than using mechanical Turk.
There is certainly room for more work. There were many papers on scaling laws in NeurIPS this year.