HNHacker News
TopNewBestAskShowJobs

float-trip

38 karma · joined March 22, 2023

submissionscomments
float-trip··on Reddit Will License Its Data to Train LLMs, We Made a FF Extension to Replace
Reddit's caches are set up to only ever return the last 1,000 of anything. So for example - you can't scroll past 1k items on /new, and if you save more than 1k posts then you'll have to unsave some to retrieve the others.

If this extension only edits comments, it'll only touch the most recent 1k. You would need to retrieve the older ones with a Pushshift replacement like this: https://pullpush.io/. But that also shows how ineffective this is. We still have public reddit archives (like Pullpush and https://github.com/ArthurHeitmann/arctic_shift) which contain comments as they were originally posted. This isn't gonna be a problem for Google.

float-trip··on Fine-tuning Mistral 7B on Magic the Gathering Draft
Related comment from gwern: https://news.ycombinator.com/item?id=38438859. Can't find the docs now - I think they were the old GPT 3 ones - but they suggested a low value somewhere around 0.01 and 0.1.

Also - why qlora rather than a full finetune? Using LambdaLabs, it'd cost roughly the same as your quote. Cheaper I think if you're willing to gamble with fp8: https://github.com/mosaicml/llm-foundry/tree/main/scripts/tr.... And fewer hyperparameters to tune as well

float-trip··on Fine-tuning Mistral 7B on Magic the Gathering Draft
That's what I ended up doing (`[Author] username [Title] post title...`)

> Adding new tokens needs a ton of data to train what the token means.

But how much? 300M tokens is fine for a simple version of ChatML with ~4 tokens. Not for 15, at least in my case. How's this relationship scale?

Just trying to offer one datapoint for what doesn't work, with the hedge that I might have just had a bug

float-trip··on Fine-tuning Mistral 7B on Magic the Gathering Draft
I tried adding special tokens for a reddit-style dataset once. The format was: `<|post_author|>username<|post_title|>title here...`

The resulting model was so much worse than just formatting everything plaintext. This was with MPT-30B, 15 special tokens, 300M training tokens, and a full finetune.

I may have made a mistake, but I haven't seen any open source finetunes successfully add a large number of tokens yet either.

float-trip··on Fine-tuning Mistral 7B on Magic the Gathering Draft
Thanks for writing up. Rather than zeroing out the loss for the prompt, did you also try using weighted loss with Axolotl? At one point, Microsoft's GPT 3 docs suggested this was beneficial when the responses are short (like you have with "Cut in.") Domain adaptation over subreddits/forums before finetuning may help as well.
float-trip··on The Mathematics of Training LLMs
There's a breakdown here for anyone interested (ctrl+f "weight flops for")

https://medium.com/@dzmitrybahdanau/the-flops-calculus-of-la...

float-trip··on GitHub Copilot Chat Leaked Prompt
The prompt for Bing Chat was previously reproduced by the same person as here, using the same trick. The Bing lead disclaimed it as inaccurate, though: https://twitter.com/MParakhin/status/1627491603731423232
float-trip··on Understanding large language models: A cross-section of the relevant literature
Two other recent literature reviews worth reading:

"Transformer Taxonomy" - https://kipp.ly/blog/transformer-taxonomy/

"Five years of progress in GPTs" - https://finbarrtimbers.substack.com/p/five-years-of-progress...