HNHacker News
TopNewBestAskShowJobs

amanrs

3 karma · joined April 12, 2020

submissionscomments
amanrs··on Character Prefix Conditioning
This is harder than it looks.

First "token-healing" doesn't work. Consider the case "app" where the most likely options are "ap|praisal" or "apple|sauce". You can't just sample all tokens that start with app, or you'd miss appraisal.

Second, it's easy to come up with a naive algorithm that samples from the true distribution. It's very difficult to make this algorithm efficient.

amanrs··on Editing Files at 1000 tokens/s with llama-70B
It edits the file based on the "plan" laid out by a smarter language model
amanrs··on Llama Is Expensive
The key misconception about many quantization methods is that lower precision = better speed.

I believe GPT-Q is not much faster than bf16 from skimming the AWQ paper - https://arxiv.org/pdf/2306.00978.pdf

It's 3x faster for a batch size of 1, but that's still over 10x more expensive than gpt-3.5

For larger batch sizes, bf16 costs dip below 3-bit quantized.