HNHacker News
TopNewBestAskShowJobs

paul_mk1

223 karma · joined August 6, 2023

submissionscomments
paul_mk1··on Towards 1-bit Machine Learning Models
Sub 1-bit has been done at least as far back as 2016 for VGG style networks (my work).

I was able to get 0.68 "effective" bits.

The idea is that in each forward pass you add noise to each weight independently drawn from normal distribution, and when you calculate snr, it's sub 1 bit. Points to the idea that a stochastic memory element can be used.

https://arxiv.org/abs/1606.01981

paul_mk1··on Fine tune a 70B language model at home
Goes back before then. This got popularized by BinaryConnect in 2015, and groups were training binary networks as early as 2011.

You are probably referring to XNOR net, and the novel piece there was also using binary activations (which bitnet is not).

So as far as I can tell, bitnet is basically BinaryConnect applied to LLMs.

https://arxiv.org/abs/1511.00363

paul_mk1··on Fine tune a 70B language model at home
It's clear that QLoRA has opened up finetuning to a wider audience with limited compute, which is a good thing.

One thing I've wondered about: what are the drawbacks to using QLoRA? For example if compute is not a limit, I'm guessing one should not use QLoRA and finetune in full precision instead?

Afaik when a model is first quantized to nf4 (before finetuning begins), model performance is degraded from baseline (see https://x.com/Tim_Dettmers/status/1661482614811918338?s=20).

Dettmers shows that after finetuning wrt the dataset, the result is as good as full precision. But afaik never explored the effects outside the finetuning data. Assuming the finetuning dataset is small, the model will just be the degraded nf4 version, right? Or perhaps finetuning will even skew the model in weird ways (trying to fix quantization errors).

Anecdotally models finetuned wth QLoRA perform well. Does anyone have any papers or a careful analysis of this?

paul_mk1··on The Era of 1-bit LLMs: ternary parameters for cost-effective computing
I don't think there is anything conceptually new in this work, other than it is applied to LLMs.

But in fairness, getting these techniques to work at scale is no small feat. In my experience quantization aware training at these low bit depths was always finicky and required a very careful hand. I'd be interested to know if it has become easier to do, now that there are so many more parameters in LLMs.

In any case full kudos to the authors and I'm glad to see people continuing this work.

paul_mk1··on The Era of 1-bit LLMs: ternary parameters for cost-effective computing
Nice to know there is a trail to relevant citations. I missed the BitNet paper and need to catch up.

Btw TrueNorth project evolved into "NorthPole" chip by the same group, and was recently in the press. From afar NorthPole looks like an interesting design point and leverages on-chip memory (SRAM)--so it's targeting speed and efficiency at the expense of memory density (so perhaps like Groq in some respects). Tbh I haven't followed the field closely after leaving the group.

paul_mk1··on The Era of 1-bit LLMs: ternary parameters for cost-effective computing
Fun to see ternary weights making a comeback. This was hot back in 2016 with BinaryConnect and TrueNorth chip from IBM research (disclosure, I was one of the lead chip architects there).

Authors seemed to have missed the history. They should at least cite Binary Connect or Straight Through Estimators (not my work).

Helpful hint to authors: you can get down to 0.68 bits / weight using a similar technique, good chance this will work for LLMs too.

https://arxiv.org/abs/1606.01981

This was a passion project of mine in my last few months at IBM research :).

I am convinced there is a deep connection to understanding why backprop is unreasonably effective, and the result that you can train low precision DNNs; for those note familiar, the technique is to compute the loss wrt to the low precision parameters (eg project to ternary) but apply the gradient to high precision copy of parameters (known as the straight through estimator). This is a biased estimator and there is no theoretical underpinning for why this should work, but in practice it works well.

My best guess is that it is encouraging the network to choose good underlying subnetworks to solve the problem, similar to Lottery Ticket Hypothesis. With ternary weights it is just about who connects to who (ie a graph), and not about the individual weight values anymore.

paul_mk1··on MK1 Flywheel Unlocks the Full Potential of AMD Instinct for LLM Inference
Comparisons between different chip architectures are imperfect. In our opinion the most fair thing to do is 1) match the TFLOPs (since these workloads are compute bound), and 2) find a similar card that can run the same size models. Since MI210 has 64GB, the closest one is the A6000 with 48GB. Many of the less expensive ADA cards can't even fit a 13B model in VRAM, so a comparison could not even be made.
paul_mk1··on MK1 Flywheel Unlocks the Full Potential of AMD Instinct for LLM Inference
You can try it yourself on SageMaker using NVIDIA. There's a free trial.

https://aws.amazon.com/marketplace/seller-profile?id=seller-...

For AMD, you'll have to wait until these cards become available on the cloud.

paul_mk1··on MK1 Flywheel Unlocks the Full Potential of AMD Instinct for LLM Inference
This is a good observation, the cards do have different memory bandwidth with the MI210 having more than double the bandwidth via HBM2e.

Note that the comparisons between the two cards (MI210 and A6000) are being made for high throughput workloads, and in this regime the performance is compute bound. So as long as the memory bandwidth is decent (as it is for GDDR6 with 768.0 GB/s) the lower memory bandwidth is not the main bottleneck. There are also other architectural differences so any comparison will ultimately be imperfect, but we found A6000 to be the closest match for the workloads that we cared about (ie high throughput workloads at reasonable latencies).

Also worth noting that there are still more stack optimizations on the table, which again can shift the bottleneck between compute and memory. In those cases it might makes sense to compare with another card with matched memory bandwidth.

paul @ mk1

paul_mk1··on MK-1
>The "-ngl 32" means that only 32 out of 35 layers are being run on the GPU, and this results in a huge slow down as the GPU syncs with the CPU, and then computes the last 3 layers on the CPU.

Thanks for the updated run configuration. It was a misunderstanding on our part about what llama.cpp considers “layers”, since layers are traditionally understood as learned parameter decoder layers (as they do in Hugging Face models). And, in this case the llama 7B model has 32 layers.

>On my XTX 7900, I get a 55% speed up on llama.cpp (to 132.61 tok/sec) when running all layers on the GPU, rather than only 32 as in your measurements.

On my 4090 I now get 128 t/s for Q5_1, and 116 t/s for Q6_k. So these are ballpark to mk600's 125 t/s for batch=1. Not surprising that different inference runtimes approach the same speed as they become memory bound for similar model sizes.

paul_mk1··on MK-1
Appreciate your response.

We compared MKML mk600 (5.2GB) against llama.cpp Q5_1 (4.7GB) and Q6_k (5.1GB) on a 4090 for llama-7B. The test is the same in all cases: we generate 128 tokens from a single token prompt (batch=1) and measure performance of the forward pass during auto-regression.

(llama-7B, single prompt, batch=1)

MKML mk600: 125t/s

Llama.cpp Q5_1: 8̶4̶ 128 t/s

Llama.cpp Q6_k: 7̶8̶ 116 t/s

Our llama.cpp test: Build (https://github.com/ggerganov/llama.cpp#cublas):

make -j12 LLAMA_CUBLAS=1

Run:

./main -t 16 -ngl 3̶2̶ 35 -m llama-2-7b-chat.ggmlv3.q6_K.bin -p "?" -n 128

Please feel free to post your llama.cpp results if they are different.

>Maybe I'm misunderstanding MKML. As I understood it, MKML is a compression step that then feeds into another framework like HF's Transformers or PyTorch.

MKML is not a compression tool that feeds into another framework. It is an inference runtime (like FasterTransformers or vllm) except that MKML is also plug and play with existing frameworks like Hugging Face.

paul_mk1··on MK-1
Hi, one of the founders here.

Attempting to address some of the comments in a single message.

To help understand why we decided not to compare to existing methods: I think it would be difficult to do so fairly, since there are many tradeoffs and different use cases. It's not always the case that one technique is bad and the other is good, it's more about the targeted design point (say, cloud vs local). We are openly offering our numbers / benchmarks and looking for early partners that are aligned with our current value proposition (hence the closed beta).

A good example is that llama.cpp is a fantastic framework to run models locally for the single-user case (batch=1). While llama.cpp supports different backends (RPi, CPU, GPU), I don't think it would be particularly fair to compare and show that MKML is better at a given perplexity, compression ratio, and speed on GPU for a multi-user case (batch >> 1), when that is not llama.cpp’s targeted use case (afaik). For example MKML achieves ~2700 tok/sec at batch 32 (i.e. 32 prompts in parallel) on a 4090 for a Llama-2 7B, with a ~4̶.̶2̶G̶B 5.2GB memory footprint, and perplexity that is ~fp16.

Also, we're not currently wrapping any open source tools or techniques for quantization. Everything is our own and there’s more news to come soon.

If anyone has specific technical questions I'd be happy to answer as best I can.

Cheers, Paul Merolla