HNHacker News
TopNewBestAskShowJobs

polishgladiator

18 karma · joined August 5, 2023

submissionscomments
polishgladiator··on MK-1
Something doesn't smell right.

Such sloppy errors with measurement and comparison (from people who are supposedly experts?), and cageyness about answering technical questions, reminds me of the era of crypto currency scams..

polishgladiator··on MK-1
Actually no -- that post shows they are not performing measurements and comparisons correctly.

These are not serious people.

polishgladiator··on MK-1
OK, so this is a case of bad measurement and comparison.

If you bothered to look at the llama.cpp output, you would see this line: llama_model_load_internal: offloaded 32/35 layers to GPU

The "-ngl 32" means that only 32 out of 35 layers are being run on the GPU, and this results in a huge slow down as the GPU syncs with the CPU, and then computes the last 3 layers on the CPU.

On my XTX 7900, I get a 55% speed up on llama.cpp (to 132.61 tok/sec) when running all layers on the GPU, rather than only 32 as in your measurements.

polishgladiator··on MK-1
> [...] llama.cpp is a fantastic framework to run models locally for the single-user case (batch=1) > [...] I don't think it would be particularly fair to compare and show that MKML is better at a given perplexity, compression ratio, and speed on GPU for a multi-user case (batch >> 1)

Ok so you agree that llama.cpp etc are great for batch==1, right?

And I agree their targeted use case is not batch==32 (because who is doing that really?)

But if we extended llama.cpp or some other faster batch==1 implementation to support batch==32, why do you suppose it wouldn't still be faster than MKML? It seems to me that if you can do batch==1 faster, you could easily do batch>>1 faster too -- it is just that no one really needed that (yet?)

polishgladiator··on MK-1
> If anyone has specific technical questions I'd be happy to answer as best I can.

What is the context size for these measurements? Is it the full 4k for llama-2? And just to be clear, when you say memory footprint, this is the entire memory foorprint right? Weights, 4k KV cache etc?

And more generally, I'm curious about the use case for running puny models like Llama-2 7B in the cloud on desktops GPUs (like 4090) with batch==32?

polishgladiator··on MK-1
Based on the integration examples, I don't think they are simply repackaging llama.cpp

Rather it looks like they are reimplementing their own quantization scheme, in such a way that it is a little easier to integrate for basic python users, at the cost of performance (compared to llama.cpp and others).

Given that the bar for integrating something with higher perf like llama.cpp isn't very high (and that's the way the world is heading -- ask any 15 year old interested in this stuff), I can't see anything of value here.

polishgladiator··on MK-1
I've been doing some hacking with Llama2 on an AMD 7900 XTX this weekend, using llama.cpp and q5_k_s quantization.

Compared to MK600 on an RTX 4090 in their data, I am measuring higher throughput and lower perplexity (again, note that I am using a cheaper GPU!)...

polishgladiator··on Bram Moolenaar has died
In the realm of text editors, Vim has been my unwavering companion for 25 years. It has become an integral part of my daily existence, an inseparable bond that defies the imagination of life without its embrace.

Farewell Bram!