[0] https://github.com/antimatter15/alpaca.cpp/issues/182
[1] https://news.ycombinator.com/item?id=35400066>Now, since my change is so new, it's possible my theory is wrong and this is just a bug. I don't actually understand the inner workings of LLaMA 30B well enough to know why it's sparse.
I haven't followed her work closely, but based on the links you shared, she sounds like she's doing the opposite of self-promotion and making outrageous claims. She's sharing the fact that she's observed an improvement while also disclosing her doubts that it could be experimental error. That's how open-source development is supposed to work.
So, currently, I have seen several extreme claims of Justine that turned out to be true (cosmopolitan libc, ape, llamafile all work as advertised), so I have a higher regard for Justine than the average developer.
You've claimed that Justine makes unwarranted claims, but the evidence you've shared doesn't support that accusation, so I have a lower regard for your claims than the average HN user.
> I'm glad you're happy with the fact that LLaMA 30B (a 20gb file) can be evaluated with only 4gb of memory usage!
The line you quoted occurs in a context where it is also implied that the low memory usage is a fact, and there might only be a bug insofar as that the model is being evaluated incorrectly. This is what is entailed by the assertion that it "is" sparse: that is, a big fraction of the parameters are not actually required to perform inference on the model.