LLaMa running at 5 tokens/second on a Pixel 6
twitter.com
twitter.com
But if it is a trimmed version, it is wong to call it LLaMa.
It does not seem fine.
It is incomprehensible and doesn’t match the results I’ve seen from 7B through 65B.
It is true that RLHF could improve it, and perhaps then this severe of optimization will seem fine.
Person A uses GPT3 to generate training data, and publish it on his blog, without representing it as human generated. Person A does not give permission for it to be used by Alpaca team.
Alpaca team comes along, scrape his blog, and uses it as training data, without permission from person A. Now this is fair use, so there is nothing person A can do to stop it, just like how Github scraped our code for Copilot without permission.
That would have the same licensing problems that they have though: that alpaca_data.json file was created using GPT3. But creating a "clean" training set of 52,000 examples doesn't feel impossible to me for the right group.
If the cloud of uncertainty around commercial use of derivative weights from LLaMA can be resolved, I think this could be the answer for a lot of domain-specific generative language needs. A model you can fine tune on your own data, and which you host and control, rather than depending on a cloud service not to arbitrarily up prices/close your account/apply unhelpful filters to the output/etc.
With a Markov chain, you're assuming a state machine where each state has independent probabilities on outgoing edges. As the number of states gets larger, you have fewer training samples for each state. When n gets large enough, nearly all states have zero training samples; they've never been seen before. How do you estimate probabilities?
Better to just say it's a stateless function of the input.
If you are talking about the video that's perfectly fluent English. There are some unusual elements to the story which probably wouldn't be there in a larger model.
I'd invite you to try that with a Markov model or even something like a LSTM based neural network and compare.
You're probably right (because why would they?) but I don't see any reason they couldn't have done this if they wanted to.
Currently typing this from my Pixel after running it countless times :)
There are so many use cases (like this) that require more RAM. And even if a use case doesn't theoretically require more RAM, getting a developer to dedicate time to optimizing RAM is time taken away from making a wonderful app.
Advanced hardware makes bullet points on advertising to sell the device; giving the bare minimum of RAM accelerates the device planned obsolescence, so that user will be forced to upgrade sooner to the next model.
Feature updates with the current IOS 16 goes back to the iPhone 8
Yeah you do lose feature updates and slowly app support after the latest version drops support, but it's not like they're dropping support after 2 years, and you can stay on it for years later if you'd like.
I'm not saying it couldn't be better but they're clearly far above the vast majority of their competition.
Old versions of iOS quickly stop working as apps demand updates, and the updates require a new iOS version.
Apple still does security updates for IOS - last was 12.5.7 - 23 Jan 2023 - that's back to the iPhone 5S
They've literally provided security updates for a 10 year old device, has any competitor even come close to that?
Functionally they're useless
So maybe if you implement the ggml.c with tensorflow/libcoral - you'd have a chance.
So we have numbers on PTB original perplexity 8.79 quantized 9.68, already 10% worse. And PPL reported per token I suppose? Because word PPL for PTB must be around 20, not less than 10.
Any numbers on more complex tasks then? like QA?
However I can see fractional bits (via binary representations) and larger models happening first before that compression step.
And then we have the sub-bit range..... ;DDDD
(unless this ggml library is doing that under the hood)
i assume it has unified memory, but maybe not little numbers...