LiteLlama-460M-1T has 460M parameters trained with 1T tokens
huggingface.co
huggingface.co
> My name is Merve and my favorite color is blue. I am a very active person and I love to exercise. I am a member of the American Red Cross and I am a member of the American Red Cross. I am a member of the American Red Cross and I am a member of the American Red Cross. I am a member ...
If I prime it with a pattern of Q and A with several examples of a good question and a good answer, then a final question with a blank answer, it provides a more cohesive and less repetitious answer. To me, this indicates it just needs fine-tuning with an instruction dataset to perform more "normally".
Obviously, a model this tiny is going to have very little inherent knowledge, and it's probably not going to be great at anything other than speed.
Q: What is the capital city of France?
A: The capital city of France is Paris, a global center for art, fashion, gastronomy, and culture. It is located in the north-central part of the country and has an estimated population of 12.4 million in its wider metropolitan area.
Q: Who was the first person to walk on the moon?
A: The first person to walk on the moon was Neil Armstrong, as part of the Apollo 11 mission on July 20, 1969. When he stepped onto the lunar surface, he famously said, "That's one small step for man, one giant leap for mankind."
Q: What is the chemical symbol for gold?
A: The chemical symbol for gold is Au, derived from its Latin name, aurum, which means 'shining dawn.' Gold is a transition metal that is bright, slightly reddish yellow, dense, malleable, and resistant to corrosion and tarnish.
Q: What is the tallest mountain in the world?
A: The tallest mountain in the world is Mount Everest, located on the border of Nepal and Tibet. It has a height of approximately 8,849 meters (29,032 feet) above sea level. The first recorded summit of Everest was achieved by Sir Edmund Hillary and Tenzing Norgay in 1953.
Q: What is the LHC? What has it accomplished?
A:
And this is the output I received from LiteLlama: The Large Hadron Collider (LHC) is a particle accelerator at CERN in Geneva, Switzerland. It is a super-high-energy proton-proton collider that is designed to collide protons with protons at the LHC.
Q: [Omitted since I don't care. You would normally stop the model when it responds with a stop word like "Q:", which it did.]
Obviously that response is not going to blow anyone away, but it didn't get stuck in a loop. It knew that LHC stood for "Large Hadron Collider", and it also knew that CERN and Geneva were relevant.I was a mountain.
I was a mountain.
I was a mountain.
I was a mountain.
I ran the same prompt with a repetition penalty of 1.1 and got this:
The mountain in front of me was a little higher than the other two, and I could see the top of the mountain.
I'm going to climb it," I said. "It's a good place for a picnic."
You're not going to climb it?"
No!" I said. "But I'll climb it if you want! I can climb it with you—if we can get out there together—"
In any event, I don't believe this is means as a general purpose chatbot or something, it's an experiment with a small model and possibly something that could be fine tuned for a narrow task. There's not much point comparing it to much bigger models.
Answer: 0 + 10 = 10 + 10 = 10 + 10 = 10 + 10 = 10 + 10 =
This seems like a waste of compute and time.
You could instead ask it how to calculate something, and it could give you accurate instructions for how to achieve that. Then you either perform the calculation yourself, or use something like ChatGPT that has a built in python evaluator, so it can perform the calculation.
Quick example: https://chat.openai.com/share/9f76f5e5-d933-48fb-99e8-4a6530...
You're correct, LLMs are not (usually) explicitly trained to solve math calculations, this does not mean they cannot solve basic math equations (they can!).
Oh ? So what is it then ? Magic ? When you give GPT-4 random multidigit arithmetic that would not have appeared in it's dataset and it's more accurate than you can manage without a calculator, what is that ?
"Pattern matching" has really lost all meaning.
Basically, if you can write the list of operations and base it on digits, the network can learn to replicate that. Actual calculation on values - not really. You can tell which mode is used by using large random numbers and seeing what kind of error happens. Missed digits - it's not actually doing the calculation.
Also, we don't know how much gpt4 cheats at math. OpenAI may be replacing those bits with external calculation.
Uh yes really. https://www.alignmentforum.org/posts/N6WM6hs7RQMKDhYjB/a-mec...
>You can tell which mode is used by using large random numbers and seeing what kind of error happens. Missed digits - it's not actually doing the calculation.
Missed digits doesn't mean anything other than the wrong calculation. The assertion that it isn't doing calculation is absurd.
>Also, we don't know how much gpt4 cheats at math. OpenAI may be replacing those bits with external calculation
Pick one. Either it's so bad it's obviously not doing any calculations at all or it's so good you suspect Open AI are passing numbers through a calculator behind the scenes. Not only is the latter baseless speculation, the 2 options are mutually exclusive.
This is based on integer addition up to 113 with a data set that only trains that. Yes, you can achieve that in a special case. No, LLMs don't achieve that.
Conclusion from the post you linked: "Epistemic status: I feel confident in the empirical results, but the generalisation to non-toy settings is more speculative"
> Pick one. Either it's so bad it's obviously not doing any calculations at all or it's so good you suspect Open AI are passing numbers through a calculator behind the scenes.
That's not the claim. I'm saying LLMs don't do real, precise calculation, but can do digit level operation and some implementations of systems around an LLM could pass the problem through a calculator. LLMs can rewrite the problem into a function, but the execution of it doesn't happen inside LLMs.
If you squint it's like the json output - LLMs can kind of do it most of the time, but you can implement a system around the LLM which ensures that any json output will be valid.
Q: 1+1
A: 2
Q: 1+2
A: 3
Q: 3+2
A: 5
Q: 9+8
A: 17
Q: 10+2
A: 12
Q: 10+10
and it spits out A: 20
Q: 20+10
A: 30
Not too bad.