Qwen2-Math
qwenlm.github.io
qwenlm.github.io
The problem involves players removing stones sequentially and asking which will win with perfect play: the listed answer definitely doesn’t list all possible types of strategies.
The answer it gives may be right; in fact I bet it is correct (the second player), but does the qwen team offer the solution as correct including the logic? And is the solution logic correct?
It makes a bunch of specious claims about parity. (Adding 4 to a number changes the parity? Therefore, the parity of each colour always changes for each turn? Therefore, the parity will always be the same as it was initially?) And then it concludes that since the parity is right, 2022 of each colour must be a reachable state.
Which, as you say, is quite possible the correct answer, but it's really weird to put it out there as an example with no comment on the reasoning.
1) Much easier to state the problem and basically all the knowledge we have of math is not in the form of Lean proofs
2) It can be applied to a much broader range of domains, math is kinda unique as in most cases verifying something is 100% correct is impossible (and getting such a signal for RL)
2) How is a solution by LLMs supposed to be verified without such a formalisation?
2) This model is already specialized for math; applying it to other domains is out of scope. And as you point out, speaking Lean (and thus being verifiable) gives you an RL reward signal that's way more precise and readily available than piddly RLHF from human reviewers. Gyms like this are where AGI will happen.
(P.S. if anyone wants to work with me on reproducing AlphaProof, hit me up on Discord.)
Similarly on the Martian question "If we transform 1 red and 1 green Martian, we get 4 blue Martians. This changes the parity of red and green Martians from even to odd, and the parity of blue Martians from even to odd." - is complete nonsense too.
"the sum of the digits of a number k modulo 2 is equivalent to k mod 2" - 12?
Basically all "solutions" are regurgitated techniques from math competitions which are used completely incorrectly but with a lot of confidence
So at 32 bit full precision, 70 * (32 / 8) ~= 280GB
fp16, 70 * (16 / 8) ~= 140GB
8 bit, 70 * (8 / 8) ~= 70GB
4 bit, 70 * (4 / 8) ~= 35GB
However in things like llama.cpp quants sometimes it's mixed so some of the weights are Q5, some Q4, etc, so you usually want to take the higher number.
It's not really all that consistent, but larger models can be compressed more without as much loss.
2002 = 10^3 + 10^3 + 1^3 + 1^3
Then, multiply through by 2002^2001, which is itself a cube since 2001 is divisible by 3.
It's a useful to hone your creative thinking and learning how to approach a mathematical problem, but it wont make you a mathematician.
I'm a bit surprised that an HN reader, who presumably has first hand experience with them, isn't sure if its just a hashmap lookup.
The irony.
7B: https://model.box/try/playground/qwen/qwen2-math-7b-instruct 1.5b: https://model.box/try/playground/qwen/qwen2-math-1.5b-instru...