Why does having comparable performance indicate having been trained on a preexisting model's output?
I read a similar claim in relation to another model in the past, so I'm just curious how this works technically.
O1's reasoning traces aren't even shown, are you suggesting they've somehow exfiltrated them?
Unfortunately, it does not have "unified memory", a somewhat "powerful GPU", and of course no local LLM hype behind it.
Instead, I've decided to purchase a laptop with 128GB RAM with $2,500 and then another $2,160 for 10 years Claude subscription, so I can actually use my 128GB RAM at the same time as using a LLM.
GB10, or DIGITS, is $3,000 for 1 PFLOP (@4-bit) and 128GB unified memory. Storage configurable up to 4TB.
Can be paired to run 405B (4-bit), probably not very fast though (memory bandwidth is slower than a typical GPU's, and is the main bottleneck for LLM inference).