From the video output seems fine.
But if it is a trimmed version, it is wong to call it LLaMa.
But if it is a trimmed version, it is wong to call it LLaMa.
It does not seem fine.
It is incomprehensible and doesn’t match the results I’ve seen from 7B through 65B.
It is true that RLHF could improve it, and perhaps then this severe of optimization will seem fine.