- The weight releases were messed up: released Lora for Llama 3.0, claiming it was a 3.1 fine tune
- Evals initially didn't meet expectations when run on released weights
- The evals starting performing near/at SOTA when using a hosted endpoint
- Folks are finding clever ways to see what model is running on the endpoint (using model specific tokens, and model specific censoring). This post claims there's proof it's not running on their model, but just a prompt on Sonnet 3.5
- After it was caught and posted as being Sonnet, it stop reproducing. Then others in the thread claimed to find evidence he just switched the hosted model to GPT 4o using similar techniques.
Lots of mixed results, inconsistent repos, and general confusion from the bad weight releases. Lots of wasted time. Not clear what's true and what's not.