Full threadlavp·What does “perform slightly better than Llama” mean exactly? A model like this needs to be trained from scratch right?View on HN