ParentFull threadrubymamis·Did anyone else notice the huge gap between scores on private vs public for ALL Jev-like models compared to LLMs (such as GPT Luna)? Doesn't it mean those models aren't generalizing so not very useful on data they haven't seen?View on HN