You don't think it's possible that an LLM's internal machinery could decide that an underused-by-humans word should be used more frequently in output than it sees in input because it maps cleanly onto a frequently needed semantic? I think that's possible