https://www.pnas.org/doi/full/10.1073/pnas.2016239118
They found representations on fundamental properties of proteins such as secondary structure, contacts, and biological activity in an LLM only trained to predict protein structure sequences. No folding, nothing explicit in the corpus itself. Yet those truths manifest as a necessity of accurate predictions. Prediction is not bound by what the data explicitly shows.
>A language model’s vocabulary is limited to the words that exist within the model’s training texts, which means a LLM can only refer to objects and relations that we humans have already discerned, named, and written about.
This is not true and is easy enough to test.