(Bidirectional Encoder Representations from Transformers)
(Bidirectional Encoder Representations from Transformers)
GPT predicts the next word by only look back at what we have seen so far. In other words, it's auto regressive.
> [W]e split training documents at the character level into a prefix, a middle part[,] and a suffix with the splitting locations sampled independently from a uniform distribution over the document length. We apply this transformation with a probability of 0.9 and to documents that are not cut across multiple model contexts only. We randomly format half of the splits in the prefix-suffix-middle (PSM) format and the other half in the compatible suffix-prefix-middle (SPM) format described in Bavarian et al. (2022, App. D). We extend Llama 2’s tokenizer with four special tokens that mark the beginning of the prefix, the middle part or the suffix, and the end of the infilling span
Tokens are super powerful :)
> As an example, our model would complete the string 'enu' with 'emrate' instead of 'merate' which shows awareness of the logical situation of the code but incomplete understanding of how tokens map to character-level spelling.
that doesn't really feel like a failure of language modeling to me
> Note, however, that the results in random span infilling are significantly worse in suffix-prefix-middle (SPM) format than in prefix-suffix-middle (PSM) format as it would require token healing (Microsoft, 2023),
Talk about a negative endorsement. I am continually disappointed in the auto correction implementation.
https://en.wikipedia.org/wiki/Bit_error_rate#Bit_error_rate_...
No. There are at least two kinds of costs. First, It takes time to search 'adjacent' domains. Second, by reducing your available acronyms/initialisms, you make it harder to map your architecture name onto those letters.
It is fun to think of some of the alternative BERT names that "could have been", such as BIDET = BIDirectional Encoder representations from Transformers.
Now my disappointment is immeassurable and my day is ruined.