Transformer-XL: Unleashing the Potential of Attention Models
ai.googleblog.com
ai.googleblog.com
Does that include the size of the model, or is this just the cost to encode the errors?
tldr: the measurements here are not meaningful for a comparison, read https://encode.ru/threads/3059-How-much-further-can-the-best...