Only in arXiv you could get away with that kind of language :). Good paper though! Kudos.
"Another direction to go from here would be to increase the size of the context window during the data preprocessing stage to feed even more contextual information into the model."
Could you comment on how the training time would scale with increasing the size of the context window? Is there a sweet spot?