what would be a typical large file you want to have in your codebase? usually this shouldn't be a consideration
Also, deep learning training data often consists of large image files, and can also be considered "source code", and in any case it can be very useful to put these under version control.
And finally it can be useful to put external dependencies as tar-files into your source tree.
For writing tests in a deep learning code base, rather than simply including a native data file (image, CSV, whatever), I've taken to writing a fake data creator class. It always feels like overkill when an alternative solution is including a native data file or two that already exists.
It is a pile of garbage, but it's better than nothing.