If a diff is not useful, don't store it in git. Doing a code review on PII is not useful.
You also shouldn't add packages in a git repo if they can be downloaded immutably in CI/locally. It adds a lot of space. Every commit is a snapshot + new code. Commits are not diffs.
Similarly, I keep small pdf manuals in git. I add them in their own commit, which doesn't have a useful diff. In exchange, they're always there and I don't need to spin up some special one-off system.
I was replying to a comment asking why PII data shouldn't be stored in git vs a database. Just tried to give a general rundown of how to think about the tools.
Totally agreed that the data should have been synthetic/scrubbed and that a small sample set (large enough to validate) would be ok.
Sometimes a PDF or a few pictures are necessary (or just much simpler). I get that. I have seen some repos with an unreasonable amount of PDFs/Docs/pictures/etc.. and that's when a script that copies them into a gitignored directory from someplace (Dropbox/S3/etc..) is a better fit.
Putting PII into Github (even if the repo is private) is catastrophic. You just made Github into a third-party data processor by accident. Good luck explaining that from a CCPA/GDPR perspective.