Note that I ripped an equal amount of commit messages per language so the results aren't based on how many projects there are per language.
I like how he had to tweak the data collection process to make the visualization method fit.
I like how he had to tweak the data collection process to make the visualization method fit.
This is normal and has nothing to do with how you choose to represent it.
It would have been meaningless to show any graph or table saying 'Python has the most messages with profanity" if the amount of Python projects is 80% of all the projects out there.
He should just collect as many commit messages as possible, then divide the profanity count for each language by the commit message count. Because that has lower standard error [and no more bias] than what he did.