Guides like these inspire me to quit being lazy and get back to writing one!
Helping others wrangle data is one of the reasons I publish my Jupyter notebooks open-sourced. A few examples my data wrangling with R:
Processing Stack Overflow Developer data: https://github.com/minimaxir/stack-overflow-survey/blob/mast...
Identifying related Reddit Subreddits: https://github.com/minimaxir/subreddit-related/blob/master/f...
Determining correlation between genders of lead actors of movies on box office revenue: https://github.com/minimaxir/movie-gender/blob/master/movie_...
I'd add that Kaggle is very good for the "other end" of data science: they generally have pretty clean data, and clear problem descriptions.
In real life the data is never clean and the problems are rarely known in advance.