There's two approaches: musicologist and popularity.
The musicologists do a serious study and have sophisticated tools I can't pretend to understand but their results sound similar: so similar that the serendipity and adventure is sucked out of it.
Also they suffer from the generality problem as well. Qualities that matter in one genre, such as airy female vocals in a minor key, or whatever, are absolutely irrelevant for another genre. Take for instance, Irish folk music where it may be signal and say, Acapella, where it may be noise. Thus considering them is both right and wrong and we're back to our problem.
If you want to categorize the music to "fix" the problem, you run into binning issues. Let's say early 90s hardcore; they can have trance, breakbeat, dnb, house and jazz sections in a single song. Good luck trying to use your genre based contextual mapping on an unsupervised model.
The popularity approach, which ignores the content, is deceptive because it appears in many forms: people who listened to X also listened to Y or Y is trending or any of a number of variations where some magnitude of humans or temporal delta is used as signal.
These all tend towards the not long end of the long tail and so you eventually get the same mediocre experience - you start with your obscure prog rock group from the 1960s and 20 songs later you're at Cream or Hendrix. The "solution" is to tamper the drift via clustering but it will tend the same directions.
In practice though, these approaches service the majority of tastes, that is familiarity, so they're fit for purpose.