I think you might see some interesting results using a non-parametric [one that doesn't require specifying the number of clusters apriori] clustering algorithm, like mean-shift. I've never seen an adaptation for discrete data like this, but it should be possible.
You could have the same tradeoff between note-distance and time-distance.