First, the original data was constructed like this: "...The next step was to format the raw HTML files into the full chord progression of each song, collapsing repeating identical chords into a single chord (’A G G A’ became ’A G A’)..."
Already this makes no sense - the fact that a chord is repeated isn't some sort of typo (though maybe it is on UltimateGuitar). For example, a blues might have a progression C7 F7 C7 C7 - the fact that C7 is repeated is part of the blues form. See song 225 from the dataset, which is a blues:
A7 D7 A7 D7 A7 E7 D7 A7
Should really be:
A7 D7 A7 A7 D7 D7 A7 A7 E7 D7 A7 A7
With these omissions, it's a lot harder to understand the underlying harmony of these songs.
The second problem is that we don't really analyze songs so much by the chords themselves, but the relationships between chords. A next step would be to convert each song from chords to roman numerals so we can understand common patterns of how songs are constructed. Maybe a weekend project.
[1] https://arxiv.org/pdf/2410.22046 [2] https://huggingface.co/datasets/ailsntua/Chordonomicon/blob/...