The difference in timbres between a baroque violin and a modern violin, what pitch the strings (if any) are tuned to, the difference in group texture a string quartet and a string orchestra, differences in contrapuntal texture between say, Bach and Mozart, or the difference in harmonic styles between Italian and German music from the same time period.
For example, if you hear a string quartet, and it's tuned to A425 instead of A440, you're listening to a historically accurate recording of something most likely written in Germany in the 18th century. Probably Mozart or Haydn. If there's any meaningful melodic activity coming from anyone but the first violinist of the quartet, it's probably late Mozart or late Haydn.
Of course, if the quartet is tuned at modern pitches, it's a setback because most recordings are still done at modern pitches, so it could be anything from early Haydn to Shostakovich or George Crumb. So you'd have to rely on other features. And then some things are just impossibly difficult to infer. Fauré was French but mostly fascinated with contemporary German music from the Weimar school from an early age. Absent specific knowledge of his music, it would be insanely difficult to ferret out the harmonic and contrapuntal features that define French music of his time.
Music History is littered with edge cases like that. In the late 16th century there was a school of composition known for extreme chromaticism and hyper complex rhythmic structures. If you play some of this music for even well-trained professional classical musicians, almost all of them will be fooled into guessing that it's early 20th century atonal music from the Second Viennese school.
I can't think of any practical use case for this. I'm mostly just curious what the limits are for machine learning at this point. I've been "learning" these kinds of stylistic differences for 30 years now, though I mostly practice with the radio now on a classical station. Turn it on for one second, and everyone in the car gets a few seconds to guess. It's possible for a dedicated human be good at this.
I guess I'm just curious what the current state of ML is at the moment with regard to a subject I'm intimately familiar with. I worked on real-time fraud detection at a credit card precessing company years ago, but a fraudulent transaction is harder for me to relate to than a piece of music. I'm hoping that this turns out to be a good learning experience for me to be able to get a batter handle on what happens when you tweak certain knobs in an ML model.
I'm also uncertain as to what makes for an apples-to-apples comparison between the classifier and a human tester. My first response was to not use an api for metadata because that seems like it's unfair to be able to instantly lookup the name and composer of a piece. But on the other hand, an experienced human will a lot of performing experience has basically absorbed all of that metadata. So it really is possible for a person to get a 90-95% or even better accuracy rate based on 1-second samples.
Anyway, that's enough of that. When I get a chance to play with this, I'll keep in touch about what I come up with. Thanks again for putting this out there.