How exactly was the data evaluated? I would assume that manually checking every speech would be too labor-intensive?
I would even go so far as to say that _not_ using LLMs for this task would be fairly odd, unless I'm missing something or the author really enjoys a month of manually classifying documents to write an interesting and well-written but not exceedingly outstanding article.