Understanding and explaining accuracy in machine learning systems
engineering.wootric.com
engineering.wootric.com
In other words, what else should I ask to validate that a 70% F1-score is better than a 90% F1-score but on a smaller data set?
To formally answer your question, the main things that matter in determining how stable your F1-Score from your test set is are: - Size of the test set - % of test set that has the label (in our case feedback tag) - the values found for precision and recall
However, if you are referring to noise as in typos and misspellings - then yes, depending on the training and/or preprocessing steps, classification systems could potentially reduce the noise in the input data to still achieve good results.