We are working on this - we don't have quite enough Irish speech data.
12 karma · joined June 25, 2024
Also, the training dataset is highly imbalanced and Spanish is the most common class, so the model predicts it as a sort of default when it isn't confident -- this could lead to artifacts in the reduced 3d space.