Pretty cool! Do you think the model would be good at other under-served languages as well? Or is it hypertuned to just these?
Currently the model is only given data for these languages so it doesn't know anything else.
À crawler and data ingestion pipeline will not help with that?