This is really needed in the NLP world. Most things are english only.
Good luck with Turkish. A very large lookup table may work for %90 cases but it will still fail. A basic `true` lemmatizer requires either a complex graph with rules, or an FST generated from it. input is searched through the graph and morphological disambiguator must be applied to result to pick the correct lemma.
Or Welsh. You want a canonical form for "wnaethpwyd"? Try "gwneud"! Some regexes and an exceptions list isn't going to cut it.
The more a language needs a lemmatizer for NLP the harder it is to write it.
Could someone explain what's hard in particular about canonical forms in Welsh?
spaCy excels here. English, German, Portuguese, and more.
I can personally suggest FreeLing [nlp.lsi.upc.edu/freeling]. It supports most languages in the iberian peninsula (spanish, portuguese, catalan (including multiple dialects), galician), and then it also includes english, italian, french, german, russian, croatian and slovene. It's also more flexible than other tools if you need to go outside those languages and customize something.
As of now wink-tokenizer (https://github.com/winkjs/wink-tokenizer) supports multiple scripts, therefore, it can also tokenize sentences in languages like Hindi, Marathi, French, German etc. We are working on extending multi-lingual support to other components including this lemmatizer.