LinguaCafe: Self-hosted software for language learners to read foreign languages
simjanos-dev.github.io
simjanos-dev.github.io
I didn't think this many people would be interested. I'll write a guide for Jellyfin, then add Italian, French and Dutch languages tomorrow.
I've added Chinese. However i couldn't find a dictionary for it yet, and it might need a custom font for Chinese characters. DeepL works with it as well. If it has issues, I will fix it soon.
I'll release a v0.4 update tomorrow or after. It makes a lot of things more simpler, I would recommend to wait for it before you install it. After that update I'll work on Chinese dictionaries and issues. It will take 1-2 days. It will have two built-in dictionaries for Chinese: cc-cedict and wiktionary.
French and Dutch are added.
And as someone who now also speaks Italian, I am even more pleased to see that Italian support will be added tomorrow.
It is wonderful to see such a useful tool released as an open-source, self-hosted project. (^_^)
[1] EDICT: http://edrdg.org/jmdict/edict_doc_2009.html
[2] JMdict: https://en.wikipedia.org/wiki/JMdict
Thank you for helping me learn Japanese! :)
Can you please explain what do you mean by actively supporting JMDict? I hope I didn't make an attribution mistake, or misunderstood something. My understanding is that I can use those files in my project as long as I follow the license guidelines.
It makes me really happy that so many people are interested in it. :)
I know Christmas is over, but my letter to Santa would include: - some Anki sync feature (over an external Anki sync server or any other solution) - a non-docker install guide - of course more languages!
I've been looking for a tool to study vocabulary this way, especially in languages I'm already fluent in, to learn more nuances or specific meanings to some words. Having tried several things I settled on the bookmark feature of my Wiktionary Android apps (Livio's, which are nice), and a small sync/script chain that would let me review words, compare definitions in different dictionaries, choose the best and edit/complete it, and make an Anki card of it. The whole process was still tedious.
Reasons it's useful: * If you've got both Native & Target Language subtitles, you can see a natural translation if you're struggling to understand something * If there isn't a Native translation, then you can machine-translate one - especially useful early on to catch common idioms/etc that aren't just the sum of each individual word. * Jellyfin also supports eBooks, although its reader isn't great - but if someone has already built their library, it would be nice to be able to re-use it somehow.
I would be very interested in seeing that particular feature expand, but I don't imagine it's at all simple!
Tangentially related, but I could see some desire for Calibre support as well, somehow. Calibre was very much designed to be completely stand-alone and it doesn't really support other apps trying to read its database, but it is possible.
I'd also really like some language-specific features, like separable-verb handling for German (see this comment: https://news.ycombinator.com/item?id=38915786) - it's relatively important and lacking support really limits the usefulness of vocab tools. It would also be a nightmare to handle for subtitles, since it's not always clear where a sentence ends, but such is life - subtitles are sadly not aimed at language leaners. For books and not-terrible Podcast transcripts, though, it wouldn't be so bad.
I thought of it as a niche feature because I thought most of the users would come from language learning communities, where most people are not into self-hosting. So even if someone would set up a server just for this, chances are they do not have or interested in Jellyfin also. But I've seen several comments about it, and it seems like a lot of people are from the self-hosting community so maybe it's more popular.
I'm also planning to support YouTube and improve on Jellyfin support, but I'll work on other issues and features first.
I definitely wouldn't expect it to be high on the list of priorities, but I do appreciate that it's under consideration at the very least.
It unfortunately does have false-positives (a complete solution would require LLMs, I believe over the much less complicated NLP algorithms - I just don't want to send whole books to ChatGPT, as that would quickly become expensive), but I found it usable, so I made it public now: https://github.com/tenaf0/lwt
I don't want to "advertise" it even more, as the NLP lib is run by academia as a free service, and I don't want to overburden it (I have been planning on hosting it myself, but didn't yet get there).
The fine-tuned mistral models are known to out-perform GPT-4 on their specific tasks.
https://en.wiktionary.org/wiki/f%C3%A4ngt_an
tells you what the word is, and gives a link back to:
https://en.wiktionary.org/wiki/anfangen#German
I chose to add Wiktionary to Kiwix Android (8GB download) for offline use. In addition, I can search by right-clicking or tap+holding on a word. All that information is available because of the (mostly manual) work done by Wiktionary contributors, but it reaches a very high standard. There is usually more digression and explanation for the usage notes in Wiktionary than, say, Collins German-English dictionary, which is a rather good thing for language learners.
There's a nice project for converting and extracting the data from English Wiktionary into JSON but it doesn't support any other languages, AFAIK, which is a bit of a shame but also not very surprising - Wiktionary is a lot more complex, technically, than I expected!
- the English Wiktionary has fewer English words than the German Wiktionary has German words, or
- the English Wiktionary has fewer German words than the German Wiktionary does?
Another example is "krächzender", which might also serve to give some idea of the particular pains in processing German text. It's not in English Wiktionary, but krächzen is, and is a verb. So "krächzender" is the adjectival form of the verb, and if you know "krächzen" and the general rules around adjective formation it would probably be obvious. But would you rely on a computer to parse those rules, or would you want a table with all the declensions laid out? And if you're building a vocab list for a book, is it a separate entry in the list, or does it fall under the verb?
Obviously, German Wiktionary only has definitions & explanations in German so it's not great for beginners, but any tool that's trying to automatically do stuff with German text would likely benefit from using German Wiktionary.
I have no idea if it's true for other languages, but I wouldn't be surprised if it's also true for other major languages spoken by Wikipedia users (e.g., French, Spanish, but maybe not Chinese).
Ah, found, it's here: https://github.com/simjanos-dev/LinguaCafe
[1] https://learnanylanguage.fandom.com/wiki/Listening-Reading_M...
I've added Chinese as an "experimental" language. I couldn't find a dictionary for it yet, and it might need a custom font type. DeepL works as well. I will fix the font issue soon.
1. Learn the translation of the commonly used everyday words
2. Learn the rule to build sentences in different tenses (Verb conjugation)
3. Keep practising in everyday conversations, starting with most simple ones and gradually learn more.
This clearly takes a lot of inspiration from LingQ, but fixes some of LingQ's more glaring challenges such as letting you use a real dictionary, instead of relying on definitions that were crowdsourced from other learners using the app. (And therefore full of quality problems an inaccuracies.) On the other hand, it sounds like some nice features aren't implemented yet, or maybe not even planned, so maybe LingQ is still a good option if you don't want to hassle with self-hosting a webapp or hunting down your own resources, and don't mind paying the subscription fee.
All in all, though, it looks very promising!
(Disclaimer: I haven't actually used LinguaCafe, but am a longtime LingQ user, so I'm not really making a fair comparison. I know LingQ's feature set much, much better.)
My current self study centers around movies/tv and Linq, which this tool seems very similar to.
I'm learning Dutch, so it's a bummer that it's not supported currently, but I'm keen to dig in and see how much effort it takes to add a new language.
We support many languages out of the box, would love to hear what's making you consider LinguaCafe over LingQ :)
I noticed a lot of signups from this post, I'll try to do another round of onboarding this month.
I study French, German, Swedish, Mandarin, Japanese, Portuguese, Latin, dabbling in Polish, though it’s hard to find shows dubbed in Latin :). Someday will get to Russian and maybe learn Icelandic as a way of getting closer to the roots of English… but alas life is not forever.
I’ve written some LLM-based software for generating podcasts (www.anyglot.com, but the server currently offline). This project showed me that GPT4 is excellent at generating content in English, and in doing various NLP tasks but not translation which was better left for Google. ElevenLabs voices are fantastic but their Japanese would invent weird Kanji readings, tho that was when Multilingual V2 just came out so maybe they’ve fixed that already.
I've added Dutch.
I've added French and Italian.
Japanese only but I am expanding it to more languages early this year.
What does adding a new language involve other than adding a dictionary? It doesn't seem like there are too many language-specific features at first blush.
1. Languages that have conjugation (I am -> you are -> she is...) need a way to recognize different forms of the same word, as dictionaries typically don't have this information. 2. In some languages, nearby words dramatically modify a word's meaning (in Spanish quito=I take away, me quito=I take off). The modifier words can appear quite far away in the sentence (especially in German I think). 3. Languages that don't put spaces between words are a nightmare. 4. Even for languages that do put spaces, the set of characters that act as a separator can differ slightly. 4. Languages like Chinese and Japanese with their enourmous 'alphabets' need UI changes to help learners with pronounciation of new characters. 5. Fonts and text entry! Does your framework support all 10-jillion Chinese characters and the 5+ different types of on-screen keyboards that people use? Did you remember that Arabic is read right-to-left?
Those are all of the main challenges I encounted implementing Spanish, French, Chinese and Japanese. I have no idea what new challenges would come up in Finish, Hindi, Swahili...
I would definitely use it if there were a Korean option!
I've added Korean as an "experimental" language. You can try it out if you would like to.
It’s great that you can track your progress in this app!
When I was learning German, I used the dictionary lookup on Kindle a lot and made a web app to extract that vocabulary as Anki flashcards. It’s available on https://fluentcards.com. The code is open source on GitHub.
Being the resident licensing pedant, I'll point out that neither of those repos have any licensing information aside from package.json and I doubt gravely that's strong enough for any contributor's comfort level
Normal desktop applications aren't self-hosted, as they aren't hosted in this sense to begin with.