On Anki's Database
natemeyvis.com
natemeyvis.com
A few extra tidbits:
- A few versions back the cards' "ease" field (that is - how a card was graded) meant something different depending on which phase the card was at the time (so sometimes "2" meant "hard" and sometimes "2" meant "okay"). It was finally fixed and AFAIK in new versions it's consistent now, but apparently the migration didn't always work properly and I still sometimes see databases where the grading is the other way around compared to what it's supposed to be, and I need to heuristically detect that this is the case and handle it.
- Initially JSON blobs were used to store a lot of data; relatively recently that was changed so that it's stored as proper tables, but not completely, so a lot of data's still in the blobs, but this time instead of JSON it's protobuf. (Which seems strange to me considering SQLite has native support for JSON.)
It's a good thing the schema's slowly being cleaned up, but unfortunately it's only done incrementally, so every time any little thing changes I need to add yet another special case to my importer to handle it, and often in various permutations too because some databases are half migrated Frankensteins. (Don't ask me how that happens; I don't know. Maybe it's an issue of people using outdated plugins with their Anki installation, or copying their database between multiple independent Anki implementations, or maybe the current phase of the moon's just wrong.)
I can't personally use JPDB (due to my own niche learning strategy, not a flaw in JPDB), but I desperately want to be able to consume the underlying data. It's just that good -- the data that you've curated is unbeatable. If you ever provide a public API, I'll join your Patreon in a heartbeat.
Yes, an API will be coming in the future! (:
> due to my own niche learning strategy, not a flaw in JPDB
Just for curiosity's sake - what kind of strategy is it, if I may ask? I have a very ambitious plans for the future, so depending on what exactly it is it might be possible someday.
I have always had the most success with Anki and Wanikani when it comes to Japanese. Trying to add in yet another paradigm for learning is frustrating. I appreciate you've put a lot of effort into helping people move from those tools, but I don't want to.
The single biggest reason is offline access. Anki works on my phone on an airplane or in an area with no mobile service. (In Australia there are lots of those).
I only started using WK seriously when I discovered the Android apps that let me do my reviews offline.
If you had a Patreon tier that allowed for Anki exports of your lists I'd sign up in a heartbeat even if it only allowed 1 download per month of something similar. I mean lets be honest what possible valid need could I have for downloading the whole data set in one go...
Compared to the time it would take me to use Subs2srs across a season of a show I'd rather just give my money to you.
I always ask this not because I necessarily want to convert people to use my thing, but because I always love to hear what features people need and what I can improve. In case of offline access it is something that's technically on my tentative roadmap, but very far off into the future, so indeed for anyone who needs that Anki's the better choice.
> allowed for Anki exports
That's something that I'm planning to add very soon actually! Well, maybe not exactly Anki exports (I haven't yet researched as to what that would entail), but just generic functionality to be able to export the built-in decks as a .csv (which I'll be happy to tweak/improve to make it easier to import).
I'm following an eclectic strategy where I isolate and separately learn spoken and written Japanese. The process looks a little like this:
1. I start with a deck of Anki vocabularly notes that I want to acquire
2. Study begins with "Speech" Anki cards from these notes (The card front is audio-only, including a clip of the word and a clip of an example sentence. The back has the English definition & a helper image). I only consider a card as being "Good" once I am able to recall & replicate the pitch accent with a steady rythm (I pipe back delayed audio from my microphone while I practice with a metronome running)
3. In parallel, I also do Kanji isolation study using KKLC
4. Each week, I manually enable new "Writing" Anki cards that come from the same set of notes (The writing is on the front. Only the word audio is on the back). I only enable a "Writing" Anki card if I have previously learned BOTH the component Kanji and the spoken word
5. I study my enabled "Writing" Anki cards in parallel with the other two tracks
I like this approach because I effectively have three separate learning tracks that I can switch between -- the variety keeps me motivated. It also helps train your ear to be able to distinguish homophones by pitch and leads you to think of 同訓異字 writings as variations of a spoken word, rather than as true homophones.
I'd definitely like to expand the configurability of jpdb up to a point where you'll actually be able to do something like this in the future. Unfortunately that's not going to be anytime soon, so you're definitely better off with sticking with what you have now. (The most immediate feature that I have planned soon-ish are pure kanji decks; the necessary customizability for the rest will come much later.)
I'm incredibly thankful that Anki exists, I don't think I would have ever learned Japanese to a high level without it. But having spent some time looking at its guts myself, it sure is a mess in there. I thought at one point about building something on top of Anki, but decided against it discovering some of the same stuff that has already been mentioned.
Does anyone know where one can find high quality public code reviews? I imagine there must be open source projects on GitHub with good public feedback on pull requests. Any ideas of specific projects to look at?
Finally, I haven't found much information on database migration best practices. Any good articles, books, or other resources people would recommend?
And now he has spent years trying to slowly undo the stuff he baked into it at the lowest levels.
It's fine to do stuff independently without knowing, but please seek feedback and advice on crucial design decisions.
Or at least isolate them from the rest of the code. Of course, it's difficult to recognize which design decisions are crucial without already being an expert, so this sort of advice is probably silly. It's probably also the kind of advice that would have been more likely to strangle Anki in its crib rather than giving us the opportunity to discuss how a wildly successful program that has helped many people learn many things should improve its data model.
I'm not saying they need an expert or they shouldn't push forward anyway, but it's good to recognize when you're out of your depth and to get input when possible. There's a lot of value from just running something by someone else to see if they can understand it and if something occurs to them that would make things simpler.
Thanks for the note! Unfortunately Eric's reference is not publicly available (as far as I know).
I'd love for more people to collect and publicize sets of commonly used code review notes.
Others here will know much more than I do about which publicly viewable projects have the best (public) feedback. Good luck!
Yes, I agree. I hope I was sufficiently clear about that. My motivations for writing this up were that (1) it's a wonderful example of a lot of things I often find myself saying about databases; (2) peoples' Anki decks are really important to them and I'd like to enable people to work with them if they want; (3) I do hit strange behavior sometimes in Anki and perhaps these issues are behind it.
I hope you're backing up your deck!
Of course, usually changing the schema shouldn't lose data, so there's no need to be afraid of it in general.
1 - A database schema requires quite a bit of thought and up-front design, which we're loathe to do in this agile world.
2 - Schema migrations are hard and scary, so we avoid doing them
I'd guess that (2) is more important than (1) here. Or, at least, it seems anecdotally right to me that programmers are very hesitant to do migrations.
As for (1), you must be right that a disconnect between a design and eventual use is important here, but I'd have said that the more central cause is just that systems change and that their eventual use is very hard to predict. I need to think more about your conjecture that there's a misfit between programming practice and what's necessary for database work; I like the idea that we're habituated to implement something very basic up front, and that this habit intersects particularly badly with database work (that is, it causes problems that aren't nearly as bad with non-database parts of computer systems).
Especially importing media files and de-renaming them was a pain, as well as handling the different types of cloze deletes (some of this is described quite well in anki's docs, for example here https://docs.ankiweb.net/#/templates/generation)
Another link on their DB structure which saved me a lot of time was this one: https://github.com/ankidroid/Anki-Android/wiki/Database-Stru...
The app I am working on is focused on language study. A big drawback of Anki when it comes to language study is that the card is tightly coupled to the content it's trying to teach you.
In our implemntation, we decided to decouple the presentation of what you're studying and how you're studying. We have the concept of a term which can be a word, definition, sentence or character. We store the progress information along with term. The term can then be studied with any of the available study methods. This gives quite a bit of flexibility in how the information is studied and reviewed and goes a lot further than just flash cards and multiple choice questions.
Another benefit is that this progress information can then be used to make recommendations to the user on what to study next.
We're working on a Thai language course right now and hope to be rolling parts of it out in the near future. Other languages are coming later. The link is in my bio, but there's nothing to play around with yet.
Authors bio links to https://emurse.io/
EDIT: Never mind, was able to get behind a computer and yes, I did use Mochi before. And it does seem you support MathJax. Awesome!
If everyone is using such libraries, the structure itself can be changed.
The AnkiConnect project[1] is about the closest that we get to that right now, but requires running everything through a server inside Anki.
I began a feeble attempt at a lib for my own purposes, but I had to take Anki db structs a little at a time just to keep my frustration in check.
There are in fact some libs for writing Anki data: https://github.com/kerrickstaley/genanki and maybe https://github.com/patarapolw/AnkiTools — but genanki, while looking quite good for one-time generation, doesn't seem to be able to update notes.
Afaik Anki itself does include Python libs for creating and manipulating db records, which can be used in third-party scripts—however dunno if they work without the full app running, and on top of that I personally keep trying to use Lua, since it runs circles around Python in terms of speed.
[1] https://github.com/kerrickstaley/genanki
[2] https://www.reddit.com/r/Anki/comments/g0zgyc/spotify_anki_l...
But yeah, that's what I've used as well.
I've seen the author reply here, if you don't mind me asking, are there other ways to represent this database? Is there a useful exercise I can perform to do so? e.g. NoSQL version or similar, thank you.
There are definitely other ways to represent this data! I think a useful exercise is:
(1) Sit down and think of a very basic representation of this data (in whatever language you prefer). (2) Figure out where in the current SQLite database the relevant information lives. (3) Write a function that is given some relevant set of rows from this database and returns an item in the representation you determined in (1). (4) Test it. (5) Think about how to persist it.
You can do some subset or superset of these as you see fit, but I do think it's valuable to think about how to represent the relevant objects before you worry about persistence details.
Thanks for your comment!
What’s the benefit of getting knees deep (other than out of pure interest)
God yes. Why do people do this? I recently moved into hardware and it's even worse! Someone abbreviated a TLA to a single letter. Like instead of CPU_CLOCK it was CCK. Madness. These aren't printed on a silkscreen!
> compute the Perron-Frobenius eigenvector of a graph of medical school Anki cards based on an automated tagging system that tokenizes the cards
> change medical education forever