Creating multi-game highscore lists in ArangoDB
arangodb.com
arangodb.com
Is that actually a good idea though? Leaderboard queries are likely to be the most common things you do with the data. You'd save a join by storing the username with the score data (as you'd never want one without the other), and it wouldn't save that much space by separating them given that usernames are only a few tens of bytes each.
That could be the name in a particular game or other data from the users profile. Or - limit the leaderboard by filtering on age, region or whatever.
Optimising a database for storage is really only a good idea when the additional redundant data is going to take up huge amounts of space. Storage is cheap. Making a query run faster isn't. Generally speaking you're better off optimising to reduce the number of joins and subqueries because they're the things that usually slow you down.
The way I like to approach performance issues like these is to cache or precompute de normalized data as needed, and make sure the master data is as clean and normalized as possible. You could have a materialized view, or a trigger that produced the denormalized version, or just store the data is Memcached.
Problem is, we must stick with the opentech versions since we are running in a windows environment.
It appears the 'fork' hacking those guys made is not very stable. Every once in a while, the fork process fails for some unknown reason, and we are actually trying to fix the code ourself to understand why this fails.
Since we have spent a lot of time on it, maybe it's time to go away as this 'stable' release from ms-opentech appears to not be very stable at all.
What would people recommend ? Is there any production ready solution that exists, with windows stability in mind ?
It seems only client-side frameworks use this storage scheme for relations. DBs always use foreign keys (even such as rethinkDB). This always gets me a tiny mismatch between client and server.
It mentions ArangoDB twice. Once in the Document Store and once in the Graph sections.
1) Keeping all data in memory -- we didn't have enough RAM to keep everything in memory
2) Extremely slow startup time -- it was basically just reading the log of statements and loading things in memory
If those two are fixed now, Arango might be worth a serious look.
"A database in ArangoDB can be larger than the available main memory. Although ArangoDB stores all data on disk for durability reasons, it is a “mostly main memory database”. This means that the working set (the set of pages that are frequently accessed) should fit into main memory. It’s left to the operating system to determine the working set and to transfer pages between main memory and secondary storage. The data that are currently not needed are kept only on secondary storage.
"This is in principle also true for indexes: a part of the index (the working set) could be in main memory and the remainder could be on secondary storage if not used frequently."
You may also be interested in the roadmap: https://www.arangodb.com/roadmap/
Sigh. Good ol' mmap. They should ask MongoDB how they went for them. Then they should read Stonebraker's 1981 paper:
http://db.cs.berkeley.edu/cs286/papers/os-cacm1981.pdf
People never learn.
To press on about issue one, in a general sense, what I like about ArangoDB is the thoughts behind it and the continous execution of those thoughts. Is memory an issue? Possibly. Is the team aware of memory-mapped files' limitation? Yes. Will they improve in future release? Probably based on the excellent work they've produced in the past.
Their implementation of graphs is interesting to the general point. ArangoDB presently stores relationships in a collection. This won't scale. Neo4J's free book on graph databases explains why[1]. Fortunately, the ArangoDB team has a graph API (I'm in the process of implementing that now for Clojure [2]). This means, when they move to a more scalable architecture, I and anyone who use the API should be golden.
By the way, I'm a proponent of ArangoDB. Just in case that wasn't clear. I've got several projects for it.
[1] - http://graphdatabases.com/ [2] - https://github.com/deusdat/travesedo
SELECT * FROM scores
WHERE game_id = :game_id
ORDER BY score DESC