Dragon: A distributed graph query engine
code.facebook.com
code.facebook.com
The technology of Facebook and Google (where I worked as a contractor) impresses! That said, I like a simpler Internet with smaller scale services like Gnu Social, and simply using email to stay in touch with family and friends. I have this preference both for privacy reasons and also prefer more one on one communication.
(filter (> age 20))
That second form to filter should be a function, but in this case it is an expression that returns a boolean and is not itself a function.
I'm pretty sure this is Clojure or something very close to it. ->> is the thread-last macro. What it does is take the first expression, ($alice) and inserts it as the last expr in the next form so that the second step (assoc $friends) would eval to (assoc $friends $alice). It continues this by inserting this expr into the last part of the next form. So the example would reduce to: (filter (> age 20) (assoc $friends (assoc $friends $alice)))
i.e. while walking, we notice a function (> age 20) and age isn't visible in scope. We thus rewrite that into something like (fn [age] (> age 20)) or (fn [x] (> (:age x) 20)).
I've taken this approach before with DSLs. It makes for extremely fast to write business logic where business users don't have to care as much about where their data comes from (they just need field names).
(->> ($alice) (assoc $friends) (assoc $friends) (filter (> age 20)) (count))
...in Gremlin is: g.V(alice).out("friends").out("friends").has("age",gt(20)).count()S-expressions: fully parenthesized prefix notation.
For distributed relational or graph databases, seems like the key trick for making queries efficient is to get related data on the same host, whenever possible.
So would be cool to dig into the specifics of the algorithms they are using, to see exactly how they are optimizing where to store the data. With a graph database, it's impossible to guarantee having all related data together (eventually, friend of a friend of a friend...will be on a different host). So needs to be heuristics based.
For tree shaped data, on the other hand, it is possible to have the root of each tree and all of its related data on the same host (assuming each tree is "reasonable" size). Google's F1 project took this approach.
https://www.usenix.org/conference/nsdi16/technical-sessions/...
- LinkedIn has been terrible for about 2-3 years since things stopped updating real time and the timeline started getting random
- Facebook randomly suggesting things from your graph based on what it thinks you like
- Twitter I heard soon will be no longer in chronological order?!?
I think there is something to be said for low complexity MySQL and memcached.
Saying that Facebook is recommending things simultaneously at random, based on a reccomendation algorithm, and also as an optimisation which wouldn't be needed if they used MySQL, is obviously nonsense.
I'm not saying MySQL > some specialized solution for scale... I'm guess I'm drawing the correlation that when companies have to make investments like this they are also trading off experience, or maybe to your point screwing up the experience using recommenders and algorithms instead of letting my feed flow.
"Crap how do we make money"
"Crap how do we retain users/keep them here longer"
The stack is to do with providing a fast performance to vast numbers + making it easier to work on for engineers than it might be at that complexity.
But there are definitely stack complexity trade offs when you get to the scale of these guys.
Also could be more correlation not causation that when you get to a certain scale and more stack complexity you also start hiring specialized staff that end up diluting the culture and spirit of a startup.