Communication Logs: Handling 250M SMS messages a day
infobip.com
infobip.com
CREATE TABLE friendship (
person1 INT NOT NULL,
person2 INT NOT NULL,
PRIMARY KEY (person1, person2),
CHECK (person1 < person2)
);
CREATE INDEX friendship_person2 ON friendship(person2);
CREATE VIEW friendship_view AS
SELECT person1 AS person, person2 AS friend
FROM friendship
UNION
SELECT person2 AS person, person1 AS friend FORM friendship;
taken from here: http://www.postgresql.org/message-id/20141111201127.d80b6bc4...What do you folks think? It seems like this would be the densest way to store the relationship, using the checked constraint trick with the IDs so you don't need duplicate records to store Alice and Bob's friendship as {A -> B, B -> A} if the friendship is inherently bidirectional (vs a followed/follower model), and doesn't preclude one from storing the data in another isomorphic structure that might be optimized for specific queries. I'm sure it's possible to replicate this with Neo4j, but it felt cumbersome programming that in Java versus expressing that in SQL. But I'd really like to use Neo4j in an OLTP environment, but still haven't had enough of an impetus yet because WITH RECURSIVE in postgres works well if the recursion depth is capped.
By the way, if you're interested in graph representations for efficient querying, check out TripleStores [1], and the Wiki list on subject-predicate-object databases [2].
[1] https://en.wikipedia.org/wiki/Triplestore [2] https://en.wikipedia.org/wiki/List_of_subject-predicate-obje...
https://github.com/neo4j/neo4j/tree/3.0/community/lucene-ind...
ps:
http://stackoverflow.com/questions/9541541/b-tree-vs-bitmap-...
https://en.wikipedia.org/wiki/Bitmap_index
Thanks for showing me triplestores! Good to know that it's better to use a database engine optimized for triples for graphs vs rolling your own in SQL.
I certainly hope your system can manage 400kBps throughput.
As a wild guess, you're off by 1000x. You may say 400MBps is still not too crazy, but it's beyond the real capacity of many systems out there.
Even if peak "instantaneous" flow is 400MBps, considering that SMS already has seconds to minutes of latency, a bit of buffering shouldn't be a big deal.
They start the article by saying they make six (not necessarily serial) network roundtrips for each message which indicates that something is horribly wrong with the design of their system. Their code needs to bill each SMS to the sender, send the SMS to the telco, and log successes and failures in a way that's searchable later. The rest of the article describes their terribly overwrought approach to these tasks. With so much enterprise software and so little actual discussion of the problem domain, there's no way they're butting up against anything interesting enough to merit a HN submission.