How Mailinator compresses its email stream by 90%
mailinator.blogspot.com
mailinator.blogspot.com
It's also nice not receiving ads in the mail every hour of every day just because I wanted to try some new YC-startup's product.
Then, keep a separate Chrome profile that's logged into the fake facebook, and your identity is safe. Probably best to lock down all privacy settings on the fake user to absolute maximum lockdown, to avoid information leaking onto Facebook's public layer that may tie you to it.
That was the strategy I used when I worked for a company that did business with Facebook games. Those things were so spammy, I would never want my actual friends to be spammed with all the viral content. To this day I keep my fake account for any and all untrusted FB-related things.
I recently wanted to create an "Admin" account for a client I was developing a web-app for as I didn't want to link my personal facebook account to my work. But with the phone number requirement the account got flagged and rejected from the system.
Frustrating!
[0] http://mailinator.blogspot.com/2008/01/your-own-private-mail...
A lot of certification authorities will allow you to get a certificate for example.com if you can prove that you can reply to an email sent to an email address @example.com. If you configure an MX record pointing to the mailinator MTAs, then anybody can request a cert and ask to validate domain ownership by contacting whatever@example.com, and can subsequently browse the whatever@ mailbox on mailinator to read and reply to the confirmation email.
A security researcher once obtained a certificate for Microsoft's live.com domain by registering an email account sslcertificates@live.com and using it to reply to the CA's verification email.
http://www.theregister.co.uk/2011/04/11/state_of_ssl_analysi...
Edit: yes, this works perfectly.
It would be nice if there was a quick-and-easy open-source version of an in-memory mail server. I'd like to run my own service, but not enough to write a robust SMTP server from scratch.
I haven't surveyed the field extensively but this has been my experience with the common recommendations. If anyone knows of a server that already exists and would make it trivial to set up basic SMTP/IMAP/POP3 with authentication let me know.
Ideally the only configuration info I would need would be something like this:
allow_imap = True
allow_pop3 = True
allow_smtp = True
accept_mail_for = ['this_domain.com', 'another_domain.com']
#users
users = ['jeff@this_domain.com', 'molly@another_domain.com']
#forwards
forwards = {'jeff@this_domain.com': 'cookiecaper@some_domain.com'}
and then, a program like mail_server_passwd would ask me for my username and let me set a password, and that would be that. It shouldn't be any difficult to bootstrap a mail server, and right now setting something that has POP3, IMAP, and SMTP up can take hours for experienced admins unless they know the mail server(s) very well ahead of time.
[Warnings: OpenSMTPD apparently hasn't eaten anyone's mail yet, but it's hardly done. By design, it doesn't include as many options as other mail servers, some of which may be useful. I don't think OpenSMTPD been ported to non-OpenBSD yet. calomel.org tends to be out of date, inaccurate or dangerously incomplete.]
Also, some SMTP servers are probably pretty easy to modify. I use qpsmtpd[1], which is probably quite easy to modify. Then, all you need is a IMPAD that can read from memory..
I would like to have my mail in something like Redis. It would give easy access from different languages etc..
I had originally created this site naively stashing the uncompressed source straight into the db. For the ~100,000 mails I'd typically retain this would take up anywhere from 800mb to slightly over a gig.
At a recent rails camp, I was in need of a mini project so decided that some sort of compression was in order. Not being quite so clever I just used the readily available Zlib library in ruby.
This took about 30 minutes to implement and a couple of hours to test and debug. An obvious bug (very large emails were causing me to exceed the BLOB size limit and truncating the compressed source) was the main problem there...
I didn't quite reach 90% compression, but my database is now typically around 200-350mb. So about 70-80% compression. So, I didn't reach 90% compression, but I did manage to implement it in about 6 lines of code =)
I made a submission here about it a few months back: http://news.ycombinator.org/item?id=3026892
In terms of mailinator being blocked . . . if that's the case, he's probably in the database http://www.block-disposable-email.com (as are my domains). The guy is pretty quick at finding at adding new disposable email domains.
I plan to add a few new domains on every now and then to see if I can at least keep some getting past his filter...
(biased creator here)
yopmail.com - any address (like Mailanator) and allows download of attachments directly
There are a ton of these available, just depends on your specific needs. I regularly use them for signing up to those sites that need sign-up before access to a file/article/image (e.g. forums).
I've also used them in testing web apps, using Selenium to grab the email address and registering a user and then going back to check the welcome email had arrived.
I'm currently floating in google near the top of page two =/
There are a few of entries on multi-threaded synchronous io vs asynch io; he's a big proponent of the former.
Thanks in advance!
CPU registers (8-32 registers) – immediate access (0-1 clock cycles) L1 CPU caches (32 KiB to 128 KiB) – fast access (3 clock cycles) L2 CPU caches (128 KiB to 12 MiB) – slightly slower access (10 clock cycles) Main physical memory (RAM) (256 MiB to 4 GiB) – slow access (100 clock cycles) Disk (file system) (1 GiB to 1 TiB) – very slow (10,000,000 clock cycles) Remote Memory (such as other computers or the Internet) (Practically unlimited) – speed varies
Hmm, there is something very wrong here. I'll try and explain in a blog post.
It probably doesn't matter though..
And yes, Redis is very fast and you gain a lot when using it compared to just a Hash in the same process.
The bigger win, IMHO, is that you gain flexibility: Since Redis (or whatever) is decoupled from you process it can run on another processor, another machine or perhaps run on many machines etc..
Not sure how in-process cache would work in node, being async and all, but yes in-process is faster. But then you have to think about stuff like:
- how do you avoid loosing everything when node crashes / restarts?
- what if another process needs to read write to the cache?
- what if you need more memory than a single machine provides (probably not going to happen).
- implementation bugs
In-process: Faster but probably harder to scale. But then again, it might be so fast you don't have to.I get the feeling you are kind of anti-Redis and I don't get why? Redis is a very cool project and could be useful for a lot of things.. It's not Redis' fault some people misuse it..
So why would you compress something that you can only decompress if its recently-reused?
How would you do mailinator with your strings in redis - and taking O(n) calls to redis to recover them to decompress an email where n is the number of lines (or consecutive lines, granted) in the email?
Yeah, you actually do.. sorry.
"So why would you compress something that you can only decompress if its recently-reused?"
Not sure I understand your questions, and I've just started looking at Redis. But I guess you could do it the same way, but the added latency may make it infeasible. But the better answer is probably that you don't: You would modify the implementation to fit Redis' (or whatever) strength and weaknesses.
You really can process mailinator- quantities of email with a simple Java server using a synchronized hash-map and linked list LRU and have some CPUs left over for CPU-intensive opportunistic LZMAing.
Trying to do it with IPC TCP ping-pong for each and every line though; well I'm not sure you could process mailinator quantities of email within any reasonable hardware budget...
Luckily you have a chance to see the error of your ways :)
But you have to remember that most people can't have important data in just one process; it's going to crash and your data is gone. The LMAX guys solved this in a cool way, but I wouldn't call it easy: http://martinfowler.com/articles/lmax.html#KeepingItAllInMem...
http://www.sorting-algorithms.com/ gives a decent overview with n equal to 20, 30, 40, or 50, but nothing smaller. Now I'm curious.
2. Has the author looked at Google Snappy? It does 500MB/sec. http://code.google.com/p/snappy/source/browse/trunk/format_d...
There is a pure-C implementation that might be easier to port: https://github.com/zeevt/csnappy
If mailinator wasn't already awesome, his writing about it sure is.
Was all of this worth it? It solved the problem of not burning through network and memory, but it was a local optima. The root problem was that this data came from another system which did not provide repeatable reads, and providing them would have been a massive effort. However, our users wanted to meander through a consistent data set over the course of an hour or so. To provide this ability to browse, we throw these records into a somewhat transient embedded H2 DB instance. The serialized format is required primarily to provide high availability via a clustered cache. In retrospect, I would have pushed for using a MongoDB-esque cluster which could have replaced both H2 (query-ability) and the need for the the serialized format (HA).
It surprised me that there were no open source projects (at least Java-friendly ones) which provided compression schemes taking advantage of the combination of well-defined record schema and redundant-in-practice data. Kyro (http://code.google.com/p/kryo) comes closest as a space-efficient serializer, but it treats each record individually. Protobufs, Thrift, Avro, etc. are designed for RPC/Messaging wire formats and, as an explicit design decision (at least in the protobufs case) optimize on speed and the size of an individual record vs. the size of many records. The binary standards for JSON and XML beat the hell out of their textual equivalents, but they don't have any tricks which optimize on patterns in the repeated record structures.
Is this just an odd use case? Does anyone else have a similar need?