Real time bidding in Erlang (2012)
ferd.ca
ferd.ca
On the implementation side, like Python it's weaker with CPU heavy tasks, but a normal program maintains low latency, unbeatable uptime, hot code swaps. For CPU heavy work, I presume most Erlangers do what we all do- C (or maybe Rust) libraries. On the language side, it's terse like Python and a small language.
I fail to see why this shouldn't be the server language of choice. Other than tougher recruiting. Learning about FP would be interesting too.
I've dabbled in Go and Node.js, and keep up on developments with Rust, but keep coming back to Erlang as what I'd like to deep dive on next.
Anyway, we have enough non-core stuff (the mostly static website, resource translations etc) floating around that's not in Erlang that having people around who used PHP or Python in a former life is a good thing. :)
We still use it for our new product (marketing tool for the music industry) but while it certainly works for us, we no longer have to support tons of simultaneous connections to our servers. The backend tech could almost be anything.
OpenX is a big user of Erlang for RTB.
Actually, I hold a position that in some sense it makes recruiting easier. It's very hard to find idiots that can be taught to write in Erlang (or Haskell, OCaml, or Clojure), so SNR should be very good.
There is also a case when candidate already knows a functional language, just not Erlang (or whatever is required). Then one might assume the candidate can learn one more functional language.
EDIT: just found out some typos
Overall it offers a higher level entrance to the Erlang world and coming from Python you may find the syntax slightly easier to digest (anecdotal).
Erlang sounds nice, but I don't like platforms that are too specialized for certain use-cases. I also never understood the "C libraries" argument - this is what made Python non-portable, this is what will kill it.
After coming to Google and working on the other side of the RTB process for a while.. well, I just wouldn't do it in anything other than something like C++ or Rust -- a systems level language where I can fine tune everything and not be interrupted by a GC. I just wouldn't mess around -- not only is it important to have low latencies, it's important to have _consistently_ low latencies.
I'd point out that people can and do use GC'd languages in environments that are dramatically more latency sensitive than RTB. For that matter, in latency sensitive applications allocation/deallocation after startup is a no-no, so GC tends to not be the major reason not to use the JVM (memory layout/unfettered access to system calls/etc tend to drive that decision).
95% of the time the latency is acceptable. The problem is that under heavy throughput the latency spikes periodically. I wish I still had some of my old graphs to show.
I just wouldn't do it again, the time spent futzing trying to tune the JVM to avoid this would have been better spent on writing code in a systems level language where allocation is predictable.
I was a big booster of the JVM as a mature platform for a lot of things. But for this purpose... nope.
My point is that most truly latency sensitive applications, in GC'd languages or not, don't do any allocation/deallocation on the hot path. Memory management is just too slow.
So GC twiddling is a red herring in those environments, though using manual memory techniques on the JVM may obviate the major reasons to use it in your particular cases. Lots of teams make the engineering decision the other way and use the JVM in usages that require more consistent latencies than RTB requires.
After working on the bidder side of RTB, I then had two jobs where I worked on the other side, _sending_ bid requests, one of which was here at Google. I learned a lot.
CPU-intensive operations pose their own problems. EC2 doesn't have amazing context switching so sometimes Akka's internal dispatchers won't be able to get CPU time quickly enough to keep the cluster heartbeats going, also causing your node to be marked unreachable by the rest of the cluster. I imagine bare-metal would be far easier to work with in this regard.
There are certainly ways to deal with these problems by tuning the GC and Akka failure detectors, but it's a serious problem with Akka.
[1] http://www.erlang.org/doc/man/erl_nif.html (see "Warning")
[2] http://jlouisramblings.blogspot.com/2013/07/problematic-trai...
[3] http://erlang.org/pipermail/erlang-patches/2011-April/002046...
NIF's are one of several ways of linking C and Erlang, and the one most similar to how Tcl/Python/Ruby/etc... link in C code. In those runtimes, if your C code segfaults, you will also crash the interpreter. You're correct that they're a bit trickier to integrate because you have to consider the scheduler, whereas Tcl/Python/Ruby don't care if you call out into the C code and it takes a minute to complete.
Erlang offers you other means of working with C code, like ports, and C nodes, which can comunicate at a distance, and run in separate Unix processes. This is a compromise in that comunication is more expensive, but the system as a whole is more robust.
Certainly take the time to learn Erlang, but when you get the chance it's well worth checking out a modern typed functional language - F# or Scala or Haskell.
RTB is an amazing challenge and way too much fun for those inclined to bang their heads against all sorts of problems you never run across in the "regular" world.
I've done some work on these systems, one of my favorite challenges is writing tools that segment impressions in real time, you have only 10ms to respond because you're augmenting the bid before others can bid. A RTB bidder typically has about 100ms to respond. When you're shooting for 10ms at 2.5m qps peaks you have to think about everything from kernel versions, network settings, pinning to processors, avoiding GC, logging to disk is too slow etc.
As a complete tangent, as someone who has worked in environments that had at least an order of magnitude tighter latency requirements than the RTB world, and who sees this misconception a lot, logging to disk on a modern OS is not "too slow". There may be very valid reasons not to log to disk, most obviously operational concerns, but latency/throughput aren't one of them.
It wasn't rtb (just normal adserver).
Wonder what your comments would be?
Maxmind - good. Not like there's many choices here, though. :)
Instead of "waking up to sync logs", consider using something like NSQ to emit events as they happen. You can scale the number of servers/processes generating messages and the number of workers consuming those messages (and committing them to your database) very easily.
You could also replace the writing of the transaction log with a NSQ event. It lets you avoid having to write and scale the log shipping stuff.
We precalculated which ads a given user was eligible for and a separate process was contacted when a bid request came in to get the info for the ad to show. We never had to do anything funky to do geotargeting exclusion at scale.
Instead of having your adserver connect to a database, have a separate process generate a working set (as JSON or whatever you fancy), compress it, and ship it to the adserver periodically. The adserver can just do a straight load from the file every minute or whatever interval you'd like. If the file's mtime is too old, raise an alert and stop serving ads if necessary. Keeping things separate and simple lets you scale more simply. Our working sets were on average about 2gb uncompressed and they could be loaded in a few seconds (C++/JSON and later Go + JSON).
Seems like it was a fun project and I hope you learned a lot!
I also suspect he refers to 300k distributed across multiple datacenters.
We didn't use anycast, though. I suggested it to our CTO many times and it would have saved us over $5,000/mo in DNS costs, but it never got done.