373 karma · joined August 2, 2009
[ my public key: https://keybase.io/dialtone; my proof: https://keybase.io/dialtone/sigs/aGkL1Q2uSwF2ucv4-0mU7ecTvvlLiofZJ0oQBvFASOI ]
The batch usecase is, among the other things, attribution and it doesn't process 30x80B events per day, it looks at site activity and campaign activity over the past 90+ days and draws customer journey for each cookie. We obviously used to do things, when we were smaller, by keeping the data streamed in a database like hbase, but not only that was horribly slow and unreliable, it was also horribly expensive and doesn't really allow for re-applying a different attribution model retroactively and interactively like our product allows today (and it's incidentally now it's also much cheaper and faster than it used to be in the past when it was a streaming solution, orders of magnitude at that even).
In any case attribution is just one of the things we do in batch, and certainly one the bigger ones.
There are 150B+ auctions each day, of those we participate in at least 80B and in those 80B there are at least 5 separate predictions, to determine the type of auction (1st price v 2nd price for example), determine the price likely to win, determine the likelihood of the placement being viewable, determine the likelihood of the user to click, determine the likelihood of the user to convert given that they click, and then we run these last 2 for each candidate (campaign, creative) that is eligible for the current auction. We obviously don't analyse the stuff we didn't buy but 80B IS the number of top level ML-generated prices from our system.
The budgeting and targeting rules don't apply to the 80B number and they are slightly different systems.
Also recently Google has shifted the auctions to unified pricing which means 1st price basically and a single floor for them to make their prices more competitive in header bidding, so this might have had effects on Google’s position in the waterfall as well given what the publisher want to accomplish.
Generally speaking the article is overall fine, there are some historical inaccuracies of sorts and innovations aren’t explained well, for example Google didn’t just start to track people because they figured, but because behavioral targeting retained customers of all sized as opposed to just the bigger ones that do brand awareness campaigns, and that was due to measurable superior performance.
[1]: https://www.economist.com/leaders/2018/10/18/americas-shale-... [2]: https://www.economist.com/business/2018/10/20/the-shale-boom...
If you like developing open-source code, languages such as Python, Go, JS, C, D, Lua, Erlang, AWS, petabytes of data, and distributed low-latency systems, this may be your dream job.
This time we are particularly interested in finding full stack web developers with good JavaScript experience and experienced developers, tech leads and data scientists with great math knowledge and coding skills. This is a really unique opportunity to get to work with a massive scale (thousands of instances on AWS), low latency (real-time bidding with 100ms max latency and 70B requests daily, real-time machine learning with 1ms max latency), mission-critical systems (this is how we make money) and enjoy working on a strong frontend development team (http://tech.adroll.com/blog/frontend/2017/08/29/how-to-run-a...).
Learn more about us here http://tech.adroll.com/blog/
I am happy to tell you more over coffee in SF or by email, dialtone@adroll.com
This time we are particularly interested in finding data scientist, full stack web developers with good JavaScript experience and experienced Erlang developers / tech leads. This is a really unique opportunity to get to work with a massive scale (thousands of instances on AWS), low latency (real-time bidding with 100ms max latency and 70B requests daily, real-time machine learning with 1ms max latency), mission-critical systems (this is how we make money) and enjoy working on a strong frontend development team (http://tech.adroll.com/blog/frontend/2017/08/29/how-to-run-a...).
Learn more about us here http://tech.adroll.com/blog/
I am happy to tell you more over coffee in SF or by email, dialtone@adroll.com
If you like developing open-source code, languages such as Python, Go, JS, C, D, Lua, Erlang, AWS, petabytes of data, and distributed low-latency systems, this may be your dream job.
This time we are particularly interested in finding data scientist, full stack web developers with good JavaScript experience and experienced Erlang developers / tech leads. This is a really unique opportunity to get to work with a massive scale (thousands of instances on AWS), low latency (real-time bidding with 100ms max latency and 70B requests daily, real-time machine learning with 1ms max latency), mission-critical systems (this is how we make money) and enjoy working on a strong frontend development team (http://tech.adroll.com/blog/frontend/2017/08/29/how-to-run-a...).
Learn more about us here http://tech.adroll.com/blog/
I am happy to tell you more over coffee in SF or by email, dialtone@adroll.com
This is a library to read and write a data format that is optimized to give access to granular events and actors within an event stream. For example this could be used to trail all of the events generated by one entity (credit card, cookie, email, account and so on) over a dataset. At that point you can choose what you want to do with it: extract features for ML, train ML directly on raw data, run arbitrary queries for outliers and anomaly detection and what have you.
And continuous data is handled by sharding TrailDBs across some fields in our log lines, this way all the related events for a cookie in a given day belong in the same shard, each day the shard mapping is the same and we can just download the same shard ids from S3 and process the files sequentially using our DSL language. With a bit of code you can make this whole process of downloading from S3 and processing completely automated, this is in fact what we do with our data pipeline[0][1].
[0]: http://tech.adroll.com/blog/data/2015/09/22/data-pipelines-d... [1]: http://tech.adroll.com/blog/data/2015/10/15/luigi.html
TrailDB is more about a different way of grouping and events, and granularly querying and analysing each trail of data. TrailDBs are materialized and stored in S3 typically.
Most advanced buyers with actual ML in their buying algorithms do this. But ML works in statistical averages on the behavior seen across all users visiting a particular site. At that point the buying process works by figuring out the expected value of a new impression and bids that value, the expected value depends on how the advertiser values clicks or conversions or impressions, so as long as the cost of the impression is lower than the marginal value an algorithm will continue to bid, and potentially win, because it's worth it.
And you can do all the A/B tests you want and you'll see that this is actually true, capping frequency. or choosing to not show an ad because the position on the site is not great, is not a good idea, the right process is to determine a price that, all things considered, is the maximum price (proxy for value and risk) you are willing to pay to be shown in that bad slot that adds marginal value for the advertiser, and marginal value is measured however the advertiser wants.
At the same time the data that you voluntarily give to Facebook, while certainly useful for look-a-like campaigns and such, is not exactly that predictive when talking about DR campaigns for which you need to establish an imminent intent of buying something. Using their various tracking pixels, that ad networks cookie match with, they start to collect also that dataset. At the moment they also have 3 different ways to cookie match with partners, all 3 basically required, this also contributes to added latency when loading pages, would be better to have 1.
The ad tech market is fairly complex and unfortunately often people make assumptions about datasets, their predictive power or reasons behind why the situation is the way it is without necessarily having done the research about it.
[1]: http://tech.adroll.com/blog/data/2015/06/26/kinesis.html