The ML side may well benefit from batch but I bet splitting out the attribution component would allow them to be more flexible and cost efficient in their approach there.
The ML side may well benefit from batch but I bet splitting out the attribution component would allow them to be more flexible and cost efficient in their approach there.
The batch usecase is, among the other things, attribution and it doesn't process 30x80B events per day, it looks at site activity and campaign activity over the past 90+ days and draws customer journey for each cookie. We obviously used to do things, when we were smaller, by keeping the data streamed in a database like hbase, but not only that was horribly slow and unreliable, it was also horribly expensive and doesn't really allow for re-applying a different attribution model retroactively and interactively like our product allows today (and it's incidentally now it's also much cheaper and faster than it used to be in the past when it was a streaming solution, orders of magnitude at that even).
In any case attribution is just one of the things we do in batch, and certainly one the bigger ones.
Doing 80B predictions a day for 500,000 conversions would make this the lowest efficiency ad platform I’ve ever seen by several orders of magnitude.
The actual attribution shouldn't be on 80B though. It's pointless analysing requests you didn't buy. There might be some value in using that data in ML to feed into a pricing algorithm, but it would be marginal and I doubt cost/benefit would ever stack up. it would technically be in breach of pretty much every SSP contract I've seen (although everyone does it).
There are 150B+ auctions each day, of those we participate in at least 80B and in those 80B there are at least 5 separate predictions, to determine the type of auction (1st price v 2nd price for example), determine the price likely to win, determine the likelihood of the placement being viewable, determine the likelihood of the user to click, determine the likelihood of the user to convert given that they click, and then we run these last 2 for each candidate (campaign, creative) that is eligible for the current auction. We obviously don't analyse the stuff we didn't buy but 80B IS the number of top level ML-generated prices from our system.
The budgeting and targeting rules don't apply to the 80B number and they are slightly different systems.