HNHacker News
TopNewBestAskShowJobs

dialtone

373 karma · joined August 2, 2009

CTO @ AdRoll

[ my public key: https://keybase.io/dialtone; my proof: https://keybase.io/dialtone/sigs/aGkL1Q2uSwF2ucv4-0mU7ecTvvlLiofZJ0oQBvFASOI ]

submissionscomments
dialtone··on New York Times phasing out all 3rd-party advertising data
This is just 3rd party data, not 3rd party ads... Nothing to see.
dialtone··on How NextRoll uses AWS Batch for daily business operations
The ML side is split from batch, this is part of our ML infrastructure http://tech.nextroll.com/blog/data-science/2018/04/26/just-b... .

The batch usecase is, among the other things, attribution and it doesn't process 30x80B events per day, it looks at site activity and campaign activity over the past 90+ days and draws customer journey for each cookie. We obviously used to do things, when we were smaller, by keeping the data streamed in a database like hbase, but not only that was horribly slow and unreliable, it was also horribly expensive and doesn't really allow for re-applying a different attribution model retroactively and interactively like our product allows today (and it's incidentally now it's also much cheaper and faster than it used to be in the past when it was a streaming solution, orders of magnitude at that even).

In any case attribution is just one of the things we do in batch, and certainly one the bigger ones.

dialtone··on How NextRoll uses AWS Batch for daily business operations
It's a bit more complicated than that given that attribution looks at the past 30+ days of data (120 in NextRoll's case) to determine if any marketing activity happened and allows the customer to adjust the attribution window to whatever they want.

There are 150B+ auctions each day, of those we participate in at least 80B and in those 80B there are at least 5 separate predictions, to determine the type of auction (1st price v 2nd price for example), determine the price likely to win, determine the likelihood of the placement being viewable, determine the likelihood of the user to click, determine the likelihood of the user to convert given that they click, and then we run these last 2 for each candidate (campaign, creative) that is eligible for the current auction. We obviously don't analyse the stuff we didn't buy but 80B IS the number of top level ML-generated prices from our system.

The budgeting and targeting rules don't apply to the 80B number and they are slightly different systems.

dialtone··on France Tried Soaking the Rich. It Didn’t Go Well
Already the case. People with assets beyond $2M will be taxed a capital gain tax (23.8%) from all of their assets as if they were sold at fair market value.
dialtone··on France Tried Soaking the Rich. It Didn’t Go Well
renouncing citizenship comes with an exit tax of 23.8% of your assets. https://1040abroad.com/faq/renouncing-u-s-citizenship/
dialtone··on How internet ads work
It's nothing too crazy, simply that for brand awareness campaigns to work, or even collect the market research data that you are talking about you need to spend significant money for a decent period of time. Something that smaller companies tend to not have while they are focusing on starting predictable growth and product market fit more than scale. Smaller companies will invest in marketing that has some guarantees of performance like retargeting, email or search rather than brand.
dialtone··on How internet ads work
That’s an average price and mostly driven down by lack of data on users browsing the site and the mix of campaigns that bought, given retargeting campaigns can easily pay averages in the $5-10 CPM ranges depending on the campaign and users and data density. So the bad persistence of cookies is what largely drives this price to be low. Obviously anyone can choose how to take that.
dialtone··on How internet ads work
The publisher chooses the order in the waterfall for header bidding. Due to its nature Header Bidding is serial in the page and only a limited number of partners can bid on the page (4-5) and are called sequentially. Header bidding was created to increase competition for premium inventory beyond Google, so naturally Google ends up being used as remnant inventory in these cases because it’s always able to fill inventory, but not necessarily at the highest price and it’s thus placed last often but it really depends.

Also recently Google has shifted the auctions to unified pricing which means 1st price basically and a single floor for them to make their prices more competitive in header bidding, so this might have had effects on Google’s position in the waterfall as well given what the publisher want to accomplish.

Generally speaking the article is overall fine, there are some historical inaccuracies of sorts and innovations aren’t explained well, for example Google didn’t just start to track people because they figured, but because behavioral targeting retained customers of all sized as opposed to just the bigger ones that do brand awareness campaigns, and that was due to measurable superior performance.

dialtone··on Deconstructing Google’s excuses on tracking protection
You are aware that Google has vowed to actively fight any sort of fingerprinting right? https://techcrunch.com/2019/05/07/googles-chrome-will-soon-g...
dialtone··on People care more about privacy than they think
respect for others that don't want to see me half naked while on the toilet? I don't lock the bathroom door at home.
dialtone··on Silicon Valley is awash with Saudi Arabian money
USA still imports about 50% of the oil barrels it uses each day, about 10m[1], and there is not enough pipeline capacity to move shale oil to refineries[2]. It's a long way to being self-sufficient.

[1]: https://www.economist.com/leaders/2018/10/18/americas-shale-... [2]: https://www.economist.com/business/2018/10/20/the-shale-boom...

dialtone··on iOS 12 released
The fact that each device has an IDFA that any app can share with literally anyone in the world, without the need to tell you, because IDFA is a quasi-persistent identifier I suppose doesn't matter.
dialtone··on Ask HN: Who is hiring? (February 2018)
AdRoll | San Francisco | On-site/remote | Full-time

If you like developing open-source code, languages such as Python, Go, JS, C, D, Lua, Erlang, AWS, petabytes of data, and distributed low-latency systems, this may be your dream job.

This time we are particularly interested in finding full stack web developers with good JavaScript experience and experienced developers, tech leads and data scientists with great math knowledge and coding skills. This is a really unique opportunity to get to work with a massive scale (thousands of instances on AWS), low latency (real-time bidding with 100ms max latency and 70B requests daily, real-time machine learning with 1ms max latency), mission-critical systems (this is how we make money) and enjoy working on a strong frontend development team (http://tech.adroll.com/blog/frontend/2017/08/29/how-to-run-a...).

Learn more about us here http://tech.adroll.com/blog/

I am happy to tell you more over coffee in SF or by email, dialtone@adroll.com

dialtone··on Tether has issued $450M USDT in past 4 days
Tethers today traded $3.8B, $1.6B is their market cap. https://coinmarketcap.com/currencies/tether/ so it's even more unreal.
dialtone··on Ask HN: Who is hiring? (November 2017)
If you like developing open-source code, languages such as Python, Go, JS, C, D, Lua, Erlang, AWS, petabytes of data, and distributed low-latency systems, this may be your dream job.

This time we are particularly interested in finding data scientist, full stack web developers with good JavaScript experience and experienced Erlang developers / tech leads. This is a really unique opportunity to get to work with a massive scale (thousands of instances on AWS), low latency (real-time bidding with 100ms max latency and 70B requests daily, real-time machine learning with 1ms max latency), mission-critical systems (this is how we make money) and enjoy working on a strong frontend development team (http://tech.adroll.com/blog/frontend/2017/08/29/how-to-run-a...).

Learn more about us here http://tech.adroll.com/blog/

I am happy to tell you more over coffee in SF or by email, dialtone@adroll.com

dialtone··on Ask HN: Who is hiring? (September 2017)
AdRoll | San Francisco | On-site/remote | Full-time

If you like developing open-source code, languages such as Python, Go, JS, C, D, Lua, Erlang, AWS, petabytes of data, and distributed low-latency systems, this may be your dream job.

This time we are particularly interested in finding data scientist, full stack web developers with good JavaScript experience and experienced Erlang developers / tech leads. This is a really unique opportunity to get to work with a massive scale (thousands of instances on AWS), low latency (real-time bidding with 100ms max latency and 70B requests daily, real-time machine learning with 1ms max latency), mission-critical systems (this is how we make money) and enjoy working on a strong frontend development team (http://tech.adroll.com/blog/frontend/2017/08/29/how-to-run-a...).

Learn more about us here http://tech.adroll.com/blog/

I am happy to tell you more over coffee in SF or by email, dialtone@adroll.com

dialtone··on TrailDB – An Efficient Library for Storing and Processing Event Data
That's just what the example script does. In the database all the data is stored granularly so you can, and we (AdRoll) do, use TrailDB to examine individual trails one by one after you select the subset that you are interested in.
dialtone··on TrailDB – An Efficient Library for Storing and Processing Event Data
Parquet is a columnar format, so it's optimized for aggregations over columns filtered by a set of dimensions. TrailDB data format is optimized for discrete event analysis without aggregation: data points are granularly grouped by their source (for example a user account and actions on a website) and your queries operate on each of these groups independently. No aggregation happens in TrailDB.
dialtone··on TrailDB – An Efficient Library for Storing and Processing Event Data
This is a tool to look at each user individually actually, not aggregated.
dialtone··on TrailDB – An Efficient Library for Storing and Processing Event Data
AFAIK Druid is for time series and it's a columnar database. Their format has dimensions and then pre-aggregated fields on those dimensions. Afterwards you can run SQL queries on that data format to get full aggregations in return.

This is a library to read and write a data format that is optimized to give access to granular events and actors within an event stream. For example this could be used to trail all of the events generated by one entity (credit card, cookie, email, account and so on) over a dataset. At that point you can choose what you want to do with it: extract features for ML, train ML directly on raw data, run arbitrary queries for outliers and anomaly detection and what have you.

dialtone··on TrailDB – An Efficient Library for Storing and Processing Event Data
We have an internal framework that compiles a state machine DSL into lower level code to execute queries on TrailDBs. However the binding in Python is fairly thin, and the same can be said for the other ones like the golang one. So before optimizing with a lower level language or a special DSL I would verify that Python or any other of the supported languages doesn't satisfy your requirements.

And continuous data is handled by sharding TrailDBs across some fields in our log lines, this way all the related events for a cookie in a given day belong in the same shard, each day the shard mapping is the same and we can just download the same shard ids from S3 and process the files sequentially using our DSL language. With a bit of code you can make this whole process of downloading from S3 and processing completely automated, this is in fact what we do with our data pipeline[0][1].

[0]: http://tech.adroll.com/blog/data/2015/09/22/data-pipelines-d... [1]: http://tech.adroll.com/blog/data/2015/10/15/luigi.html

dialtone··on TrailDB – An Efficient Library for Storing and Processing Event Data
It's not quite the same thing, PipelineDB is a database to execute streaming queries at massive volumes. That is definitely one of the challenges at AdRoll.

TrailDB is more about a different way of grouping and events, and granularly querying and analysing each trail of data. TrailDBs are materialized and stored in S3 typically.

dialtone··on TrailDB – An Efficient Library for Storing and Processing Event Data
The approach here is very different. It's not that window functions are bad, but that the traditional SQL database, or even column oriented database, doesn't have a disk format that is versed to this type of analysis for the high volumes that AdRoll has. TrailDB can filter through 10s of millions of events per second on a Macbook, just storing billions of events in a more traditional DB or warehouse would use far more than the space that TrailDBs use. On top of that it's relatively easy to build visualisations on top of these databases with an interactive UI to do exploration. So the window function being complicated was perhaps just the spark but there are real advantages in speed, efficiency and expressiveness in this type of database.
dialtone··on Google: End of the Online Advertising Bubble
What Google knows about you, or Facebook for the matter or any other buyer, can be used for targeting and as user features in the ML model which is used to determine the price the buyer is willing to pay for a given impression in the auction.

Most advanced buyers with actual ML in their buying algorithms do this. But ML works in statistical averages on the behavior seen across all users visiting a particular site. At that point the buying process works by figuring out the expected value of a new impression and bids that value, the expected value depends on how the advertiser values clicks or conversions or impressions, so as long as the cost of the impression is lower than the marginal value an algorithm will continue to bid, and potentially win, because it's worth it.

And you can do all the A/B tests you want and you'll see that this is actually true, capping frequency. or choosing to not show an ad because the position on the site is not great, is not a good idea, the right process is to determine a price that, all things considered, is the maximum price (proxy for value and risk) you are willing to pay to be shown in that bad slot that adds marginal value for the advertiser, and marginal value is measured however the advertiser wants.

dialtone··on San Francisco Bubble
At 3.92% interest rate for 30-years fixed with a 20% downpayment that's about $3,400/mo and at $900k property tax is another $6k/year, insurance is less than $100/mo so you net out at less than $4k/mo in PITI. And maintenance you'd hope to check what's wrong with the house before you buy but a new roof costs about $20k so unless you replace a roof a year I doubt you'll be spending over $1000/mo for maintenance, at least I'm not. And 20% is not a sizeable downpayment but the minimum to not have to pay mortgage insurance, a sizeable downpayment is 30%, which in a 30 years fixed means paying less than $3500/mo PITI. Of course if you go with ARMs you'd end up paying even less since interest rates are even lower there. If you can afford the downpayment you'd be paying less than rent in SF most of the time.
dialtone··on Facebook Ads Are All-Knowing, Unblockable, and in Everyone’s Phone
That FB can do just fine without tracking users outside of FB is true but debatable given how widely adopted is their WCA product, and the fact that they are looking into using their like buttons to track people's activity online to be used for their advertising machine learning.

At the same time the data that you voluntarily give to Facebook, while certainly useful for look-a-like campaigns and such, is not exactly that predictive when talking about DR campaigns for which you need to establish an imminent intent of buying something. Using their various tracking pixels, that ad networks cookie match with, they start to collect also that dataset. At the moment they also have 3 different ways to cookie match with partners, all 3 basically required, this also contributes to added latency when loading pages, would be better to have 1.

The ad tech market is fairly complex and unfortunately often people make assumptions about datasets, their predictive power or reasons behind why the situation is the way it is without necessarily having done the research about it.

dialtone··on AWS US East is experiencing high error rates on several services
You can relatively trivially build multi-master cross-region replication in DynamoDB by using kinesis and writing to kinesis instead of DynamoDB directly. On the consuming end of Kinesis you then fan out to all the DynamoDB (or whatever other database you want to use) regions[1]. Admittedly this only works with some relatively relaxed constraints on the latency you can see, intra region latency can go up to 1 second although rarely, while cross-region is around 3 seconds. An important role is also played by the structure of your objects and how accepting they are of concurrent updates coming from different regions (which is the main reason why the default replication in DynamoDB is not multi-master).

[1]: http://tech.adroll.com/blog/data/2015/06/26/kinesis.html

dialtone··on AWS US East is experiencing high error rates on several services
This really is the first major failure of DynamoDB since when I started to use it, when it was in beta over 4 years ago. Not sure many companies would be able to run a multi-million QPS scale database with such a great availability record. In any case we just moved traffic away from us-east-1 and into us-west-2 and all is fine still so you should design at least your critical systems for regional failures.
dialtone··on Why my car cost more than taking Uber everywhere
900 miles per year... You definitely don't need a car at 900 miles per year. Get zipcar or rent a car when you occasionally need it.
dialtone··on Google Shows How To Scale Apps From Zero To One Million RPS, For $10
This is not the point of the test, the test is about showing you that the load balancer in GCE can handle that many requests per second and with a single IP address. Whatever the machines are doing behind doesn't matter since the load balancer job is to handle a ton of traffic. This is practically the only case in which responding with 1 byte makes sense in the test.
← PreviousPage 3 of 6Next →