From what I saw, the data privacy rules and client contracts are followed scrupulously, and infractions are noted and remediated. Encryption slowed us down at least 10X,and bureaucracy another 10X.
While frustrating at times, I appreciated that most if not all actually wanted to stay well within the legal guardrails.
Of course, this is one person's experience, and obviously can't apply to every company. Genie and cork, toothpaste, and all that.
What many companies do is they sprinkle some magic "anonymous" pixie dust on their data and then tons of laws no longer apply, even if the data is trivially identifiable stuff. And location data is often of that nature.
And engage with third party matching services so that personally identifying information is salted and hashed until the print shop
Etc.
I've never had anything to do with ad tech, but the (few) certification processes I've been involved with were largely documentation and "yes, we do have backups. no, unauthorized persons cannot access our servers" promises without anybody actually auditing/testing.
Do they have independent auditors regularly look at the tech and operation to verify that they actually do conform with the laws and don't just claim to?
Further, any time there is a broad news story (e.g. Wells Fargo's bogus accounts) or legal item in pipeline with broad impact, you can be assured every company looks to make sure they aren't in bad shape.
- Companies collect a number of "records" keyed to an email address. How they do this varies from running microsites/forms that collect data directly from people, to scraping linkedIn and resumes.
- Data vendors will share this data with each other by hashing the email address. Almost always with md5, and rarely salted.
This means you can enrich data both anonymously (by relying on the fact the email address is unsalted) or in an identifying way (because one of the parties has the real email address). In both cases, it's just a LEFT JOIN.
I've been approached by a company that trades e-commerce transaction data, for personal data tied to IPs.
As in, you give me your customer data, and I'll tell you who visited your site based solely on the IP address.
We declined, but it's tempting.
They'll buy the md5-email-to-cookies from one provider (e.g. Lotame, Liveramp, etc) then use that to onboard email+contact data they've purchased from companies that have email address-to-personal data (e.g. MVF, ZoomInfo, etc).
IF that was done as purely a lead qualification step, it's a good way to take legitimate content syndication done by a third party into direct marketing, but there's no technical reason they need to be leads or have any level of qualification -- and only a weak market force (poor conversion rate) that prevents it from being more widespread.
I still want to try dropping a letter in a mailbox from a different city with just our postal code written on it and see if it arrives.
Its because nobody cares until Congress starts hauling in software developers and "data scientists" from tech startups, everyone is content watching Zuckerberg be the face of society's discontent.
As long as they keep asking the wrong people the wrong questions there is no need for everyone else to acknowledge their role in the problem.
Advertising is applied sociology. As such, advertisers want to aggregate large data sets into large segments that are easy to manipulate statistically. (Where the central limit theorem starts working.)
There is no demand for personal data or de-anonymization because that stuff doesn't sell.
The personal data collection is done by Google, Facebook at al not for advertising purposes. They're collecting it because they view it as a resource and a currency in the future de-anonymized world. (Think China's "social capital" except on a larger scale.)
Source: I've worked in the ad industry for over 15 years.
Given your 15 years of experience, what resources do you recommend to HN so that we can learn more? Can you give us a "life of an advertising bit," eg: A person visits a website on their phone, that information is accompanied by x data on their phone, goes to the initial ad server, this information is compiled against data from sources a,b,c, etc ...
Especially because the web is full of Privacy notices people agree to, and I guess in some of those people actually agree to have their anonymous browsing data connected to their real identities.
Like I said, knowing real identities is the last thing on the list of ad tech priorities.
If shadowy entities are collecting "real identities" then it's not for ad purposes.
Say what?? I’ve also worked in the ad industry and deanonymized personal data is shared and sold routinely. You speak of statistics and large segments but every advertiser I’ve interacted with is either doing individual-level targeting or striving towards it.
Only if they're clueless.
For example: Nike really wants a dataset of "people who buy expensive sneakers for fashion purposes".
This dataset is probably hundreds of millions of anonymous people, and not personal data. If there was a way to get this dataset directly, Nike would do that in a heartbeat.
Unfortunately, as of 2019 the only way to get something like this today is by, e.g., crossreferencing credit card purchase info with Twitter browsing logs, which leaks a shitload of sensitive private data.
For ad purposes personal data collection is a bug, not a feature.
I’m not necessarily talking about demographics, but rather clickstream data, and anything categorical that you can get your hands on. You join that to your CRM and build a model to predict buyers. A really good predictive dataset for marketing purposes is simply a list of time stamps and names of visited domains. With the right feature engineering, that becomes an excellent proxy for demographic data, current buying appetite, and a whole lot more.
At the end of the day, you don’t even necessarily need to know what the data means as long as it’s predictive. And there are plenty of brokers out there who will let you test their data for free with an agreement to pay if you end up using it at scale. All of this revolves on using PII for matching.
I’m sure what you’re saying is true for some marketers, but there are billions being made on PII keyed data.
Let me repeat again. PII is a crutch used for matching, because current matching/segmenting technologies are crude.
Advertisers don't want PII. What they want is target audiences with predictive power, which means data sets where the central limit theorem holds sway. (I.e., thousands and millions of people lumped together.)
If advertisers could get at these segments directly without PII, they'd do it in a second.