You’re not anonymous. I know your name, email, and company.
42floors.com
42floors.com
https://panopticlick.eff.org/ <-- check how unique your browser is.
Instead of a script to embed, these firms could provide an API to identify users from the server side. The scripts that captures the profile would be served by the sites themselves rather than from third party services.
Toast.
A possible solution would be anonymize the browser fingerprint, at least in private mode, ie lie about the details of the system.
Google, Mozilla, Opera, can you hear me?
--
Maybe it's forgotten. Maybe it lies. Maybe every time I rev Firefox Nightly I change identity.
What is true is that every time I leave an email address, it's tagged with the name of the site where I left it.
But note that a constantly changing fingerprint doesn't make it useless for tracking - if a site can keep any kind of cookie to persist between browser updates, it could add the updated signature. Then when you purge cookies and persistent storage, a site can re-add the cookie to keep on identifying you if your signature hasn't changed.
You'd have to purge all persistant storage at the same time as an update to avoid this, and even then (or if you never had persistent data to begin with) your IP or even geographical location will likely be enough to identify you again.
I use Google Apps (possibly not a good idea but soo convenient :/) with a catchall. I didn't catch any offender so far...
However (off topic), spammers are spoofing addresses as if they were coming from my domain. I receive two to three dozens automatic replies from mail servers (this address does not exist...).
I've properly set up DKIM and SPF records, making it obvious that these mails are spoofed, but I'm afraid my domain will end up on grey/black lists... Anyone out here familiar with this kind of issue?
Obviously you can never give out the "something@example.com" address and then assume that everything that goes to that address must be spam, but I've had legitimate contact from companies who have had to email me by removing the + portion because their internal email system wouldn't allow addresses with a + in them.
If you do go this route, I'd recommend using a whitelisting approach. I do get a lot of spam sent to random addresses at my domain.
^[a-z0-9_.]+@(?:[a-z0-9-]+\.)+\.[a-z0-9]+$https://en.wikipedia.org/wiki/Email_address#Valid_email_addr...
Unless the nightlies have a different behavior than the releases, the patch level is not reported anymore. The changes went into 16.0.2 which was released on 10/26/2012. initial report on b.m.o[1] reads as follows:
Steps to reproduce:
1) Load http://www.delorie.com:81/some/url.txt
Actual results:
The User-Agent header exposes the security patch level as either a minor version
number or as an alpha/beta/pre indicator. This data is exposed twice: in the
Gecko version and in the application version.
While it is of value to expose this data to e.g. AMO, exposing it to all sites
makes the browser more fingerprintable (see https://panopticlick.eff.org/ ) and
doesn't serve a purpose more important than user privacy. Point releases don't
change functionality beyond security and stability fixes, so sites shouldn't be
sniffing the patch level anyway.
Making trunk, alpha and beta builds look like release builds for sniffing
purposes reduces sniffing-related failures that waste time when treated as
functionality-related regressions by mistake.
Expected results:
Expected the version numbers to show the major version of the most recent
Firefox beta that Mozilla has shipped and not to show the security patch level
or an alpha/beta/pre indicator.
Additional information:
Internet Explorer doesn't expose the security patch level in its UA string."
[1] https://bugzilla.mozilla.org/show_bug.cgi?id=728831"Within our dataset of several million visitors, only one in 857,908 browsers have the same fingerprint as yours."
As it doesn't allow for plugins, my fingerprint should (cookies aside) match that of any other <popular device> user.
So maybe the solution here is coming up with a 'secure browse' profile that every browser reports the same fake fingerprint.
Security in numbers.
There already is: https://www.torproject.org/projects/torbrowser.html.en
This also has the advantage that no other solution has: it completely hides your location as well, whereas even with a "standard" browser, your IP address + time zone alone can do a lot to identify you.
Just mentioning this lest anyone get the wrong idea that setting your browser to update frequently might be a defense.
Indeed; the user agent is part of the fingerprint.
Google, Mozilla, Opera, can you hear me?
== This. The system needs to be fixed. Need to know (only) vs nice to know info exch, etc.
This is another problem with having an advertising company (Google) supply a browser that is very popular (Chrome). In fact, Safari and IE are also run by companies with large presences in the online ad market.
I doubt Google in particular will risk antitrust suits by blocking these kinds of very, very unsettling but unfortunately legal trackers, which in part are not so technologically different to GA but combine a few more bits of tech which makes them awfully invasive. We might be able to hack technological solutions together here but this stuff rarely makes it out into people's mainstream browsers.
The most important way of securing people's data over the next 10 years is going to be by way of the browser and the mobile OS, but the thing that is most easily achievable is to have a solid browser that people can trust on to implement privacy-preserving technologies. The only browser I can realistically see doing that is Firefox.
Maybe taking a page out of the enterprise play book and using a proxy, like Squid, would make sense. From reading the Squid manual, it seems like it could play a role as it is quite extensible and sophisticated. Making it easy to setup and customize would be pretty difficult from what I can tell, unfortunately.
First, the precise browser version and OS can probably always be identified by checking for supported features, bugs etc. even if the extreme measure would be taken to remove the user agent string.
Add the screen resolution, IP, timing and request patterns (+) and we are all screwed.
(+) e.g. rule out users that are using other sites at the same time. Note that it would be possible to determine if a page is in the currently focused and visible browser tab and forward that information to the tracker.
The trick is not to remove information, but to poison it.
For example, Panopticlick sees that I have dozens of "system fonts", enough to stand out. I want my browser to lie about the fonts I have, based on settings I choose.
There are many details about my browser and system that are irrelevant to what most sites need to do so lying about them should not interfere with viewing a site.
Force them to do detailed packet timing and their costs will go up, and it will become less economical for black hats to play around with your personal data.
I don't know if nuking the user agent string is a horrible idea, but it's less of a problem today than it was 5 years ago: today, a website can assume all browsers conform pretty closely to a standard. Only really advanced features require user agent sniffing (arguably, if you're sniffing the UA you're doing it wrong).
I think we should make that kind of fingerprinting opt-in, not opt-out.
Rather than trying to hide everything another tactic is to provide random misinformation (different user-agent strings, only presenting a subset of fonts and plugins, etc). Enough to defeat the fuzzy matching that does go on.
Sure you've got to be careful that you don't do things that may break some sites that rely on this information remaining stable during a session, but that's got far less common with the frequent browser upgrades that go on nowadays.
imho the best strategy would be to copy one behaviour everywhere, so that there could be no way to differentiate between users.
+-------+ +----------------+ +------------+
|Browser|<==A==>| Visited server |<==B==>| ID service |
+-------+ +----------------+ +------------+
The client data acquisition is done through A (AJAX), then that info is sent through B (API call) to get the identity. The browser doesn't interact directly with the ID service.The data acquisition scripts would be served directly by the web sites.
you can certainly make something that is "ghostery proof", but (1) this isn't it and (2) it would be more complex to deploy and so gain less traction.
It wasn't about this very company, but tracking in general, which can be implemented without serving code from a third party server.
If solutions like Ghostery become the norm, there are still workarounds, and they will catch up if nothing else works.
Not really. If they share data on the server side, they wouldn't be able to share a cookie - they would have to rely on other means to identify you, such as IP address etc. Not entirely impossible, but not as precise either. And that is spoofable through proxies etc.
These scripts will only track you and not receive information.
I've worked on a similar tracking snippet/system for http://www.projectcounter.org/ and this was one of the first things we attempted.
If you block the leadlander domain(s) the script will obviously not run and also consequently won't be able to send fingerprint details back.
In Firefox, the plugins get them to 1 in 860,000 which leaves only 3 possibilities in their DB of 2.5 mln, even though Firefox loads only QuickTime and Flash.
It must be the combination of codecs I have installed. How do I go about cleaning that up?
The browser UA is only one component of the fingerprint, and probably not the most important one.
Plus one for this. I wonder if a plugin alone could change enough info to fool the trackers?
It's pretty easy to guess company name from IP address, especially if you don't care about accuracy. You can kinda sorta do this in Google Analytics under Audience > Technology > Network. That seems to be roughly what they're doing in the screenshots posted. IMHO, this is not the most serious privacy issue on the web.
I would be very curious to hear exactly what percentage of visitors it is able to supply Name and Email for (and how many of those fields look bogus). This sort of individual-level tracking across sites is obviously possible, but I don't think it's common. Google/DoubleClick do not, as far as I know, do any sort of tracking at the level of an individual's name or email address (And why would they? It's asking for regulatory problems and it doesn't really help them much -- they target ads to groups of similar people based on demographics, not to particular named individuals.)
> It's a fair question and one that I asked myself. If the entire service is a fake, then it is an extremely elaborate one because the name and emails of the individuals it did indentify (which I noted was a small percentage) were real.
One one level, I can see why sites do it. On another, one inch higher level, I can see how any site implementing it is so shortsighted that I'm amazed they didn't immediately go bankrupt as soon as they started.
So I signed up for a demo account and installed (and hastily removed) the tracker
Like the parent, I have no idea how this information could have been obtained. It lists search terms, how could a 3rd party track clicks from SERPs to a website not running their tracking code?
Was anything mentioned about the browser used? Maybe when "auto-fill" browser options are enabled for a user there's a way to access that data.
The second a prospect submits a web form, all that previous web activity is tied to their email address (and any other info you collected via the form). You now have a real lead.
I don't see any privacy issues with this.
What I would see an issue with is if the tracking company were sending the IP address and cookie back to a central database to query "Does anyone _else_ know who this visitor is?" and then provide PII any company who uses the tracking service.
The moment you start giving my PII to a company that I didn't voluntarily give it to is when I feel a line has been crossed.
That appears to be exactly what's happening. The email mentions "access to our entire network of identified data ([...] we can identify any visitor [...] if that person has filled out a web form from any other website we are tracking)".
In the case of marketing automation, all the data lives within the system and is used by the company - rather than giving that information out - a very different proposition.
Surprisingly, AdBlockPlus doesn't seem to block it.
Edit: actually it's LeadLander.com as pointed out by NiekvdMaas here http://news.ycombinator.com/item?id=4891764
Hardware manufacturers & telecomms seem to feature heavily:
(a selection) ... Adobe, Dell, IBM, AMD, Box.net, Cisco, CSC, Comcast, Freescale, HP, Lenovo, Motorola, Novell, Qwest, Salesforce.com, Siemens, Symantec, Verisign, VMWare, Vodafone.
And there's several anti-virus/anti-malware companies listed there.
UPDATE: The LeadLander.com site also lists their customers - Microsoft, Motorola, Red Hat and Cisco, among others.
https://easylist-downloads.adblockplus.org/easyprivacy.txt
you'll see that Demandbase is there.
Pharmatrak eventually won on appeal though, arguing that they had no intention of collecting personal information, which exonerated them because only intentional eavesdropping is a crime.
The company in the OP's article could make no such arguments though. I suspect that their main difference is that they make no assurances of confidentiality to the websites using their software the way Pharmatrak did. Which 1) is just really creepy, and 2) sets them up for trouble with users in California, because California's wiretapping statutes say that it's a crime unless both parties agree to it. [3]
[1] http://cyberlaw.stanford.edu/packets001737.shtml
[2] http://en.wikipedia.org/wiki/Web_bug
[3] I'm not sure if this applies to police, but it definitely does to private parties: http://www.citmedialaw.org/legal-guide/california-recording-...
Edit: Added third reference.
"We do not share any information about you or your company to unaffiliated third parties, except as necessary to administer the communications we offer and as permitted by law. We may use a third party service provider to for communications; that company is prohibited from using our users’ personally identifiable information for any other purpose. If you follow us on Twitter, Facebook or on other social media services, we may use information provided by these services to customize our communications to you. We will not share the personally identifiable information you provide with other third parties unless we give you prior notice and choice." - http://www.sencha.com/legal/privacy/
Nearly every company using LeadLander is breaking the law because their posted privacy policies do not state that they are giving a third party your personal information, and that third party is giving it to others.
Edit: It looks like http://formalyzer.com/formalyze_call.js is the specific js file that uploads personal information. Of the sites I listed only clustrix.com is loading that (on the contact form). The other sites seem to be using LeadLander without the form tracking.
IANAL, YMMV, etc.
Doesn't that wrap things up? They'd just argue they've shared it with an affiliated third party.
It turns out that sencha.com might not be sending personal information. clustrix.com appears to be, their privacy policy says:
"The Personal Information we collect is not shared, rented, or sold to any third-parties. We may provide your Personal Information to companies that provide services to help us with our business activities such as shipping your order or offering customer service. These companies are authorized to use your personal information only as necessary to provide these services to us." - http://www.clustrix.com/privacy-policy
I'm not a lawyer, but as a normal native English speaker I read that as they are not going to send my name, email, and phone number to another company, who will in turn share it with with anyone who pays them. But that's what they are doing. They are selling your personally identifiable information.
[0] http://www.marketo.com/small-medium-business/inbound-marketi...
[1] http://launchpoint.marketo.com/strikeiron-inc/747-strikeiron...
Loudly tell them that their spying is unacceptable, then actually follow up on that statement. Ghostery is awesome, but that's a proactive measure. We're talking about appropriate reactions.
But anyways, I agree with what you're saying. If we care about privacy, we have to be loud about it. I just thought it was worth pointing out that facet of Ghostery.
[1] Yes, you can still see the traffic in the web server logs, but I don't see evidence of many companies still doing that. Google Analytics and the like seem to have completely replaced server logs for traffic analysis.
But Ghostery is not able (correct me if I am wrong) to disable the server from logging you. And automatically reading server-logs is not so difficult at all.
The most funny thing here is, that in Germany, you should anonymize an IP-address, when tracking, because of the law, that is concerned with privacy.
But the server logs your full IP non the less.
On the original post: The technology advertised to the author would be totally illegal in Germany. And if I would ever encounter (via Ghostery) a site that uses them and has a German base, I would inform the authorities against them.
I just hate this philosophy of bending/breaking the law/common sense, just because it is possible and might bring in some Bucks. And just because pressure from users might change the regulators minds in the future. It just feels so totally wrong, so disrespectful against fellow human beings, that imho everybody, that has something to do with things like this should be deported to somewhere like North Korea, or the likes. Or like in the middle ages should stand in the pillory (and not in a virtual one).
We can tell the world all day long this is Bad and Unsafe, but within six months it'll be more popular than ad retargeting and the meebo crapbar (because, hey, analytics!).
Doesn't look like it, at least not intentionally. They are trying to capture name, email, phone, and company. Source: http://formalyzer.com/formalyze_call.js
To name a few -
http://trends.builtwith.com/analytics/LeadLander http://trends.builtwith.com/analytics/Hubspot http://trends.builtwith.com/analytics/Marketo
There's a lot of them out there now and mostly all of the big ones are continue to grow in popularity.
But I could be completely wrong; is the guy who relpied with the rx Darren?
Proof: http://o7.no/Z0huP7
I get emailed by them for every startup I'm involved with and that first email is mostly the same every time as you can see in that screenshot. (Compare it with the one posted in the article and you'll see).
They seem to be targeting startups and make it look like some big VC firms are visiting your site to get you interested. I'm not sure how they come up with the 'search terms', but I guess they could just look at your META-tags or make them up.
In their email they do say it's a "mock example", but still I find it very deceptive.
The article goes into depth about how much personal information is sent along to advertisers including a popular dating site's apparently anonymized information about drug use, and sexual orientation.
I think we need a non-profit service that defines a set of privacy licenses (akin to CreativeCommons' licenses) which companies can opt to label their websites/apps with. There would be no policing/auditing [2], but companies found to violate the privacy licenses would be obliged to donate a sum to an organization like the EFF.
That the privacy policies would be encompassed by one simple privacy licence badge would allow users to quickly and easily identify a company's privacy policies. I believe users would gravitate toward using services that display this license.
Edit: it appears such a service is in the works - http://privacycommons.org
[1] http://online.wsj.com/article/SB1000142412788732478440457814...
[2] The auditing process would likely become complex, costly and corruptible
facts (detected by a ghostery at 42floors.com): ClickTale, Facebook Connect, Google +1, Google Analytics, MixPanel, Optimizely, Twitter Button
...by 42 floors. They're still telling all those networks that I visited the website.
On the other side, most startups including YC ones, use some sort of tracking for analytics to improve usability and internal flow, so advocating against all trackers and for all users installing a blocker is a double edge sword.
I don't see a moral issue with retargeting because at its heart it's anonymous - all we know about a user is a string of sites and maybe search words. However, as soon as that data is correlated against personal information, as soon as the real world data and the digital paper trail are correlated and identifiable it becomes sufficiently creepy to me. Who knows - maybe 5 years from now this will seem innocent and benign compared to the mind-reading banners on the bus stops but this seems like a line in the sand I am willing to draw today.
It wouldn't completely work here (e.g. EFF's panopticlick could still fairly uniquely identify me, or IP address would give away info if I'm not going through my VPN), but it improves things.
It feels kind of extreme, but it's worth it to me. My experience is not broken that much, and I feel like various sites are aggregating less about me. These tracking technologies not such an issue now, but I foresee at least the possibility of abuse in the future, so I figure I'll do what I can now if it's not too much hassle.
Lastly, at its heart most of this is about advertising, something I know I'm very susceptible to (try as I might to convince myself I'm not). So the better I am at blocking out these things, I think the less money I'll spend in the long run on frivolous nice-to-haves.
(edit) Does anyone know of a Firefox/Chrome add-on that strips referrer info from cross-site requests? That'd be the simplest way to deal with all externally hosted .js and images that double as trackers.
2 = always send; 1 = send only to same FQDN; 0 = never send.
I can't imagine that I'm alone in this train of thought.
Anyone know if GA's privacy policy firewalls the data collected by GA from AdSense and other parts of Google?
I know they modified the privacy policy a year or so ago to integrate data across all their products but does it include GA?
"Do I _have_ to give you my number before I buy this?"
"yes, but it's for return purposes only"
Of course I received 'promotional' txts the next week. I was hesitant to give it to them for just this reason, and because I acknowledged I had a phone number I felt obligated to give it to him. Dick Smith is a member of a larger chain it's no stretch of the imagination to hook up CCTV cameras to an OpenCV instance and send txts to customers when they walk in.
No matter the law, morals people hold, or customer wants large companies are always motivated by profit margins. The Consumer Guarantees Act, the Privacy Act, the Bill of Rights Act all become murky when you're dealing with new technology, and law will find it hard to keep up.
Dick Smith used to pull this all the time in Aust before they got pulled up over it. I always told them "sorry, not available" and they just moved on.
Today, I get an email from a site that I visited yesterday and haven't heard from in 6+ months. It's too much of a coincidence for me to assume it's random so I dig into their website a little and they're using one of these services.
TL;DR: even though I'm relatively paranoid with giving out details online, one of these networks seems to have successfully identified me and provided my email to a website that I visited, who then reached out and tried to sell me shit.
I was appalled when I saw that we could identify, not only visitor's names, but their friends, access public photos, and all of their profile information. All of this without any action on the user's part and before there were any privacy controls.
This is just the next logical step-federating data collection across multiple sites, not just FB.
I'm obviously in the minority since FB has grown tremendously in the past 2 years but I've not looked back. I dread the forthcoming lack of privacy and anonymity our world is heading toward.
Facebook Connect is the most benign of these sorts of things there are-- it's access to data, and the implementors of its widgets and API-- have an onus to protect it.
Now, of course, there's plenty of bad actors out there, and I'm sure it's sold and exchanged, but technically and legally speaking, you're forbidden from doing so.
I remember being disturbed when I saw that recently, and immediately sought out and installed a social widget blocker.
They were just two examples I could think of off the top of my head though. As the other commenter said about TWP, the practice is common. I see my name and other social data displayed on sites I've never signed up to regularly.
This is a bit over my head programmatically but that doesn't seem possible. If Facebook is serving something to visitors on my site, surely there must be a way for me to capture that data?
But when I see the anger that these types of plugins, especially AdBlock, produces in content publishers I wonder if we're not headed towards a new RIAA/MPAA-style battle front. As online publishers of all kinds get more established and consolidate their power they could start lobbying to regulate against these plugins. It might seem farfetched now but so did paying a tax to the RIAA for blank media, until it happened.
“When a user opts-in to GhostRank, Ghostery sends the following information each time a tracker is encountered:
the tracker identified by Ghostery
the blocking state of the tracker
domains identified as serving trackers
the time it takes for the tracker to load
the tracker’s position on the page
the browser in which Ghostery has been installed
Ghostery version information”
Nothing about the URL you visit. Do you have reason to believe they're lying? http://l.ghostery.com/api/page/?d=news.ycombinator.com%2Fitem&l=304&s=0&ua=firefox&rnd=7639974
http://l.ghostery.com/api/page/?d=news.ycombinator.com%2Fitem&l=426&s=0&ua=firefox&rnd=5747246
http://l.ghostery.com/api/page/?d=news.ycombinator.com%2Fitem&l=346&s=0&ua=firefox&rnd=8989043
Why does Evidon need to know about pages that have no trackers? The FAQ says the domains serving trackers will be identified, not the complete URL path (minus query string parameters) for pages that have no trackers.We'll make a blog post soon, too.
If I'm understanding it correctly, the vendor is offering the following service: place a JS snippet on your website. When a user visits your site, data will be pushed to the vendor's server about the user and what they do on your site (probably keyed on IP and as many other things as they can use to fingerprint). In return for you sharing this data with the vendor, the vendor will give you all of the data on this same user that was contributed by their other clients.
Here is an extreme (yet possible) scenario. You go to a medical forum that uses this software and create an account using your personal email address and real name, both of which you select NOT to be displayed to the public. You then post a message asking about a specific type of back pain you're having. A few hours/days later, you're browsing for a gift for someone, and visit the website of a salon that also uses this software. They can identify that your browser visited medicalforum.com, see the email address and real name you created an account with (since they were passed your form submission directly, without regard to what privacy settings you used for the forum), and see the topic you posted on back pain. So just to be helpful, they email you an advertisement: "Hi {your real name}, we see that you're having some back pain - bring this email in to {salon} for 15% off a massage!"
EDIT: To add, how do you know that you can trust the vendor not to display seriously private data? What if an online store uses this JS, and the vendor has your credit card info, possibly not-so-securely stored? Your information becoming public would be as simple as Asshole Q. Pirate making a fake site with some link-bait, and creating an account with the vendor.
Say you own a sports website, a fashion website, a political website, and a gaming website. The user only specifies a tiny bit of information on each website. Each bit is collected into a single user profile from which they can refer to do things like figure out what product advertisements to show them. They use the same techniques to identify users that don't have accounts, and still collect their viewing/interacting habits and add them to the profile.
Sometimes they'll send you an e-mail telling you to check out their gaming website if you're not signed up, because the comments you write in their other websites' forums have to do with gaming. Sometimes they just sell the information to a gaming company. In the case of Target, they might send your teenage daughter a list of baby products for the little one you didn't know she was expecting.
This is not some horrifying violation of privacy. There is a price for all the free shit you get from the internet. Usually it's paid for by all the personal information you leak onto the net. They're just mopping it up and selling it back to you.
What's the big deal?
In the US we had this thing called "McCarthyism". For a while back in the 1950's you could very easily be fired from a professional job for having read certain materials (mainly those from the Evil Empire of course) or having had certain political discussions when in college.
Just wait a few years until you need to find a health insurance plan that will agree to pay for the expensive medical treatments that will let you live a few decades extra. We'll see if your past web surfing and consumer habits make you worth keeping around.
"Buys K-Pop from Amazon.com" "Visits RedTube at least 3 times a week" "Spends at least 15 minutes in gay porn section" "Made campaign donations to Republican National Convention" "Attends Clivesville Baptist Church"
* no this is not really his web profile. But this is the kind of things web bugs leak. And while your life my be so boring that it doesn't matter to you, may other people have information they would rather not share with every stranger on the net.
(1) And of course society in general used to be and still is pretty anonymous. I can easily buy a newspaper in almost perfect anonymity through regular channels, apparently I need to take special precautions to get the same status online.
I.e. if I have to fill out a form somewhere, I would not only submit it once, but several times (ideally automated), ideally with realistic data, i.e. other businesses in my area (so geo-location won't raise a red flag then).
If I visit the next website which employs the same network, they can't really identify me - they have a big set of businesses I could possibly be (or they just take the last one, which would be fake).
At least currently, they do not seem to verify whether the filled out form can be properly validated, i.e. if the user clicked on a confirmation mail or similar.
Anonymity by obscurity :)
Ignoring the part where you can be "tracked" by company, but that's just looking up public IP records.
After that, I looked at the URL-bar and it took that long for me to click.
http://www.youtube.com/watch?v=6-ZLw2Q7U2M
and the company: http://www.immersivelabs.com/
http://occupycorporatism.com/disney-biometrics-and-the-depar...
Edit: also remember that all a VPN gives you is an encrypted tunnel between your PC and the VPN end point. This means your ISP can't snoop on your traffic and ad providers can't see your real IP (and thus location) but that's all it gives you. It isn't a solution for anonymous surfing. If you want that, use tor.
Something like this is a quiet, terrible thing slinking about unnoticed until it is rather too late. I believe these things have more potential to cause harm than any missile built and kept in stasis.
The developers at these kinds of places don't need to know what they're building. They have many tasks assigned to them and one of them is to write an API that collects a single piece of data. Many kinds of data are collected from many places and put into a database. Reports are made and cross-referenced by an analyst. Final reports are generated and fed to a guy who deals with direct marketing or advertising or sales. Any of these jobs could also be done by contractors or third parties.
You can't just tell people how to make a living without understanding what the hell you're talking about. That's my $0.02 anyway.
(P.S. people that work on missiles are often academic researchers and work for both the private and public sector on the same thing for many different clients, and aren't told what it's used for. the more you know...)
Again, let's not argue over defense contractors or some damn fool thing--when you work for Google, when you work for AT&T, when you work for Palantir or HBGary or whoever, you don't get to say "lol not my department I made swing apps and file dialogs" when you find out they've done something bad.
We need to speak out when people work on harmful technologies.
I'm talking about things like identifying if somebody is gay or republican or kinky and using the information for profit. Aside from selling it to background-check websites and the like, and the fact that it's information people willingly give up about themselves to entities unknown, I have trouble understanding how you can be so offended you think people should lose their jobs rather than develop potential parts of a potential system that could maybe harm someone at some point.
Your assumption about the "cleverness" of developers is misguided. If a guy is told to write a small piece of code which simply takes HTTP requests from JavaScript and plugs it into a database, there is no idea what the fuck that could be used for. The guy maintaining the database also may not know what the fuck he's looking at, it may just be numbers. Are you really so willfully ignorant as to believe every single outcome of every single human action is cut & dry?
I'll even accept your assertion (for the sake of argument only, mind you!) that engineers at a company might only work on some small fragment of JS munging numbers in a database.
At some point, though, an engineer needs to implement the API for a saleable product using that information, or code up a dashboard with element names like "#user-site-history" or "#tracked-profile-visits", or at the very least see the marketing materials the sales folks use to show that the product is competitive due to this information gathering.
Your assertion makes publicity even more important--eventually, some engineer or admin is going to have to get their hands dirty and that is when they need to speak out.
~
To go back and answer your "so what if we have targeted advertising" directly: there is currently no heavily established legal framework of which I am aware that protects metadata about users gathered for the purposes of advertising. I do not know if Google or Facebook is prevented from giving up (for whatever reason!) the results of their ad engine's analysis of user browsing to anyone at a whim.
We (Americans, at any rate) are very lucky that our government at least goes through the motions of liberty enough to not overtly round up deviants and send them off to the camps or send drones after them--this is far from the case in various other countries.
As far as the idea that the information is given up willfully, we're talking about techniques and technology that are really only ten or fifteen years old...the average consumer has not had time to build up any sort of reasonable intuition about what they are sharing or not sharing, or how that information can be linked to other facts about their lives. To say that they've "willingly" given up this information is, I suggest, somewhat misleading.
That said since a (very) long time I'm using separate Linux user accounts to: check my professional email + G+, surf my personal email + G+ + FB (my FB is using a fake but plausible name) and a third one to surf the Web.
The one surfing the Web is linked to a fake online identity: entirely made up, with fake friends / fake G+ circles, fake StackOverflow / OpenID and basically fake everything.
I then only ever surf using a transparent proxy for anything "work related": the IP can't be linked to my fake IP.
It's not difficult to set up: I did set up the transparent company Web proxy (VPN would to too) myself and basically Linux user accounts take care of the rest.
Now I'll start using different browsers too and, why not, maybe Tor in one of the account.
I take it I could take all this a step further and whitelist websites that my "personal" account is allowed to connect to (using iptables' owner-uid mod).
I find it a little annoying though that cookies are not sandboxed by tab, rather than by window/session. If I log into FB in incognito, and then do a little more private browsing in a new tab the FB cookies are still accessible in the other tab.
I guess I could be more vigilant with my browsing habits but I think this is a fair feature to implement at browser rather than forcing user to jump through more hoops to protect privacy. On a side note, when will chrome finally offer API hooks to allow NoScript to be developed for it?!