Ad Blockers Are Also Changing the Game for SaaS and Web Developers
snipcart.com
snipcart.com
If you want to know who is visiting your site, try reading your server logs.
Most ads would be unblockable if you made the ads come from your domain and have them indistinguishable from your normal content in the URIs as far as I know.
This will mean advertisers can't count exact hits to their ads (or at least would be foolish to do so) so they will probably have to employ some kind of web crawler or HIT services to randomly sample the sites who are supposed to be serving their ads to make sure they are.
But eventually the blocker technology will become much better at blocking page element ads more easily and automatically. Then I guess they will have to think of something else.
Or does client side mean the user downloads JS which does some investigation and reports back?
It's no more pointless than using Adobe Typekit for fonts or pulling jquery from a CDN. Both of these use a lot more bandwidth than an analytics ping, BTW.
The use of bandwidth is a moot point. If advertisers hadn't abused the user's good will we wouldn't be where we are today, but there's no putting the genie back in the bottle.
Users don't really need webfonts; text just needs to be readable. And they don't need to grab jquery from a CDN; it can come off the server. Heck, they don't need to grab anything from a CDN--the app server can serve images just as well as Cloudflare.
Do users really need DDOS protection? No, that protects the server, not the user. And it banks their legitmate request--needlessly--through a proxy server, which certainly adds latency. Not to mention that Cloudflare can see everything they're browsing.
Do users need single-page JS apps at all? Why not just render HTML4 on the server like in the good old days?
So why do developers use these technologies? Because it makes the user experience better. And so do analytics. Without analytics, developers are flying blind. And no, Apache/Nginx logs don't capture the same data--especially with more advanced JS-heavy sites.
Users want websites that load fast and are easy to use. It is impossible to build or improve such a site without data upon which to base decisions. That's why analytics are a necessity.
No, it can't, at least not in general. That's what others here are trying to explain to you. It can be very useful, and in both the visitors' and the host's interests, for someone operating a site that has a lot of client-side interactivity to see what's really going on, for example.
Actually, GA only provides you with a [meager] subset of the data that Google gathers. Now it's their data.. not yours, and not your customers. They won't even give up IP addresses so that you can check it against your own logs.
What would happen to your website (or millions of websites) if one of the CDN's that you rely on started quietly issuing evil code to a few, targeted users? Would you notice? Would your users?
I don't think that it'd be too hard for us to add a simple API call that pokes data about screen resolution, browser agent string, language, etc upon load or login, and it'll be far more efficient and private than us sending random data off to GA or similar where they frequently don't even provide us IP addresses of our own site visitors so that we can correlate the data against our own logs.
The data that GA gathers is highly valuable... to Google. They only provide you visibility of the tippy tip of the iceberg.. but ultimately it's your customers' and your data, not theirs.
Don't compromise your users with third-party includes, even Google Fonts (which is still our last holdout on the website.. hm, someone should make a simple web app that gathers names and styles of fonts and provides a zip w/ pre-generated CSS.)
CDN's sound great but they're a huge privacy hole. Ask yourself; what's the profit model? Are they really just an opportunity to gather valuable data on other people's websites and browsing habits? (yes).
Please don't leak your customers' data.
(sorry if I'm an annoying old fart in the rest of this remark but I went back to desktop software 15 years ago (yes that sounds weird) and the last 2-4 years I have more and more troubles understanding discussions on modern web development as everything that was once considered Very Bad is now not only encouraged, but taken as the natural state of affairs - e.g. javascript for core functionality, the 'css is bad, do it in javascript' movement, 'semantic markup should not even be attempted', 'frontend frameworks', ...)
So yeah my question is not sarcastic, I'm just asking for some context.
I suppose there is a job of work to be done downloading the top 1Million web sites and seeing what crap comes through the door - but would be nice to know what the OP is replacing and with what so I know what is derigeur these days
Maybe just that you can look at Piwik analytics - open source and you can host it yourself.
Html/dynamic pages, images, CSS, js, fonts, analytics all from the same server (or at least same domain), and as much squashed together to avoid requests, that was the 'best practice' when I was still 'current'. I don't really understand either what else there would be.
Also, Google Fonts might be Google Analytics in disguise. Who knows.
I use Piwik for tracking, but the self-hosted version. I don't even use newsletter services. I bought a cheap newsletter plugin for Wordpress which I use as an autoresponder email course.
But delivery is all that matters they say. And yet, all Mailchimp and Aweber and whatnot goes 100% to my spam folder automatically. I believe the delivery argument is a myth.
The best part: Decision making is much easier. "So, your product can't be installed on my own server? Bad luck, I won't become your customer."
You can have the fonts and not have any requests leave your site by using something like this[1]. It downloads the Google font data so you can serve the font files and CSS from your own site.
What would happen if EC2 started quietly issuing evil code to a few, targeted users? Would you notice? Would your users?
What would happen if Digital Ocean started quietly issuing evil code to a few, targeted users? Would you notice? Would your users?
I used http://www.localfont.com to retrieve the Open Sans font I was previously using from google font. Maybe that's the kind of webapp you are looking for.
Most of ours would. Subresource Integrity means that all of our Firefox and Chrome users would get mostly blank pages if our CDN tried to pull anything. It's really hard to justify dropping our CDN when they do so much for our load times for people outside of the US (where our servers are located).
> Ask yourself; what's the profit model?
Well, we pay them, so I kinda thought it was obvious. I suppose they could be selling data as well, but it doesn't seem like a great strategy to endanger so much of their userbase when there's already a clear and profitable monetization model.
Of course. They can just ask people.
Let's not pretend that all this data is for developers. It's only for a) advertisers to shit more on your users, and b) sales to micromanage the site into getting more conversions, usually at the cost of utility.
Just prepare your web{site,app} and have the webserver serve it to my browser. If it is too large for my screen, my browser has these newfangled things called scrollbars to deal with it, and I will see that my screen is too small to comfortably view content you graciously share with me.
The same way as desktop developers do. Test things yourself and get you mother (or some non-techy person) to try it out and see how much she/he swears when attempting to use it.
Well, did - I regret to tell you that desktop developers are now using analytics as well, using stuff like DeskMetrics and Trackerbird.
Somebody please kill HTTP-DASH
You do give up the ability to get live stats, but you get better performance and the ability to track more visitors (read: people like me who have blocked GA for years...)
Added bonus: Referrer spam is automatically blocked by default.
https://piwik.org/log-analytics/
EDIT: Almost forgot GoAccess, if you are okay with a terminal app and want live stats (can also be scripted to generate HTML reports) - http://goaccess.io/
Also, please consider adding your Google Analytics use to your privacy policy. They require it, even though most people ignore this requirement and there is no enforcement.
Side note, I just tried @ replying to your companies post on Twitter and got a @your account may not be allowed to perform this action' http://i.imgur.com/9uoEn9B.png Strange, never saw that before...
That's an important point; I'm seeing referral spam in my GA reports which I never noticed before 2015. I manually discount it when compiling reports, and my understanding is that these spammers hit GA UAs at random without even loading your website.
Btw, tail -f /log.log is always fun for live data...
The same way you can gather all available stats on server side and push these yourself. You can override IP too. Obviously not everything will be available, like screen resolution, but you still can capture a lot of stuff.
I've personally never even tried using server-side analytics because I assumed there would be so much noise it wouldn't be worth the trouble.
Hosting my websites on Linux environments, I never stopped using server-side generated statistics (based on Apache's access logs) with tools like AWStats. Used together, I feel that client and server-side stats give a much better image of what really happens on my websites.
But in the end, I feel awstats isn't enough, and I'll take a look at other solutions that have been pointed out in these comments. Thanks!
Why would the author think tracking/analytics is "totally unrelated to online privacy issues"? Baffling
As a publisher, if you want to be profitable, you have to load a bunch of crap from Google, Criteo and others. And, I feel the quality of the JS loaded are slowly degrading. That's shameful point. Advertising networks should buy the best JS talents and release top quality JS. Until then, people should use Ad blockers.
Anecdotal evidence is myself who recently removed Adsense and only direct sells ads. Positives of direct selling is that I get 100% of the sale and I have better control over the ads on my site (static image graphics, no popovers or interstitials). The former makes me happy, the latter improves the experience for my readers.
That said, I do use ad blockers from time-to-time myself because some sites have gotten so ridiculous they are absolutely unusable - I'm talking to you Epicurious and Bon Appetit!
There is no way to scale running your own internal ad network unless you have scores of folks to manage the marketing of your property, the managemnet of your ads, contracts, receipt of payment etc.
There are some possibly problematic blocks with the words 'analytics', 'log', 'event' that might be used in a log viewer, for example.
Also worth noting that I don't think EasyPrivacy is on by deafault in uBlock.
It is. See the readme: https://github.com/chrisaljoudi/uBlock or https://github.com/gorhill/uBlock
The arms race is only going to continue if trackers play the game. I think we'll probably see server side analytics instead, with Google and so making Apache/nginx/express modules and middlewear.
This is another "tragedy of the commons" situation, just like what we've seen recently in advertising, where escalating behavior forced a backlash. If you want to prevent a future backlash against analytics, stop aggregating the data.
Others here have suggested "processing the server logs", but is there some sort of locally hosted thing I can add that will help me get the same stats google analytics does? Or are there any third party hosted analytics services that provide similar services without being aggregated?
> third party hosted
That is the exact thing that needs to be avoided. By using such a service, you are allowing that service to aggregate browsing logs.
It doesn't matter if any particular site logs it's OWN requests; it is expected that if I ask you for a page, you (as the 2nd party) may choose to remember that transaction. Without aggregation, any service only knows about the people that choose to interact with that service. This mirrors fairly closely traditional expectations where e.g. a shopkeeper knows that you walked into their shop, but most people would find it more than a little creepy if that same shopkeeper allowed a 3rd party to kept detailed notes about their customers.
The problems start when you decide to let other people eavesdrop on what should be a two-party transaction, especially when they have access to a lot of these interactions. By aggregating logs, the knowledge about someone changes from known that they used a particular service, to knowing their pattern-of-life[1] (and more).
> what should a site administrator use
I'm truly sorry that there are not a lot of options (that I know of) for better server-log analysis. This area has suffered a lot of damage from the Service As A Software Substitute[2] monopolies.
I suggest pressuring vendors for better analytics software. There may be a market for better local-only, no-services-involved server-log analysis tools. Until such tools exist (or are found), you're in a hard place, because lack of tools is not justification for betraying the activities of your users to a snooping 3rd party.
[1] https://en.wikipedia.org/wiki/Pattern-of-life_analysis
[2] http://www.gnu.org/philosophy/who-does-that-server-really-se...
- Mixpanel (17 matches)
- Heap (heapanalytics.com^$third-party only)
- Segment (11 matches)
Disclaimer: co-founder of Snowplow Analytics, a first-party event analytics platform (https://github.com/snowplow/snowplow). I see 2 entries in the list related to Snowplow, and 26 for Piwik (https://github.com/piwik/piwik), another first-party solution.Even though the error message suggests that they blocked it themselves, somehow that thought doesn't come across to users before they start complaining that our product is broken. Given that google loads APIs mostly in the background through several chains of dynamic JS, there isn't much to do about this, really.
The only problem with server side analytics is that they're pretty limited with functionality unless you have your own internal analytics system which can track alot more information than just page views and general traffic stats.
Presently, surveillance and monitoring is accomplished through third-party requests. Including, yes, sites' own monitoring tools -- Google Analytics, New Relic, and related services.
If the monitoring can be brought in-house, and assembled back-end on the server side, so can the advertising. Which will means that the present generation of site-and-domain based blocking will eventually become less effective.
Have fun storming the castle, kids.
For the sites that can't be blocked with those, there are even more advanced blockers via Greasemonkey/Tampermonkey JS scripts.
I've found a set of rules which are quite helpful at removing / reducing online Web annoyances myself, including killing variants of interstitials, popups, and flyovers. To the extent I'd strongly recommend current Web devs avoid use of those and similar terms in their own CSS.
(ProTip: if you're calling an element "nag" or "tease", it probably shouldn't be there in the first place.)
https://snipcart.com/blog/ad-blockers-saas-web-developers?ut...
"clean" it to:
https://snipcart.com/blog/ad-blockers-saas-web-developers
by:
https://github.com/diegocr/cleanlinks
for Mozilla Firefox
Maybe for sites which require registration (and hence explicitly accepting terms) that could be argued, but there's certainly no implicit agreement to download and process all the stuff the site is offering to the browser.
And "entitled" is just a lazy insult.
I don't mind if a website owner gets that stats for himself. What I (and a growing number of people) mind is that huge corporations collect that stats and then track me down and crunch the data to show me "personalised ads" (of products I already bought) and categorise as a person of interest group XY and then selling that data including my email address to evil spammer YZ.
It is quite the sense of entitlement to think you have a right to that information. You have a right to log what people ask you (the server logs); logging anything more just makes you a creepy peeping-tom.
> I don't mind
That's nice. That doesn't mean everybody else agrees.
You're right, in general - it is the aggregation at Google (et al) that is the real problem. My point is that it is highly presumptuous to assume everybody is ok with a particular type of logging.