Shadow traffic: site visits that are not captured by typical analytics providers
blog.parse.ly
blog.parse.ly
I remember when the ONLY analytics were those you could derive by analyzing your http logs. Which have useful information in them. Things like source IP address (which can be geo tagged), a bunch of HTTP headers (which are full of information too), and a timestamp which tells you when it came in and from where. Not to mention session cookies which take zero javascript to implement.
I've been retooling my site slowly to only use these analytics (less the cookies) because I value people's privacy while browsing as much as my own. During the transition I've been comparing what I can pull out of the logs vs what Google's analytics gives me. Sure, Google can do wonders, especially if the person is coming from a browser where they are logged into Google. But, as the article points out, they miss everyone running noscript and/or other privacy enhancers like Privacy Badger from EFF.
I don't feel like I'm going to miss the Google added insights.
> Option 2 – Server-Side Tracking
These guys can write java script that animates a web page so that it looks like a turning page and its "too technical" to pull data out of a server log?
And for those saying relative direction is all that matters, I guarantee you the behavior of users with adblocker installed is very different from those who can't be bothered or don't know how.
If so, what open source web server analytics tools do you believe is best? E.g. https://goaccess.io
Obviously there's a gap between what trackers say and reality, bigger for some demographics than for others.
Were reporting using Firefox.
Also, not surprising, Firefox security leaves a lot to be desired.
Huh? Blocking google analytics tracking is a positive, not a negative regarding security.
How proud was the 15-year old me with my first .com domain, having over 100 visitors per day. Little did I know that the actual number of visitors was much, much less than that.
There is no reason to design your website in a way that makes your legitimate analysis use cases depend on Client-Side computations.
If Server-Side tracking looks too complex for you, you might want to reevaluate the balance of technical knowledge in your enterprise.
One disadvantage is that if your visitor/customer has javascript turned off, you get no data. This was a concern in the early days of client-side analytics, but not really any more.
A more modern disadvantage is that ad blockers might prevent your analytics script from running. However, this is only a problem for client-side analytics packages that are hosted by ad companies, like Google Analytics. It's not a problem with the concept of client-side analytics in general.
EDIT to add:
Another advantage is that only measuring things in the browser makes it a lot easier to exclude non-browser traffic like bots and spiders from your reports.
That's also a disadvantage because you can miss server-only events like "hot-linked" images or PDF downloads straight from Google. On balance, though, we care a lot less today about hot-linked files than we care about excluding automated traffic.
And in my experience, culturally, client-side packages were a huge help in getting management off of pointless vanity metrics like "hit counts" and caring more about human metrics like visits and time.
Counting the size of the executed JavaScript, because that’s what matters more than the compressed transfer size:
ga.js is 45KB, matomo.js is something like 50KB. The “new breed” of trackers are currently commonly 2–5KB (though generally if they were written more carefully they’d be well under 1KB), but they’re sure to continue growing because such is the nature of code.
By using client-side tracking on a different host name, you’re making the browser establish a new HTTPS connection—which can happen in the background so long as you do it properly, so it’s not of itself a serious performance issue—and parse and execute probably 50KB of JavaScript, which blocks the main thread while executing. On a substantial fraction of the devices your site will be running on, that’ll block for more than a hundred milliseconds (to say nothing of the couple of hundred more of CPU time that parsing took, which wasn’t blocking, but was taking away from other things it could be spent on), and on older and slower devices it’ll be adding several hundred milliseconds.
Seriously, parsing and executing JavaScript is slower than you realise. If your site uses JavaScript of your own, using just one type of client-side analytics is probably slowing useful page load down by 0.1–0.5s.
If you want to know click paths you'll need additional data in your logs.
If you want to know how much time the average user spent reading an article on your blog, you are probably out of luck using logs.
You’re making a completely unfair comparison here. Client-side analytics is performing aggregation as it goes; server-side analytics can do just the same, and serious packages in that space do do that. As an example that readily springs to mind, you can feed server logs into Matomo and use it just as effectively as when you feed client logs into it via matomo.js.
On click paths, they’re largely just an artefact of aggregation, so long as you can track individual users (which you admittedly won’t get out of the box in server logs, and perhaps that’s what you were referring to, but you can definitely make it happen). Multiple tabs thwarts doing the “navigated from page A to page B” form of path tracking correctly purely server side, but that’s not a realistic form anyway and isn’t what you’re likely to use; rather you use “page A was loaded, then page B was loaded” and just guess paths to be adjacent loads, which is perfectly compatible with server side analysis.
How much time, yeah, that’s one that you can’t do any sort of good judgement of server-side. Not that client-side tracking is particularly excellent at it. All up though I suspect that so long as your bounce rate isn’t too high you’ll get decent enough figures from reckoning the time until the next page load and eliminating statistical outliers. If bounces are too high the figures could be misleading because bounces are likely to behave differently. But people often put far too much effort into trying to track everything, where it’s commonly quite sufficient to take a smaller sample and extrapolate to the whole.
I don't know how to help you if you can't look at log time stamps and figure out click paths for users. It was a solved problem twenty years ago.
It really doesn't matter how long it took someone to read and article on a blog. They either saw your ads or clicked affiliate links or they didn't. It doesn't really matter if they finished an article or bounced right off the page. It doesn't matter if they took ten minutes or twenty minutes to read an article. You don't know if their kid interrupted them half way through or if they're just a slow reader.
That kind of shit is just meaningless made up metrics. It's the Gish gallop of web advertising. Throw a wall of bullshit metrics at site owners or advertisers to justify them paying for the privilege of listening to that bullshit.
Shameless self-plug: if server-side analytics is too complicated for you, consider using the tool that we just launched to help with that (other functionality gets included as well): https://www.nettoolkit.com/gatekeeper/about
This is much less likely to be blocked if you self-host it — breaking requests to your server will break the app, whereas blocking common cross-site tracking services is popular because there are few drawbacks for the user.
I guess your options here are, do collect metrics in JS and hope that whatever the reason that 20% of visitors don't show up in Google analytics isn't preventing them from using your site.
1. Server side analytics
2. Client side analytics for information akin to what you're asking that Server side misses
3. Crash analytics for client side errors
As mentioned though, you're only getting partial info from some of those options. It also gives you a chance to decide which of these you _really_ need and hopefully eliminate anything you don't
% of a video watched for example would be broken on the server side due to player buffering.
I would consider that to be private information regardless of the reason. Why should eg Youtube know where I stopped in the video?
> would be broken on the server side due to player buffering.
Would it, really?
You don't think that you can tie together "this much of video was buffered and _possibly_ displayed" is useful information?
You don't think that "60 seconds of a 10 minute video was buffered and _possibly displayed_, and another 30 seconds of buffer was requested every 30 seconds for 5 minutes" is useful information?
You don't think that you can determine that the user stopped watching the video after between four-and-a-half to five-and-a-half minutes of the video had played?
I suppose if nothing else it's good to know so you can immediately beef up your numbers +20% in your slide deck for VCs.
What if you're trying to get stats about people who don't convert or otherwise give you a signal that they're using the site?
I've run into sites where things like signup or checkout are blocked behind an analytics tracker (Adobe used to recommend running theirs in a synchronous navigation-blocking mode) which meant that any problem with that service was completely invisible unless they contacted you to complain.
I also remember people wondering why Firefox users stopped using their site when they shipped the release which enabled tracking protection by default.
1- present a new concept to readers (shadow visitors)
2- show how this concept is scary and bad for your business (your analytics is off by 20%!!!)
3- present 2 options, which by the way are free, but immediately shit all over them (Server logs! But that’s hard and complicated! Edge logs! But getting these are hard!)
4- present your company’s product as option 3, which surprisingly, have no downsides and isnt shit upon
5- profit
What disingenuous garbage. You should be ashamed Parse.ly.
The right way to do this is do steps 1 and 2. Then show in detail step 3: how to solve the problem with easy options, ideally with free and open software. It’s ok to show edge cases, corner cases, or just shear scale issues that makes these options challenging.
The difference is that good product marketing pieces show people how to solve the problem and offer a solution to do that at scale or in a automated/hosted way so the customer doesn’t have to deal with it.
If your product marketing content’s message is “you are screwed unless you buy our product “ you are doing it wrong
Script kiddie attack bots are generally fairly obvious as they hammer away at things like /wp-login.php for days on end regardless of what error codes the server returns.
Most other bots are pretty evident just by looking at access patterns. Just identify their IPs and drop them from your analytics.
https://twitter.com/amontalenti/status/1165262620959617025
Specifically: I noticed a huge difference between the metrics we were reporting on my blog post in Parse.ly, and the metrics being reported by my personal blog's Cloudflare CDN (caching the content).
Ironically enough, this traffic was all coming from HN and the post was itself about modern JavaScript[1].
Since then, we've also been hearing from a lot of customers about various scenarios where traffic is either under-counted or mis-counted. For example, something that has been tripping us up lately is that our Twitter integration relies (partially) upon the official t.co link shortener[2], and yet, due to modern browser rules related to W3C Referrer Policy[3], the t.co link's path segment is often not transmitted to the analytics provider, and thus the source tweet for traffic cannot be easily ascertained.
I firmly believe in privacy and analytics without compromise[4], so the team is trying to come up with ways to at least quantify shadow traffic at an aggregate level, and to ensure legitimate user privacy interests are honored, while making sure they don't break legitimate privacy-safe first-party analytics use cases.
As a developer, something that concerned me recently was realizing that Sentry, the open source error tracking tool with a SaaS reporting frontend and a JavaScript SDK[5], gets blocked in many conservative browser privacy setups. Though the interest to user privacy is legitimate, I think we can all agree it'd be better for site/app operators to know when certain browsers are hitting JavaScript stack traces.
[1]: https://news.ycombinator.com/item?id=20785616
[2]: https://help.twitter.com/en/using-twitter/url-shortener
[3]: https://www.w3.org/TR/referrer-policy/
[4]: https://blog.parse.ly/post/3394/analytics-privacy-without-co...
In light of that, presenting "existing analytics services like parse.ly" as one of three solutions on "how to measure shadow traffic" seems borderline disingenuous. If you can do it, why not say so plain and clear? If you can't do it, why do you mention yourself as a solution? Or is it only other services like parse.ly that can do it?
It also rubs me the wrong way how both the blog post and your comment has an undertone of "it would be better if users didn't have as much tracking protection". Just take the framing of your last sentence as an example...
Here they are:
- Server-side proxy: The blog post I linked about The Intercept uses this setup. Basically, a web server run by our customer captures all the traffic inside their cloud or hosting environment. That traffic is then logged and proxied to our data capture server, with some data scrubbed before we receive it (e.g. IP address removed), with data sent via our server-side protocol.
- First-party custom domain: We spin up a server and HTTPS certificate and the customer points their own subdomain (via a DNS A or CNAME record) to that server, which serves as a proxy. We originally built this facility to clarify data ownership in a GDPR context -- where the customer is a data controller and we are a data processor, so the controller owns the domain where data ingest happens.
This and the prior solution have the side-benefit that Parse.ly couldn't do any cross-site linking even if third-party cookies were enabled in the browser. We never do this anyway, but both of these setups make it technically impossible due to the browser rules around cookies and domains, which is a nice security improvement. But it also raises other issues, like the fact that the customer setup is more complex, with more moving parts.
- API connection to CDN. This one is not actually productionized but was merely prototyped. We'd pull basic per-day and per-page CDN server request logs, and compare that to our pageview counts to understand the delta, which is likely mostly shadow traffic. The upside of this solution is that it might be pretty easy to setup for customers, the downside is that we'd have to build connectors for a lot of CDNs, and through market research we have learned that larger customers might use multiple CDNs at once (believe it or not).
- "Fallback" logging of blocked page loads. This one was also just prototyped, but the idea is that some JavaScript code would detect whether Parse.ly JavaScript SDK was blocked from loading, and if so, a basic privacy-safe "this page's analytics were blocked" event would be sent to a domain owned by the customer, perhaps one that ensured scrubbing of all details other than the "fact" that a block event happened at a particular timestamp. We actually prototyped this particular idea on our own marketing site because we ran into issues with Marketo & Parse.ly data vs our server logs and even our lead capture forms. (That is, situations where a lead was captured for someone with "zero pageviews", because their session was shadow traffic but their form fill nonetheless happened.)
Re: your comment that you sense an undertone of, "it would be better if users didn't have as much tracking protection", I have no such personal or professional belief, and I can assure it isn't a view held by our company/team. I understand the motivation for tracking protection and we even suggest use of Mozilla Firefox's tracking prevention option in our privacy policy.
But there's no doubt that it is leading to confusing data discrepancies for site/app owners, and I think site owners have a right to a basic understanding of how much of the traffic they are paying the hosting bills to serve is actually perceptible to their observability/reporting, even if the only detail they get about that visit is "the visit happened", similar to the level of detail they get from server logs or CDN logs as a matter of course.
> I have no such personal or professional belief, and I can assure it isn't a view held by our company/team
Fair enough, it might just be my interpretation that is colored by the context of the content being written by an analytics business. Although I'm still at a bit of a loss what the implication is in the sentence I was referencing. Could you spell that out for me?
That is, aim was to introduce this idea of "shadow traffic", since most of our customers/prospects don't even know it exists! (Whereas, for example, "bot traffic", which can inflate analytics numbers, is a well-known problem.)
The sentence you are referencing is about Sentry error tracking, right? All I was intending to say is that sometimes, tracking protection throws the baby out with the bathwater. The end user wants to avoid creepy ads and privacy leaks. But, instead, they are blocking error tracking tools, whose primary purpose is to catch and fix frontend coding bugs. And then when those same blocking rules become browser defaults, it can end up in a situation where whole classes of users don't have errors tracked/logged, merely because the site owner (quite reasonably) chose to use a SaaS for error tracking/logging, rather than, say, rolling one's own self-hosted system for that commodity use case.
I don't want site owners to see that something has failed on my machine -- especially if it's something unique about my setup. To suggest otherwise is to miss the point about privacy. So no, I don't agree.