Cloudflare Analytics review
markosaric.com
markosaric.com
We are totally serious about building a world-class, privacy-first, free analytics product. At risk of HN cliche, this is our "early work". We are actively working to fix many of the rough edges mentioned here; if we had waited to fix all of them before shipping, we never would have shipped!
For folks who haven't seen it, I suggest checking out our launch blog post[0] which gives some more context around edge vs browser analytics (spoiler: we do both!), why we count visits the way we do, and how we handle bot traffic.
We know we have work to do on the "jagged lines" problem. For some low-traffic websites, we might show noisier, low-resolution data than is ideal. (We've artificially constrained our analytics to query a maximum of 7 days at a time because this problem is exacerbated with longer time ranges.)
My colleague Jamie wrote a nice blog post about how and why we sample data [1]. In short: we have an existing customer base of 25 million+ Internet priorities, whose traffic volume spans 9 orders of magnitude! Sampling data is an elegant approach that allows us to serve fast, flexible analytics for all our customers. Sampling shouldn't be feared, but we know we can do better in some cases. We've recently merged some deep-in-the-weeds improvements to ClickHouse [2] that should result in improved resolution. And we're currently working to store full-resolution data for the smallest websites.
Happy to address any other specific points that folks have questions about.
[0] https://blog.cloudflare.com/free-privacy-first-analytics-for... [1] https://blog.cloudflare.com/explaining-cloudflares-abr-analy... [2] https://github.com/ClickHouse/ClickHouse/pull/14221
Well I would say the opposite, sampling should absolutely be feared. In a lot of case sampling is not an issue, home page, or popular page but in others, including checkout pages, , product pages, and low visibility pages, sampling can make massive difference. When working with sampled data you should always keep it in mind
I've been reflecting recently on how problems like this only exist for companies with extreme scale (similar to how microservices came about to solve FAANG-sized problems). This is a non-issue if you go with a product like plausible (or my personal choice: GoatCounter) for your analytics, because in that case you're essentially just paying them to manage an instance of their open source software for you on a multi-tenant server (I'm guessing here). And if it does eventually become a scale problem for plausible to the point where they start complicating their architecture to solve it, you can self-host or switch to another plausible provider.
If you set out to solve a simpler problem, you can use a simpler solution.
Cloudflare Analytics is server side, we all know that server side analytics is good to get number of GET requests but is almost unusable for anything else. But in their case it seems they are not even trying to filter bots
Page views is always the biggest complains people have when they change analytics tool, "why you have 100k more/less pages views than GA, adobe Analytics?" -> "We don't count pageviews the same way" The reality is everyone is filtering bots a different way and nobody is really doing it well, as a result anyone numbers are wrong. The raw number of pageviews usually doesn't mean anything, that's better to look at difference
There's a 'dead-zone' between the page loading and the analytics script executing, any visitor who leaves during this period won't be counted.
When there's more than one analytics script one will execute before the other so there are abound to be mismatches
GA events have an optional field called "non-interaction", which defaults to false if not present[1].
Pageviews are considered interactive hits. GA's technical definition of a bounce is any visit with only a single interactive hit, leading to the layman definition of bounce rate being "people who viewed two or more pages".
Events though, are also interactive events by default (non-interaction = false being a double negative). So the moment you implement event tracking, if you don't explicitly control that parameter, the definition of "bounce rate" expands to include "the number of people that triggered my event tracking" and your bounce rate will become next to worthless the more liberal or trigger happy your event tracking becomes.
A general rule of thumb is to make "passive" event tracking non-interactive, so that it doesn't impact bounce rate. Things like scroll tracking, timers, etc. Those are all all subject to passive behaviors such as background-ed tabs, impulse scrolling when you first get to a site, etc. And not necessarily representative of active engagement with your site/content.
Then make "active" event tracking (such as clicks tracking on interactive elements or outbound links) interactive. That way bounce rate becomes more representative of active engagement from users. Then optionally, and deliberately, make certain passive tracking events interactive for the express purpose of impacting/influencing bounce rate. For example, a website with lots of long form content but few interactive elements may want to make their passive "30 second" timer event an interactive event, based on business logic that someone spending 30 seconds on an article is considered engaged with it and, having hit that threshold of passive engagement with the page, should. no longer be considered a bounced visit. But every other timer would be set to non-interactive, and scroll tracking would be set to non-interactive (so you don't get false positives from people that scroll all the way to the bottom immediately).
Sending non-interactive events continues to increment the duration, so you still get that value from sending the non-interactive events. But without it skewing your bounce metric.
That said, you rarely see that much care (or understanding) put into it in practice. Familiarity with how implementation impacts interpretation tends to be relatively rare, and a "bad" implementation leads to juiced metrics. So few analytics/marketing teams are even aware they can actively define business rules around what is a bounced visit and what isn't, so most don't and take whatever the system spits out as gospel, implicitly leaving it as "undefined behavior" subject to the whims of however it's implemented.
And even if you educate those teams on such things, no one wants to eat the hit to their vanity metrics that result from truing up an existing implementation that has such issues. So even broaching the subject tends to just lead to consternation and continuation of the status quo[2].
[1] https://support.google.com/analytics/answer/1033068#NonInter...
[2] The bitter grumblings of someone that pokes this hornets nest on a frequent basis, and potentially a skewed perspective rather than objective fact.
The biggest thing to realize is that you can control that particular calculated metric in the first place. At which point, you can decide subjectively what constitutes a bounce for your business/site, and explicitly match your implementation to that business definition. Which may even fluctuate between sections of a site, but rarely aligns with what you'll get out of the box leaving everything with Google's defaults.
FWIW, this isn't an issue for server-side (or in our case, "edge-side") analytics.
> But in their case it seems they are not even trying to filter bots
We try very hard in fact :) We just have some work ahead to surface this better in our analytics UI.
For more detail: we actually assign a "bot score" to each request. Customers customers of our Bot Management product can use this information to block bots at the edge. We are working right now to show the distribution of Bot traffic in the UI so that anyone can see how we are classifying traffic.
This is so true! Hopefully, the issues mentioned in the post will be fixed soon, considering this is a quite new product. I can see high value for small personal websites, as they would be able to get something for free and focus on user privacy.
Meanwhile, the mentioned issues might be quite beneficial for other companies, such as Plausible or Fathom. They are quite mature on the privacy-focused analytics market, and offer things that Cloudflare analytics doesn't. People might start on Cloudflare Analytics (free) and eventually migrate to something more full-featured.
>> The main danger I see after paying for and trying Cloudflare Analytics is that they may not believe in this at all.
I disagree with this part of the post because developers who opted for Cloudflare Analytics simply because it was "privacy first" are likely to explore other alternatives before selling their souls to GA.
Interestingly that was the exact reason why Google Analytics got big as they had basically no bots in their statistics (in the old days running js was so complex no crawler or bot did it). And now we‘ve done a 360 degree turn and be in the same position with inflated server logs.
Is there any good analytics solution which is only server side? Getting some good statistics but not exposing to visitors they are tracked (even in a friendly mode) would be nice.
I'm in the same boat. About a year ago i migrated over to Matomo (fka Piwik) and away from google analytics for all my (and family, friend) web sites that i manage. While i really don't have many big complaints of Matomo, i really wish i didn't have to use javascript for this. I started looking at - but did NOT implement - GoAccess [http://goaccess.io]...but don't think it fills what i need. Honestly, for the websites that i manage - and in these days where i want to preserve both my privacy and the privacy of my web visitors - I'm wondering if i should just very well shut off all tracking on the front-end, and try using GoAccess or even AWStats, and process stuff offline. I'm not running any ecommerce site after all. And, removing yet another little element (javascript) from websites helps with performance for visitors. Sorry, not helping much, maybe give a look at GoAccess...? Good luck!
>Google Analytics has a big issue with adblockers and on this site alone more than 25% of visitors block it.
>Plausible avoid it by not collecting any personal data in the first place.
Seems like this is wrong. The default filters on ublock seem to block plausible's script just as much as google's.
Could likely be argued that the greater ability to filter out bots with a script makes it preferable as compared to server-side analytics. But you are definitely under-counting legitimate visitors if you go that route.
Cloudflare is a profitable public company.
Okay, this is nice to distinguish but I would kinda like to know both numbers: what's the raw load coming through CF to my website and what are the "real human users" -- and what tools can CF provide to block the former while allowing the latter? If CF can provide easy to configure and understand WAF rules to filter more of the bots and other raw traffic that I don't want then I'm going to save on my AWS/Azure/hosting bill.
I actually use Clicky and self hosted Matomo for most of my properties. I would love to use something like Plausible but to me (SEO pro) the most important feat is lacking: landing pages and average amount of actions and/or time spend per visitor per landing page.
Hopefully their push will make the issue more prominent to website owners and help Google Analytics alternatives to grow.
Those are two completely different metrics that you cannot compare!
We don't show the number sessions/visits at the moment. Most seem happy with the 'unique visitors' number and we don't want to add extra data to the UI unless it's truly useful.
If you want that, I suppose GA or Matomo are for you? Has to be something to differentiate them, pointless to expect or want them all to be the same.
If anyone is curious, that howtotomakemyblog.com referral site in the CF dashboard seems to be a domain that redirects back to the same domain that's hosting this blog post.
It didn't appear in his analytics tool's referrals. I guess it's because his tool follows redirects and ignores listing same site referrals while CF does not?