Lightweight Alternatives to Google Analytics
lwn.net
lwn.net
Via Docker & docker-compose it's quite easy to install and keep up to date and Matomo is open source, well maintained, very well behaved and pretty hands off.
And I configured it on my websites with cookies turned off [2] and with IP anonymization [3]. In such an instance you don't need consent, or even a cookie banner, because you're not dropping cookies, or collecting personal info. Profiling visitors is no longer possible, but you still get valuable data on visits.
Note that if you want to self-host Matomo, you don't need more than a VPS with 1 GB of RAM (even less but let's assume significant traffic) so it's cheap to self host too.
And I disagree with another commenter here saying Analytics is just for vanity. That's not true — even for a personal blog analytics are useful to see which articles are still being visited and thus need to be kept up to date, or in case content is deprecated, the least you could do is to put up a warning.
And if you write that blog with a purpose (e.g. promoting yourself or your projects) then you need to get a sense of how well your articles are received. You can't do marketing without a feedback loop.
Some examples: I maintained a Vim ChangeLog for a while (which is quite some work), and turned out no one was reading that, so ... why bother?
In another case, I wrote an article about "how to detect automatically generated emails" and I thought it wasn't actually that interesting and no one read it so considered archiving it, but turned out quite a few people end up there through Google searches etc. and I ended up updating it instead of archiving it, as it was clearly useful to people.
My wife as several e-commerce and our weekly "walking through the matomo screens over a glass of wine" has learned us that there are important niches. And what those niches are (vegan smartphone covers, Fairphone flip-cases, fairtrade and environmental-friendly mouthmasks).
Sure, you also need customer interviews and old-fashioned market-research, but your webapp and website is telling you a lot about your users.
And sure, often, you don't need any metrics. But just like Carpetsmoker above, I see a lot of value in metrics. Just don't fall in the trap to collect "you-never-know" metrics: that is privacy-invading, a liability and requires a scale beyond anything you really need. You don't need data-lake, distributed ETL processes and whatnot to find out that there are products in your webshop selling better than others because you are doing well on natural searches for that product.
You get IP and user agent in those logs if you want to roughly track visit to conversion metrics.
0 - https://gist.github.com/mike-seekwell/83ac75c82a943e287a7abe...
1 - https://cloud.google.com/appengine/docs/standard/python/logs
So now I pull out GoAccess (which reads the server logs) from time to time. I find that my Atom feed is the vast majority of traffic to my site, which Matomo couldn’t tell me. I should implement pagination on the feed and see if that helps. (Or limit the number of items in the feed, but conceptually I rather like everything being accessible from the feed. Wonder how many feed readers support pagination?)
So I changed to GoAccess too. I don't check it too often, just when I want to see the impact of some spam/publicity posting around.
It would be very useful. Like you said, it's nice when everything is accessible from the feed.
<link rel="next" href="…"/>
Deliberate semantics were defined for this as part of AtomPub, https://tools.ietf.org/html/rfc5023#section-10.1 (before that, it made sense that it would mean this because of the relations registry, but nothing had been defined). It’s clearly applicable to Atom syndication in general, but it’s definitely more useful to AtomPub. I have no idea how wide client support is.Also I disagree about the slowness.
The script is loaded asynchrously, it does not block the page and I measure my loading times, which are really good actually. Just did a measurement and my front-page loads in 271 ms and this includes all network requests, including Matomo.
I don't think this is a real concern, but rather a premature optimization. If GoAccess works for you, great, but that's not something I can use due to CDN.
But since you’ve raised the script part, 50KB of JS loaded from a new host is perhaps surprisingly much work, especially on slower devices. I find the difference between running no JavaScript at all and running Matomo’s client script, even asynchronously, to be easily visible.
Is your audience just the people in your locality, or the entire world?
Do you have a way to filter out your own visits in this case? On small pages I find that my own clicks and events during testing contaminates the statistics.
But I was hoping for a simple DIY guide for setting everything up? I mean with Google Analytics a dummy can set it up very quickly. I know it'll take more work with matomo, but I need a little more details then just "use docker".
I know I need a VPS, something like Digital Ocean.
I know I need the matomo docker image.
I have a full deployment of Matomo in Kubernetes on gitlab [1] if anyone would find it useful (includes correct settings for running in multiple pods).
1: https://gitlab.com/pcgamingwiki/webanalytics/-/tree/master/
You unzip it. Click through a setup wizard. Done.
(I also run in readonly noexec in an fpm chroot but that's not necessary.) I set it up for a couple of clients since the Piwik days and it's been pretty much set and forget, apart from the occasional upgrades.
There are plenty of advanced functionality which probably few people understand and use. For my personal projects I am fine with log analytics which I mostly use goaccess for.
Matomo is just a set of php files. Upload them to any php hosting with a mysql database, in a matomo or matomoanalytics folder and point your browser to your domain name/matomo/ and the install setup should begin.
If you have 0 experience with docker and just needs matomo then forget about docker and start from the php file with a standard php/mysql host.
Does this mean each page hit cannot linked to be any other? For example, can I see that a visitor viewed a particular sequence of pages?
I mean, I'm not making the argument that analytics are useless, but this seems like the worst possible example. You can do this trivially with a script to analyze your server (e.g. Apache) logs. And you don't need "a VPS with 1 GB of RAM" for that - which is four times the RAM of the VPS my personal website has run on for the last half decade.
This approach also uses no client-side javascript to collect data, so you wouldn't have to alarm users with potential privacy threats, because nothing is stored other than what's in the HTTP headers.
There's lots of take for granted there. It's no wonder copy and pasting a <script> one-liner caught on.
There's also some very useful information that's hard to get from HTTP headers, screen size being the most obvious one.
I don't think it's fundamentally more privacy-friendly. The real problem with JavaScript trackers from GA and ad networks is that they'll try to profile you with tricks like font metrics, audio API, set cross-domain cookies, and whatnot. But that doesn't apply to either Plausible or GoatCounter (and in the case of GoatCounter, the count.js script is intentionally unminified so it's very easy to see what exactly it does if you care to do so).
GoatCounter runs fine on constrained environments; it has a RES memory size of about 30M, and for the first few months goatcounter.com ran on a $5/month VPS which was mostly just sitting idle.
I run a SaaS and what matters for me is paid subscriptions. "Visits" (even if by humans, which is hard to tell) really do not matter much. Yes, I do want to increase conversion rates, and run bandit experiments, but I'm better off doing that myself.
What also matters are search terms, but Google's search console (or tools, or whatever it's called this week) provides that.
Turning off Google Analytics was hard to do psychologically — the Fear Of Missing Out is strong. But it turns out I'm not missing out on anything, except some dubious vanity data. And I'm making the web a better place in the process.
The actionable part never occurred to me, but makes so much sense. What action could anyone really take, based on the data presented by Google Analytics? On top of my head I can actually think of anything you could easily get from server logs.
I know Cloudfront logs can sometimes drop, but is there a more important reason you're talking about?
Based on what? Server side logs records what the server itself sent. That sounds like it should offer far more accuracy than what a JS only solution would do.
That being said, if the code for analysing the server logs is lousy, it's not going to help. :)
> Server side logs aren’t an accurate way of doing analytics...
What is your thinking behind that statement, as it sounds incorrect?
Note that I'm just talking about distinguishing people, not identifying or tracking. A basic question like "how many people came to my website today" is more accurately answered with client-side analytics than by analyzing web server logs.
How can client side analytics be more accurate, when ad blockers stop visits from even being registered by client-side analytics?
1) That you actually care about this metric. I don't, I do not get paid by the number of minutes spent on pages, I get paid by the number of signed-up subscribers who use my software to make their workflow easier. I can (and prefer to) use bandit testing to measure the performance of redesigns.
2) that it can be reliably measured, which I don't think it can.
With regards to point two, being reliably measured sounds to me like perfection is the enemy of adequate. Perhaps in low volumes you can't measure certain stats like this reliabily but in large volumes I think it's useful.
Even if the numbers are off by quite a few per cent, I can easily see how some site operators might benefit from knowing e.g. that visitors close one tutorial much quicker than the others.
I am able to measure everything 100% accurate. But this is really really expensive. It's a trade-of.
The point is, there are clearly actionable points GA can offer. It's disingenuous to think there are not.
- Conversion rate grouped by browser and browser version to find browser specific bugs. - Total revenue per user per marketing channel / search keyword / ... to optimize budget allocation. - Revenue by mobile OS version share to decide testing procedures to not optimize for users that don't contribute to your bottom line. - Client side loading times per user location and provider to optimize infrastructure placement. - ...
It is easy to track all aspects of subscriptions including recurring revenue. The data quality depends on the data you send to google analytics and is not "dubious" but in your own responsibility. And google analytics is really good in separating bots from human visits.
If google analytics did only report vanity metrics to you, you most likely did not use it the right way. Maybe you missed the segmentation tools to find groups of users for whom your service did not work out?
That certainly wasn't my experience.
But more generally: I optimize touchpoints using automated Bernoulli bandits with multiple variants. And I track the metrics that really matter (like signups, MRR, churn, etc) very, very carefully. My point was that just adding GA doesn't bring much value, and makes the web worse.
Another example of how page visits don't matter: it's easy to get a huge spike of HN users clicking through. But if my SaaS has no relevance to HN users (except as a technical curiosity), this doesn't matter at all. It won't change my revenue, so it's irrelevant.
Unless you run a site with ads, focusing on page visits doesn't make sense: it's like measuring the performance of a supermarket and getting excited about increasing traffic on a nearby highway. Could it influence your sales? Possibly. Is it actionable? Nope.
The IAB list is still a helpful baseline (if overpriced if you want to lease the list itself[3]), but all it's doing is applying suppression based on things like known bot IP ranges and user agents. It's far less effective than it used to be since it's so incredibly easy and cheap nowadays for anyone with an interest to spin up a rendering bot. Now you've got to supplement it actively with your own set of filters and heuristics if you really want to get rid of bot traffic polluting your data.
[1] https://iabtechlab.com/software/iababc-international-spiders...
[2] Scraping bots were cheap and easy to run before, but didn't render the GA javascript code so never showed up in analytics. It's only bots that use something like a headless browser to render the page that show up in GA, and those have only become commoditized and cheap/easy relatively recently.
[3] Google applies the IAB list to your traffic for free if you check the setting for it, but if you want to use the IAB list yourself you have to pay $4k - $14k annually to lease it from IAB.
You can actually set analytics to exclude traffic from a certain ISP or set of IPs on the View's Filters page. There is some legit traffic that comes from data centers, but only a tiny fraction.
They can, and they do. Here[1] is a random example of applying that to a supermarket (the other pages on that vendor's site shows some of the other ways the tech can be used).
Solutions tend to use a mix of video camera data and cellular data (wifi and bluetooth beacons[2]) to track in-store movements and behavior.
Done well, they can also connect individual store journeys to specific point of sale transactions, and understand hoe variations of in-store behavioral patterns ultimately influenced actual transactions. Which opens the door to running in-store optimization experiments very similarly to how you'd run A/B optimization tests for websites/apps.
[1] https://retailerin.com/en/retailerINfor/customer-behaviour-a...
[2] https://www.thinkwithgoogle.com/marketing-resources/retail-m...
That said though once I know which marketing tools are effective, there's nothing more that GA does that CloudFlare couldn't just tell me anyway (i.e. am I getting more or less traffic) and I'll probably drop it as it's one less dashboard to look at - like you said that conversion to subscriber _is_ the ultimate metric for success.
This won't work for everyone and will be of little value to many. However that tool just doesn't seem to exist on most if not all of the lightweight alternatives. Matomo is the only alternative I have seen implement this feature, though those with more experience with the alternatives will hopefully show me I am wrong on that.
But the fundamental problem is that analytics provides data and information, when what people want is knowledge and wisdom. But you need to do the work to get it. No amount of analytics is going to tell you that your Facebook ads are underperforming their potential because your buy button is hidden by an overlay in the built-in Facebook browser, especially if Facebook is your best performing channel overall. Analytics can’t tell you what’s not there. There’s no substitute for having developers and marketers actually purchase your product, themselves, on the channels your customers use. And too many companies don’t do that.
Just a hypothesis I have, that most e-commerce companies are leaving millions to tens of millions of dollars on the table by not giving all their employees a corporate credit card and having them purchase their product on it, say, once a month. They’re losing out on far more than they’d end up paying for the few instances of fraud.
I then broke up the main page to many smaller pages and notice people still leaving right away but my documentation page got more overall clicks since the main page didn't massively overload them
I guess analytics don't give you answers but know what pages they tend to click on and what happens when you make a page more simple/more heavy you can figure out a solution
However I didn't end up finding a solution and I'm planning to rewrite my project. I have a better idea how I should introduce it to people next time
Also, early on, I found the Referrer tracking to be useful for discovering the reach of my projects and to help me get into conversations with users on other sites to help them with using my software. But that feature of GA eventually became useless when Google did nothing to address Referrer spam.
There's one feature that 100% of our subscribers use. I mean, it's the main thing we do and everyone who subscribes needs that function. Even without analytics, we know that we have to continually make that feature faster and smarter to stay ahead of our competition. We know that if we fall behind our competitors, our subscriptions will dwindle.
But, then we have a bunch of other features that help customers solve some edge case problems around that main thing. It's very important for us to know which of those functions are being tried, used, and reused (or not). Not all customers use these features, but some are tools we have that nobody else in our space does. So, having insights on which ones are getting traction and which ones need improvement help us spend our marketing and engineering time better to attract new customers and retain our existing ones.
I can't imagine not having any analytics. I feel like we need them to continually make small course corrections that ensure we're providing the best value to our existing customers.
> It's very important for us to know which of those functions are being tried, used, and reused (or not)
Aren't these 2 at odds with each other, or else how can you tell when the same person re-uses a feature? Surely you need some kind of user identifier for that?
Additionally, unlike google analytics, we do support the browser's "do not track" flag. So if a user doesn't want to be tracked at all, we completely respect that.
Removing it felt momentous and insane. But in November 2018 I finally plucked up the courage and removed it. The crazy thing is, until this article appeared on the top of Hacker News reminded me, I had completely forgotten that I had removed it. Far from the world ending, it turned out to be the most inconsequential thing imaginable.
(I remember pouring over web server logs in Analog and AWStats 15+ years ago. Now I honestly can't remember why. I think it was some combination of vanity... and because everyone else was doing it. I suspect for most web developers GA was just the natural evolution of that muscle memory.)
I'm self-employed, so I have no boss or shareholders that need pretty reports with bar charts. In my case my site is deeply database driven and I can build engagement statistics directly from real data using complex SQL queries.
And while there's only a few such 'reports' that I check regularly, most of them are temporally incongruous—I think that's how you'd describe it—in that they look at what happened in the past contextualised by what's known in the present. (E.g. tracking engagements from new/irregular users, while they were new/irregular users, but which subsequently became regular users.)
For us, we have generated a lot of revenue by measuring what works and what doesn’t. That’s why analytics are worth it for a lot of people.
Just wonder, what's awful about AWStats?
Sure it's dated, and "analog" in a way that it's log-based, not JS. But it does not send the tracking to the third party, can be used offline.
No tracking. Privacy focused. Lightweight. You embed from your own domain. They even do site monitoring now!
not a knock on you or fathom, but it seems like you're in fast-response sales mode here (which is totally fine)... the above is a particularly empty statement. why would any next release not be the best to date?
maybe say it should be an exciting release, which is similarly anticipatory without being meaningless sales-speak.
We build privacy software so it felt slightly hypocritical to use a privacy-intrusive service like GA. So far so good.
I went from 0 to Fathom in under 20 mins and for our basic requirements it works really well .
Good job Fathom team :)
> Our on-demand, auto-scaling servers will never slow your site down. Our tracker file is served via our super-fast CDN, with endpoints located around the world to ensure fast page loads.
This suggests that this solution is not self hosted. Is there a solution like this which is really self hosted? This service is one small change away from actually tracking.
Edit: Piwik/Matomo[1] appears to be the most mature one. [1]: https://matomo.org/
I think the main differences compared to Matomo is that it's simpler (less features), but provides for much cheaper some of their premium features (heatmaps, session recordings).
Let me know if you have any questions about userTrack or any suggestions! :)
The open source project is barely maintained at this point- they update the readme and get the occasionally pull request, but it's not really being developed.
I unfortunately switched to Fathom back when they were telling people they were committed to open source, so now I'm looking to migrate off to something a bit more trustworthy.
I wish they'd offer more plans between the first 2 cheapest, though. My open-source project is hitting the basic plan limits and the next offer is too expensive for me.
Personally, I think that Fathom strikes a good balance between privacy and usability, but it does still use tracking (or at least it did when I was looking at it a few weeks back) - the difference is that it uses fingerprinting instead of cookies. I think it's implemented in a privacy-focused way, but it does look like they are ignoring some of the EU ePrivacy guidance, which explicitly states that consent should be obtained before using fingerprinting, even if PII can't be reverse-engineered from the fingerprint.
As I say, I think their implementation makes a lot of sense, and even as a privacy advocate myself I think those particular pieces of ePrivacy guidance focused on fingerprinting is excessive. But the EU doesn't seem to agree.
You'll know this but some people reading might not: Under GDPR, there are multiple legal bases for processing and we rely on legitimate interest. PECR / ePrivacy is the grey area for us and other services.
Having said all of this, we're fortunately moving away from requiring any compliance at all... by avoiding the complexities all together. We're rolling a refactor to our data collector over the next few weeks, and we won't have to have these conversations about grey areas anymore :) We've hired a top-tier privacy consultant and are going to be deploying a huge update, putting us at the top of the list for compliant analytics. Every single privacy-focused analytics service is in a grey area right now (some think they're not but they are). We will be the first to move out of this GDPR / ePrivacy grey area dance.
As you say, you see the logic behind the implementation we had, but we're dealing with politicians who don't understand the difference between Google Analytics and privacy-focused analytics. And that's fine, the work they've done has lead to better privacy for everyone, so we appreciate them.
That sounds like you are trying to pick and choose the bits you want to hear :)
There have been several ammendments since the original ePrivacy guidance. There is at least one such directive that is very explicit about fingerprinting specifically. If doesn't use ambiguous language, it states clearly that consent is required for fingerprinting.
As I said, I personally think it's just bonkers, and I think your service is absolutely in the spirit of the ePrivacy rules. But you can't say the rules on fingerprinting are not clear.
I'm keen to see what you've got coming, as the only way I see to avoid consent is not to associate identifiers with users at all - so each page hit would be a completely independent object. Can you say anything about your plans here?
And we definitely agree that it's bonkers.
I can't say anything here until we've got our press release out.
Ouch, I kind of wish you hadn't said that, because it sounds like you're straying dangerously close into weasel words and deliberately incorrectly interpretations. Sorry if that sounds harsh, but what I've read is very clear.
As before I like your solution, and I think it's absolutely in the spirit of privacy. But the guidance is really clear here, and gives examples of fingerprinting. Nobody said a fingerprint has to be a permanent identifier; as far as I recall, Fathom does use fingerprinting to identify individuals, so that a sequence of page views can be attributed to a single visitor. I understand that those fingerprints include a timestamp, and so are only valid for some time (2 hours, or whatever it is).
I get it that I’m asking for a free service, I just kind of wish they never offered it if they were going to ditch it. I don’t make money off my sites. I wish they had a less than X income version to self host. Oh well.
If you are making money and willing to pay I have a feeling Fathom is great.
If you want, you can contact me on Twitter to tell me what the free version should contain for you to consider using it, or if the current price for the lifetime version is too high for you.
Can't recommend a company which pulls crap like that. :(
Here's a comparison I've always looked at. Here are 2 solid database tools:
Table Plus - https://tableplus.com/ SQLite Browser - https://sqlitebrowser.org/
Table Plus is about 3 years old, SQLite is 6 years old. I have used both but Table Plus is 10000x better and more popular. Everyone talks about it and they innovate fast. They work full time on it too, $59 a license. They are so active, and have so many great features. Full time salaries help with that. SQLite has $102 in Patreon support. Table Plus makes that in 2 sales.
Look at Sequel Pro. One of the best products in the game but it never built a sustainable business model and it wasn't economically viable to continue, and look at it now.
I'm happy to say that we're achieving our goal every day. Our goal is to build a sustainable, profitable business (profitable businesses don't typically close down) that we can we can work full time on every day, moving people away from Google Analytics. For every 1 negative comment we get, hundreds / thousands of people signing up for Fathom.
We've spoken about it here too: https://usefathom.com/podcast/opensource
Who gives a crap? You guys started out as OSS and said it would continue to be. That turned out to be bullcrap, just looking for exposure.
> I have used both but Table Plus is 10000x better and more popular.
sigh Table Plus supports many different types of database. If people are looking for a macOS specific solution that does that, Table Plus might be a decent choice.
"SQLite Browser" on the other hand is specific to a single database. Though is cross platform and fairly popular in it's niche.
That being said, where are you getting "more popular" for Table Plus from though? Measured by what things?
> SQLite is 6 years old
Huh?
> SQLite has $102 in Patreon support.
You seem to be meaning DB Browser for SQLite.
Do we seem like a commercial project to you?
We recently added a Patreon link to our download page, after forgetting to put it out there much. The Patreon was originally created to cover our server costs (about US$75/mo from memory) some years ago, a goal we reached in 3 days.
That being said, we're trying out various things for a SaaS model as well because - like yourselves - we recognise success there will help with sustainability.
But we're sure as hell not going to bait and switch anyone.
> Look at Sequel Pro. ... and look at it now.
Seems like a commercial product whose organisation went broke, released the software as OSS, and couldn't get traction for that either?
People do appear to have picked up the project again recently, in an attempt to move it forwards:
https://github.com/sequelpro/sequelpro/issues/3679#issuecomm...
> Our goal is to...
Sure. You guys started out by bullshitting people. Unfortunately, the lack of integrity doesn't seem to be directly hurting. :(
One can hope it'll mean missed important opportunities for yourselves later, but who knows. You seem to be really leaning into the sales thing here on HN though.
> We've spoken about it here too ...
With an established history of lying, who cares?
Other places (eg Redash) do so, and are doing extremely well.
The change from FOSS to non-FOSS, especially throwing away the goodwill you'd started out with... makes no sense. :(
Popularity of DB Browser for SQLite wise, you might want to check our stats page:
https://sqlitebrowser.org/stats/
We're doing decently well, and our GitHub organisation has more active developers now than ever before. :)
In the GitHub issue you mentioned "since people are confused [..]", but people are confused because your communication has quite frankly not been very good. Even today just looking at the Fathom README it's not at all clear that Fathom v2 (or "PRO") is an entirely different codebase/product which merely shares the name and branding, and not much else, and it's not clear that the "Lite" version is essentially unmaintained either.
Is that info about the original author being unhappy somewhere public?
Plausible is pretty good, found it useful to monitor traffic and usage for small projects.
We are very firm on our values. We will never sell your data. We have many ways to get your raw data out of our system (API, download links, ...).
Our collection script [2] is open source and today we are also adding source maps to our public scripts. Open source does not guarantee that a business runs that same software as their cloud based option. We are looking into services that can validate what we collect on our servers. We never collect any IPs of personal data [3].
Great to see more products that care about privacy, I hope they will really care and commit to their values for a long time.
[1] https://simpleanalytics.com
Around 1999/2000 there was a rise of ISPs needing to install reverse proxy caches because the growth of consumer access meant they were getting seriously contended on upstream access. I was working at the time at a UK 0845 white label ISP called Telinco (was behind Connect Free, Totalise, Current Bun and other 0845 ISPs), and to my knowledge we were the first in the UK to install a Netapps cache. It was the moment we realised (by checking the logs to see if it was working), just how much porn our customers were accessing.
Those caches blow server side analytics to pieces, because frequently you wouldn't even know the user had hit the page. What server side analytics was useful for is what we'd now call Observability: they gave reasonable Latency, Error Rate and Throughput metrics, which combined with some other system logs might also give you a sense of Saturation.
As such, they were not too useful for marketing. Google Analytics was the first product that allowed high fidelity analytics even if reverse proxy caches (and even browser caches), were all over the place.
And here we are. In a World where we are tightly surveilled by corporate entities in order to try and get us to click on things. Bit sad really.
I'd encourage people to think about what they need these analytics for.
If it's marketing, you might just as well using GA: it's the best product out there. We just need to lobby for better regulation (at least GDPR and cookie setting popovers give us choices on that regard now).
If you're stroking your ego, consider whether such an invasive technology is worth the price, and if you need those numbers.
If you're making sure your infrastructure can handle the traffic, use server side analytics alone. Parse your logs using the huge number of tools out there able to do that in near-realtime, and leave your users' browsers free of tracking cookies and javascript.
To ensure your origin server is hit on subsequent HTTPS requests, you would still need to configure response headers for cache validation [1] to be
Cache-Control: no-cache
instead of Cache-Control: private
[1]: https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Ca...You either ship all your logs to one place (and hope that place doesn't go offline) or ship your logs to multiple places and hope both destinations are in sync. We've opted for #2 right now (hint: it's not perfect) but it's made me think about writing an alternative.
Rather than shipping all the logs all around, my plan is to have each source (i.e. web server) run a process on it's own logs, and use something like Redis to store the aggregated statistics.
https://joinup.ec.europa.eu/collection/eupl/introduction-eup...
OSI-certified, copyleft, non-viral, GPL-compatible, SaaS-aware, multilingual
The short answer is that there are certain protections that ensure interoperability, and that linking to software does not make it a derivative work.
That directive is usually understood to be about reverse engineering in order to build compatible software: "to obtain the necessary information to achieve the interoperability of an independently created program with other programs" being a key bit.
It's not immediately clear to me - a programmer but not a lawyer - that this has any bearing on whether linking creates a derivative work.
Have any other experts, or courts, weighed in on whether this analysis is sound?
I've thought about this for quite some time, and decided I'll use a slightly modified version of the EUPL which removes GPL from the compatible license appendix. Just haven't gotten around to that for no reason in particular.
The only case when you'd get better analytics from a _service_ is exactly a GA-like setup that can track people as they go from one website to another. That is, the real value of an analytics service is derived directly from its ability to invade people privacy, at scale.
Granted, migrating to another service is usually simpler, but it offers NO insights into the traffic that you can't get from parsing server logs and in-page pingbacks. You do however get a 3rd party dependency and a subscription fee.
I was once making a service that provided cross site widgets for companies to embed. Obviously it was beneficial to track people as they go from one website to another, but at that point it was beneficial to do it with our own service.
For example, if you validate forms with JS you might want to track form submissions and validation errors.
My personal domain[0] was taken by domain squatters (forgotten bill in debit card shuffle, bought up within seconds of expire) so for now I have to host on github.io. Thoughts on an analytics service?
Of course you can run your own analytics on AWS or similar and have no issues with handling traffic, but that means higher costs / difficulty in setting up and maintaining it.
Note that even in Google Analytics, this requires extra set-up, has limitations, and tends to be pretty fragile in practice. GA identifies users by a first-party cookie and tracking cross-site visits requires decorating links with cookie values.
If you're interested just in aggregate traffic from one of your sites to another, rather than something that requires full-path analysis (like marketing attribute), then you can get that from looking at referrers. This should be more-or-less equally available in GA and server logs.
It also had some weird quirks like generating duplicate entries or randomly failing to parse some log lines (you seem to have quite a few of those "failed requests" yourelf by the way).
There also doesn't seem to be a good way to display statistics for multiple virtual hosts. Even if you change your log format to include the host, you just get an additional table in the dashboard, but still can't look at the other metrics for each host separately. You'd have to run multiple GoAccess instances to achieve that.
I figured it’s fine for my needs since I literally have nothing on my domains. I could see it being frustrating for power users.
"Analytics" is rarely useful or unuseful because of the tool. These tools need to be treated as data collection, not reporting.
If your goal is to inform certain decisions, track success or identify problems... a spreadsheet (or napkin) is usually where that happens.
Say you do analysis systematically, make a list of questions and use your tools to answer them... usually you find that the tool itself doesn't matter much, and GA doesn't answer most of your questions out-of-the-box anyway.
Say you want a "funnel." That usually consists of a handful of data points. GA usually doesn't have them by default, without tinkering configuration, etc. Decide what they are beforehand. Understand them. Use GA (or whatever) to get the data.
Finding the tool for the job is much easier once you know what the job is. GA is extremely noisy, bombarding users with half-accurate, half-understood reports.
Edit: One major thing I am unhappy with in Matomo is event tracking. GA makes it much easier (in my experience) to track conversions and events, and presents the data in a better way.
My idea was to focus everything on "segments". So for all the data you can quickly create user segments and instantly filter the data to see only stats for the users you want, or compare stats between segments.
There is a public dashboard that you can check, I would love some feedback if you have the time :)
I still get monthly emails from Google about the analytics for this website. Apparently it's getting 200-300 visitors per month still. I have replied back to Google vie email about this several times but never heard any reply. I wonder what site they are tracking?
Filtering spam and getting useful data on GA is a never ending job that Google keeps making harder. (re removal of Service Provider / Network Domain [1])
[1]: https://support.google.com/analytics/thread/27808046?hl=en
[0]: https://funnybretzel.com/self-hosted-analytics-using-sqlite-...
I'm a fan of GoAccess. Unfortunately the queries are pre-made and there are nearly no options whatsoever. You can't (yet) filter by date for example.
One thing I realized is, that on small sites, like my blog, an overwhelming amount of traffic comes from search engines or bots which are looking for vulnerabilities. Filtering them out takes a lot of time in any self-hosted or self-made solution.
If not, then logs of webservers are the only 100% reliable place (if available of course), so old-style tools like awstats, Webalizer, etc [1] should have a rise in popularity again.
[1] https://en.wikipedia.org/wiki/List_of_web_analytics_software
In this case, the source of truth is cloudflare's loadbalancer, but you have to pay them to get full analytics.
For us nerds, we can always block things using DNS level blocking but Fathom’s custom domain feature has done really well for the majority: https://usefathom.com/blog/custom-domains-embed-code
Only completely self-hosted (your domain, your tracking server) solutions are resilient to adblocking
https://webkit.org/blog/8943/privacy-preserving-ad-click-att...
(Full disclosure: I work for Snowplow Analytics)
Is there a highly-opinionated tutorial that shows how one can get some vanity metrics out from Snowplow?
Other products:
- Objectively lack features
- Potentially incur extra costs in money/time
- May be a small barrier in m&a
- May carry additional risks/attack vectors if self hosted
Trying to ween off big tech is commendable, but likely detrimental to a business.
Relatively high risk, low reward.
I'm happy to have my mind changed. I can see a case for user hostility, but most sites I imagine don't have an audience sensitive to this at the moment anyway.
From an idealogical standpoint, other cloud stat tracking services would only function if not many people used them. And I would also imagine feature creep would be inevitable and lead them to becoming an inferior version of GA.
GDPR is the European privacy law. It protects European citizens so it applies not only to European companies but any company that does business in Europe (having offices or advertising/selling there).
Google does not give much assurance regarding their GDPR compliance... their text on that subject is mostly CYA and then they make it your responsibility to decide how to use it in compliance (if at all possible).
The GDPR gives you a small window to count visitors through cookies as long as all private information (even IP) is anonymized... OR you can go do a more traditional tracking with their explicit agreement. This last use case is completely useless in terms of visitor statistics, but analytics companies sometimes dare suggest it (as in "this is the way to do things right... so our product is compliant and it's not our responsibility if you break the law").
That aside, I run international non-profit sites and GA is a bad look... and with good reason: Using social network sharing buttons, GA, CDNs, etc. gives too much power to track people to a few companies.
Totally agree, but are there "acceptable" CDNs, like unpkg? What about Google Fonts?
> The GDPR gives you a small window to count visitors through cookies as long as all private information (even IP) is anonymized.
If it's not too much to ask, could you expand on that a bit, or share a link? I guess you mean it's ok to send a browser fingerprint for unique visitor stats without having to ask for permission, but I'm not aware of any legal debate let alone court decision with respect to that.
Edit: obviously I can't read ("through cookies"), but cookies for unique visitor counts aren't "functional" are they, so my interpretation is that those cookies need consent; I'd love to hear otherwise though
The GDPR does not discriminate based on citizenship.
It applies if the organisation providing the service is in the EU/EEA* OR if the user of the service is in the EU/EEA* (to the extent that the data reference their activity in the EU/EEA).
And the UK, but thanks to brexit there's a parallel UK GDPR in place so... take that in to account.
I do think that for the average user, using GA might be fine because it's free, easy to set-up and does its job. That is unless they care about all the possible consequences.
I agree for many features are still lacking, but as a counter-argument 1) not everyone needs those features (not every product needs to solve 100% of the use cases), and 2) a lot of these products are still quite new, and are actively working on adding a number of those features.
It is self hosted, has support for desktop apps, mobile apps and web apps at the same time.
[1] https://count.ly
[2] https://marketplace.digitalocean.com/apps/countly-analytics
> Disable SELinux on Red Hat or CentOS if it has been enabled. Countly may not work on a server where SELinux is enabled. In order to disable SELinux, run "setenforce 0".
https://support.count.ly/hc/en-us/articles/360036862332-Inst...
It's pretty simple to extend too: I've added basic client-side (JavaScript) error reporting on top of it, and I'm thinking about using it for Content Security Policy reporting too.
I am wondering if HN is interested in hosting analytics like plausible that is open for us to see. Sometimes I do wonder how many page view do HN get per day, where are we all from etc. For example the plausible demo site. 35% are using macOS. But only 15% uses Safari.
For those interested, one other FOSS analytics tool is Shynet [0]. Modern, privacy-friendly, and detailed web analytics that works without cookies or JS. It also looks pretty slick. Disclosure: I’m a maintainer.
I would have though that there would be several decent packages offering www + app analytics by now, but as I wrote, options were quite limited. Some of the options mentioned in the subject here looks like good options for just website analytics, but I'm not seeing much as far as "app analytics" (custom events) goes.
Companies focusing on user level tracking today provide a different set of tools one might be used to and that can of course be compliant, see https://www.hotjar.com/.
I went with https://app.usefathom.com which tracks _aggregate anonymized_ data.
They have the option to self host, but I'm sending them money to support the project. With today's launch, I'm really happy with the product. Will continue using it.
I tried bringing together the most useful analytics features (user segments, heatmaps, session recordings, tags/events) in a self-hosted platform with simple UI. A/B testing feature is also coming soon. I built the platform with the optimal use-case being improving conversion rates on landing pages.
My goal now is to prove and teach (even to non-technical users) that self-hosting is easy nowadays when you can create a VPS running your desired software in just a few clicks.
I would love to hear some criticism or why you wouldn't want to try something like this.
Fathom analytics and simple analytics cost ~100$/year.
Plausible costs ~50$
I really liked and almost settled with plausible but I just saw goatcounter right now. It's free for personal / open source projects. That's so nice for small projects like many people here are building.
https://github.com/usefathom/fathom https://github.com/plausible/analytics
Here is my research: https://til.marcuse.info/webmaster/alt-analytics.html
I ended up going with GoatCounter.
EU be like you have to put up this banner to tell them you are spying... it super annoying and everyone hates it... and you be like "sure, I love me some spying".
I suppose the purpose of why certain collections of data are put together is where all the anger rightfully comes from, but calling for a complete armistice on all (even innocuous) forms of data collection is a little churlish.
You do it in microcosm too, you know. How many pictures do you have on your phone, that contain people you don't know?
Spying is the only solution for spying, yes. Spying is the only solution to over-eager people wanting mostly useless metrics.
It's not much different than me checking my Twitter "likes" constantly.
I was surprised matomo wasn't listed here. Does anyone know if that was intentional? Seems like it fits the criteria of the post and the goals of open source.
Yes, this was intentional as this article was focused on "light-weight" analytics. I think they'll do an article about Matomo at some point in the future as well.
I'm putting the finishing touches on a `tag=XXX` parameter that allows you to record a tag (like a pageid), and then filter the maps by it (not publicly documented yet, but will be in the next couple weeks).
Doing so is extremely helpful to understanding what events and data you actually need for your use case.
Don't use analytics. You really don't need it. No. You really don't. No, No. I promise you. Just stop.
All tracking is evil. All ads (except those inside a store for a product inside the same store) are evil.
Having a small static blog hosted on GithubPages, GA was the only option for me. (Not going to pay for analytics while my blog has like, 10 visits a week)
For work though, we use GA and I can’t really imagine switching. We actually use the event stuff, so it would be hard to switch away.
Detecting and deterring spam and sketchy behaviour while using open source software could be an interesting technical problem area.
I get the point for any commercial venture but for most sides I simply don't see the added value. If you want to know whether people like your work add a comment section or a newsletter signup - why do you need to spy on your users with intrusive tools, send their data around the globe to kraken like google, just to have a few statistics?
I assume many people have sites for their small businesses. Imagine you have a restaurant, you want to be able to know how many people reach your site, how they reach it, why they don't contact you (do they leave after seeing the menu? or the opening times? or the location? or photos?), are your contact forms properly working, is your website loading fast enough, etc.
It uses your server logs for simple analytics.
Ive changed to Goat Counter, but thanks
There's no UI for my signals application, but it should give me raw access to all the GA metrics people typically look for (pages, referrers, user agents, etc). Storage could be pain point. Compute might also turn out to be, but if that ever becomes a problem, then I'll probably have much more to think about...
What am I missing here?
Then you have redundancy of data. How are they backing up historical data? Are they running with failovers? How will their analytics do in the event that they get a hug of death from Hacker News or Reddit? There are so many factors to consider.
I don't think we should discourage people from rolling their own if they enjoy it. Heck, I've built things that existed. That's how we get better. I'm just sharing why a lot of developers won't roll their own.
I'm collecting feedback and looking to open it for beta in the next couple of weeks, but there's already a preview signup link in the newsletter, where I share my progress on building the platform on a weekly basis.
Have GCP? Hit this button and you get your docker container in your Kubernetes cluster doing all the stuff for you, pretty awesome.
The Firebase free tier seemed perfect for my use case. It is far from being perfect, but Good Enough For Me™
I used Matomo before, but simple dashboard in style of plausible.io is more useful for me. I have a little traffic on my sites.
Edit:
Also it doesn't seem to be tracking referrers correctly.
I moved to Netlify analytics from Goat counter because I liked the idea of having server side analytics, but at $9 per month these are extremely overpriced.
I think I will just go back to Goat counter.
It's not that I really need statistics, but sometimes it would be nice to know if there is even anything going on or you're just screaming into a void.
But I refuse to spam visitors of my pages with GA.
https://github.blog/2014-01-07-introducing-github-traffic-an...
We have lots of agencies who use us, and their clients may differ from yours, but here’s what’s helped them: https://usefathom.com/blog/switch
Just be aware that, while that is true for HN, and some subreddits, that still doesn't apply to most people. It's the same as saying, based on comments on HN, that people are moving away from Chrome when its market share is not going down.
GA also counts - social interaction, events, Ecommerce and many more.
I wonder if it's easy to just send events to something like Firebase and then query that DB.
A graphql api or not doesn't really matter since that's just an abstraction for querying a service or database.
I really mean something like an open source Amplitude or Keen where you can dump events into it and then aggregate/group them in any way you like on demand.
https://usefathom.com/blog/google-analytics-seo
Funny as it sounds, using Fathom instead of GA could increase your SEO rankings :)
I've used it in my personal projects and have never had any issues. It's great to have no vendor lock-in and full ownership of user metric data. If you go this route, please be responsible with the data and follow all relevant regulations & guidelines (ex: GDPR) regarding its storage and usage.
Why?
Because there are several open source projects that if joined up in a relatively simple way would provide a full RUM / Analytics solution.
For collecting the analytics: Akamai Boomerang, which is a descendant of Yahoo tooling https://github.com/akamai/boomerang and BSD licensed
Then insert a collector that will produce Prometheus metrics and write log lines. The metrics will provide the timing information and in many ways will be richer than Google Analytics, and the log lines will provide the potentially high cardinality of string based data like user agents, URIs, etc and Grafana supports PromQL queries against logs such that you can gain metrics from the log lines too.
Then add in a free Grafana Cloud https://grafana.com/products/cloud/ and configure Prometheus to scrape from the collector, and Loki to consume the logs.
This is an end-to-end cloud hosted RUM / Analytics solution that for a single user would be free and one can even add alerting.
The missing link is that collector, to consume the Boomerang output and produce Prometheus metrics and log lines for Loki.
All of this is open source and can be self hosted, the only piece you would have to host today would be that custom collector to receive the Boomerang requests.
Check out a comprehensive comparison of GA, GA360 & Piwik PRO at: https://piwik.pro/blog/piwik-pro-vs-google-analytics-compreh...
So I disabled it and went back to awstats. I've been using awstats for over a decade, and for my personal site and projects, it pretty much gives me the majority of the data I really care about.
I might look at shipping more complex nginx json logs to logstash/elastic search, but then I'd need to visualize them in Kibana and that just seems like a lot of heavy weight containers to run for stats I don't really need.
We’re fully bootstrapped, actively rejecting millions of dollars in venture capital, and we are sustainable. That is the key. We are priced fairly, and at a level that allows us to ensure the longevity of our business.
We are used by governments, small businesses, multi billion dollar companies and individuals. Everyone cares about privacy and legal teams love us.
We are fully GDPR compliant and use zero cookies. A lot of people have read our article on cookie-free tracking, but that article is outdated now. We’re also launching a new method over the next few weeks which is game changing, which we’ll blog about.
Our infrastructure is highly available and runs across multiple servers. We don’t run our services from a single VPS, and have availability in multiple availability zones for everything. We pay premiums for our infrastructure because we take our customers data very seriously.
We allow you to set-up a custom domain in less than 2 minutes, comfortably passing ad-blockers. Or if that’s not your cup of tea, you can enable honor-DNT and respect ad-blockers.
We are built to handle billions of pageviews a month, we’ve poured hundreds (thousands?) of hours into refining our aggregation script, and we’re the leading solution on the market.
Don’t forget, we also offer unlimited uptime monitoring as part of your plan, sending alerts by SMS, Telegram, email and Slack.
Finally, we run a popular podcast called Above Board, where we talk about business & privacy.
If you haven’t already checked us out, you should.
From the GitHub repo [1]:
> At present, Fathom Analytics Lite is not PECR compliant due to the fact that it uses an anonymous cookie. Our PRO version is PECR compliant, and we'll be making changes to this codebase some time in the future to make it compliant.
The open source version seems to be lacking behind and might generally be treated without much love, given that their website doesn't even link to it (which is understandable from a business point of view, but doesn't inspire much confidence in its future maintenance).
[1]: https://github.com/usefathom/fathom/blob/69baac5c4a4d96880a2...
When we focused on building a business, we were able to become a much more viable competitor to Google Analytics.
The reason we haven't written off Fathom Lite is because we've always had plans to come back to it this year and put out an update. Will we be launching new features? No. Will we be fixing bugs and ensuring it's a solid product for individuals? Absolutely.