Matomo: Open-source analytics platform
github.com
github.com
Only thing that bothered me is that most Ad Blockers are blocking Matomo as well. I did build a little Script to circumvent that, you might find it handy as well: https://gumroad.com/l/matomo_circumvent_adblock
I use it on my website. Check if your ad blocker is capable of blocking it: https://simon-frey.com
I think we can all agree there are different levels of acceptable tracking and use of that data- but the degrees of acceptance are going to be different depending on the user and service. I don't consider bypassing my restrictions to run unauthorized code to be an acceptable tracking method and raises serious concerns about how the data will then be used.
Now, I have a vested interest in this as I work on one of those tracking tools, but it actually collects less data than those Apache access_logs that people have been keeping for 25 years. Plus, the JS is unminified and easily examinable if you want (as is the HTTP request), so you also have more insight in what is being collected exactly.
"It's using JavaScript" and "it can do [..]" are massive red herrings; browsers are actually fairly sandboxed and there are millions upon millions of lines of code on your computer that can do much more than JavaScript inside a webpage.
Yes, and then you would be charged with assault. It is great that you work on a tool that respects peoples privacy. I suppose I failed to put an emphasis on trust. With server side logs, less trust is required because there is less that can be done. Paired with VPN, I can have reasonable belief that server side logging is not logging anything unreasonable and it does not require trust that they are not fingerprinting me. As you say, just because someone can do something doesn't mean they will - but trust is required, especially if there are no repercussions if that trust is violated.
Ah, right, creators want their content to show up for my search keywords, Google won't let them have pages only visible to Google bots (though even that is changing with the rise of paywalled sites), and they want the money from that same Google showing ads from their ad network.
Google initially promised to deliver a search for the open web unencumbered. It has become a sort of paywall itself (accept our ads or our search results will be useless pointing you to pages that only work if you have ads enabled).
Sure, it would be fair if they haven't pushed out the competition acting entirely differently ("we have no ads", "our ads are clearly marked" to current "see if you can tell a difference between an ad and your search results").
I go as far as to send all the tracking parameters through a custom server script before they are proxied to GA and Matomo. That way, I can change the script and parameter names at will, making them much more difficult to block. For example, Matomo-related blocking rules are as follows:
/matomo-tracking.
/matomo.js$domain=~github.com
/matomo.php
/matomo/$domain=~github.com|~matomo.org|~wordpress.org
/piwik-$domain=~github.com|~matomo.org|~piwik.org|~piwik.pro|piwikpro.de
/piwik.$image,script,domain=~matomo.org|~piwik.org|~piwik.pro|piwikpro.de
/piwik./ping?
/piwik.js
/piwik.php
/piwik/*$domain=~github.com|~matomo.org|~piwik.org|~piwik.pro
/piwik1.
/piwik2.js
/piwik_
/piwikapi.js
/piwikC_
/piwikTracker.
I experimented with tracking on the same site and the overhead is not worth it for me. Central solution for all my projects works quite reliable
Any tracker can be made to work around ad blockers by making callbacks to the site itself and having a small shim there that forwards these pingbacks to the actual tracking service. But even then they still can be blocked based on the request contents.
PS. Here's how your website looks in Firefox - https://i.imgur.com/uFKEB4X.jpg. That's with uBlock off. No console errors.
You purchase a support license to help me to continue working on MCAB. MCAB itself is Open Source and can be found on Github: https://github.com/simonfrey/matomo_circumvent_adblock
It's kind of my job to take random apps, deploy them and manage properly, so I'm not a clueless user here. I could press on and figure it out with more time, and I understand they'd be happy with people using the cloud offering / paid support instead. But I also feel like a working docker-compose (or comparable) setup is table stakes these days for an open-source service.
See Loki+Grafana for a good example: https://grafana.com/docs/loki/latest/installation/docker/#in... - it's not a production setup, but it's a valid "play around with it in 2min" setup.
version: '3'
services:
mysql:
image: mysql:8
environment:
MYSQL_ROOT_PASSWORD: root
MYSQL_DATABASE: matomo
matomo:
image: matomo:4
ports:
- 4000:80
This lets me in at localhost:4000 and I just enter "mysql" as the DB host, "root" as the username & password and "matomo" as the database name, and it's basically done.Of course, I probably have to point it out or someone else will, that it's a bad idea to be using the MySQL root user, instead of creating a user with the rights that Matomo needs: https://matomo.org/faq/how-to-install/faq_23484/
A lot of the challenges faced with a ‘from scratch’ install will revolve around which PHP version and extensions to install and how to get Nginx to talk to FPM. Neither of which are trivial for someone wanting to test/evaluate without much prior knowledge.
If we're talking about 10s of thousands, you're gonna need to invest in some SSD (€50-70/month probably).
If we're talking about dozens of sites, some of which have millions of yearly visitors (and a bunch of plugins and reports that need to be generated), then you're gonna encounter some issues and have to spend a considerable amount of time optimizing every part of it, and hardware cost will rise to a hundred or two per month.
I guess it depends on how essential your tracking is, and how you've implemented it. It shouldn't be added in a way that can take out your site unless there is some business critical reason to track.
Then I'd ask, just how critical is the tracking? If losing a few hours of data is going to throw off your product development, do you have enough data to be making decisions? My experience is that bugs and misconfiguration of experiments is common in most orgs, so even if the system is up to capture all data, product managers check an experiment a week later to find they have only 50% of the data.
You have a very good looking UI there. I really love the simplicity.
Yet, though at least it isn't cloud based, it's still quite scary what kinds of things it will tell you about your visitors.
I use it on a personal project site and it works very well.
The French data protection authority issued a piece of code (JS) which must be used to avoid collecting the user's consent. I don't know about other data protection authorities in the EU but it shouldn't be much different.
This is not how the GDPR works. If you are collecting personal data, or if you are dropping analytics cookies on someone's device, you need consent. No ifs or buts.
I never said no personal data were collected but, if configure properly, the processing of data falls within the legitimate interest basis.
https://github.com/LINCnil/Guide-RGPD-du-developpeur/commit/...
/edit Ignore me. I appear to have misunderstood the changes when I last read this. Sorry
The ePrivacy Directive requires consent for reading or writing from a terminal device. This includes anything with cookies, even if they're not personal data. While the ePD refers to GDPR for its definition of consent, it is a separate piece of legislation and many things that are true about GDPR are not true about ePD (such as being able to invoke Legitimate Interest instead of consent).
You're right to point to e-privacy, to which consent is central. But the latest draft of its new version states that (art.8): 1.The use of processing and storage capabilities of terminal equipment and the collection of information from end-users’ terminal equipment, including about its software and hardware, other than by the end-user concerned shall be prohibited, except on the following grounds: [...] (d)it is necessary for audience measuring, provided that such measurement is carried out by the provider of the information society service requested by the end-user or by a third party, or by third parties jointly,on behalf of theone or more providersof the information society service provided that conditions laid down in Article 28, or where applicable Article 26,of Regulation (EU) 2016/679 are met
So Matomo can still do without the user consent (from what I understand, the relation between GDPR and e-privacy is no easy business).
We are in agreement. It seems I wasn't clear enough in my original post, but this is my overall point. GDPR doesn't require consent, but consent is required because of ePD.
> latest draft of its new version
> So Matomo can still do without the user consent
The new draft is not law yet. It's been 6 months away from passing for several years now. In the meantime, fines are still being issued under the existing law. Google got fined a hundred million euro last month in France, and that fine was very specifically ePD and not GDPR for a variety fo reasons.
It also depends on the jurisdiction. For example the ICO has been clear that using a cookie based analytics tool requires a GDPR level of consent, without exceptions.
I'd have to re-read it to be sure about analytics cookies, but I don't think it says a whole lot about that off-hand. This the the ePrivacy directive.
matomo is very comparable to google analytics in terms of reports. matomo has some things that seem a little easier to get to; like visitor flows.
However, matomo seems to just give up on big data, complex reports. Similar reports in google analytics take a long time to complete, 10 to 40 seconds, but they at least complete; eventually.
Are these alternatives fully able to replace Google analytics?
I sort of thought Google analytics would tell you more about your visitors since with Google cookies, they could map them to other visited websites, centers of interest, age group, etc.
Are you loosing all that when switching to a less intrusive analytics platform such as this, or is Google analytics not leveraging their ability to disclose more about the visitors?
Matomo purely tracks analytics: who visited what page, for how long, from what device, from what location, from what inbound website, and what outbound links did they click. It also provides a log of pages requested per session so you can analyze people's flows through your website.
It's certainly not a replacement for Google Analytics if you use it to collect background information on your visitors. Even though Google's information is very broad (you mostly get ranges and the interests aren't that reliable), some marketeers use it to make decisions about their marketing strategies. Matomo won't help you there, your alternative would probably be Facebook or another big tech tracking solution.
It does provide a replacement for the type of tracking that I personally find acceptable, assuming the IP addresses are anonymized sufficiently. Matomo recommends shortening IP addresses to /16 after analysis, which I consider good enough, but that's a setting administrators can change.
That data is mostly garbage and only getting worse.
It's off by default. Turning it on gives you basic demographic data but also means you consent to sharing your GA data with Google to use for advertising purposes.
You do need to explicitly enable this in the GA dashboard, and ask users' consent under the GDPR.
[1] https://support.google.com/google-ads/answer/2580383?hl=en [2] https://support.google.com/google-ads/answer/2497941?hl=en
If you can replace GA depends on your needs. GA collects more personal data, you get better insight of your audience. This is important if you do online marketing and like to see how well your campaigns perform. GA does track visitors across days and you can therefore see if someone came back after a week and made a purchase.
In case you don't do that or are simply not interested in specifics, all the alternatives are good enough right now, I think. You can still tell how visitors navigate your page, what content they visit most and all that stuff. We are currently thinking about what we can add to gain more insight for businesses, without invading privacy as Google does.
Something you can't do when leaving google land.
$4/month if you pay annually or $6 to pay monthly, but free during beta.
We are actively working on it right now, but the core is working well and is open-source: https://github.com/pirsch-analytics/pirsch
Project: https://umami.is
I couldn't find any benchmarks. How many concurrent requests can it handle without errors, simple hello world.
Never mind I found a post. https://blog.rh-flow.de/2016/01/10/benchmark-helloworld-in-p...
As expected node is not a thing to handle large amounts of requests
Why would node be able to handle so much more HTTP requests than Apache or Nginx? I think the throughput is mostly dictated by implementation.
But, seems that at least in the TechEmpower framework benchmarks, es4x (JS) ends up on position 9 while the closest PHP framework ends up at 13. Now it's just a small benchmark with specific tests, but I do think it's easier to make NodeJS handle large amount of requests than PHP. Although again, you can definitely do large amounts of requests with PHP too. I've spent about 5 years on each, found that getting good performance out of V8 is easier than out of PHP.
PS: I have also been building something similar, but not completely open-source: https://www.usertrack.net
Once you purchase it you get full access to the original server-side code (PHP, MySQL).
For the client-side part you only get the bundled JS/HTML/CSS (the original client-side source code is TypeScript, React), mostly because otherwise I would have to provide all the build tools and document better the code, tooling, building, releasing, etc.
Open source has a specific meaning - is the Software released on an open source license. (https://opensource.org/licenses) For example if you pay enough, you get ms windows source as well - that doesn't make it "not completely open source". Your project doesn't seem to be open source at all.
I have seen many other products that are marketed as "open source" because you get the source code after you purchase it, so it is literally "open source", but not "open-source" as in released under an open-source license.
I am personally not marketing userTrak as open-source and I will stop using similar terms if other people do have a strong opinion about what "open-source" actually means.
You are NOT allowed to:
Redistribute in any way any of the userTrack files or any parts of the userTrack's source code (with the exception of the public tracker JavaScript files that have to be included on your site).
Install userTrack on someone else's server.
Continue using userTrack or offering userTrack access to others after this license agreement has been voided (either via a refund, license period expiration or legal action).
This is not open source (or even "fair code" as redis etc advocate for). Providing the source but under a license like this is usually referred to as visible source or shared source
The way userTrack is currently distributed is as any other digital product (you pay for it and you are not allowed to sell or redistribute copies of it) with the mention that the server-side code is un-compiled and un-obfuscated so you can transparently see what it does, how it does it and change it if you want.
I am not sure that fully open-sourcing it is the way to go as I've seen so many projects die or disappear because the maintainers didn't have a lot of incentives to keep improving it or simply no longer had time to work on it. I also think that it's fair to pay for something that brings value to you also knowing that by paying for it you support its further development.
Great product and an excellent demo!
I would love to make userTrack free if I can find a sustainable way to work on it. Most other similar open-source software offers a "hosted" version to get revenue, but my goal is to promote decentralization and self-hosting in general, so me focusing on the hosted version would go against my goal and beliefs. I really want to see a feature where any non-technical person can choose a few products and have them running on their own VPS/server in a few clicks. This would have many advantages for the clients AND for the developers:
* Clients pay a lot less for a products
* Developers must focus more on product and performance, leading to higher quality products.
* Hugely increased privacy for the average internet user and for the own data of the client using the product
* Better performance (each client has their own server so it is more likely to have more resources)
* Better latencies (each client can choose to use/host their product on a local datacenter)
* Better data transparency, easier migrations and fewer vendor lock-ins (if you own the server and the data on it you can most likely always export it in some form)
I think there are many other advantages for both companies and clients. The current SaaS environment makes it really easy for companies to ask huge amounts of money for services just because they want to, as the client has no real alternative unless he is really technical and can spend days installing and maintaining a self-hosted software that rarely gets updated.
Sorry if it seemed like I was complaining about the pricing change. I was just wondering whether I remembered it correctly from here (https://news.ycombinator.com/item?id=24207129)
It's a great product. People will pay for it.
Thank you for the kind words, I do love working on this project and I hope to be able to continue working on it. Existing customers absolutely love it and keep recommending but I am still struggling with finding a pricing structure that makes sense for everyone.
I do hope that one day I will find a way to make userTrack free for everyone, but looking at Matomo, making it open-source seems to drastically slow the development of a project as there are so many people involved and so much more decisions to be taken. Apart from that I would still have to earn a living somehow, but if I get a job and keep userTrack open-source I won't be able to spend too much energy on maintaining it and I hate not being able to make a product as good as it can be.
https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Re...
Anyhow that's as far as I got. The next and largest step would be analyzing those records aka creating all those beautiful graphs. Also creating a management interface for sites and generation of tracking code. While it's open source (search for gopiwik) it still uses the mgo.v2 driver and I just yesterday added go module support and did some minor changes.
I thought, why reinvent the wheel, let's just grab the piwik.js and build a receiver. Turns out it's a bit more complicated than that.
Nah.