Replace Google Analytics with a shell script
go350.com
go350.com
A ton of them still exist : things like webalizer, goaccess, and the venerable analog. See the self-hosted category here (https://en.m.wikipedia.org/wiki/List_of_web_analytics_softwa...)
Like the article, I recently realized that, while I did use it, I didn’t take advantage of google analytics at all so decided to nuke it from my sites. I had additionally been running webalizer for over a decade, but recently replaced it with goaccess (took about 10 minutes to set up) which is more actively maintained and have been quite happy with the basic metrics it produces. Sure it won’t stalk a user as they browse my site but I don’t care, just having an idea of how many hits I get, from which countries and which browsers is enough for my few personal and hobby sites.
I personally find it to still be too revealing in some cases, which is why tools like the nginx IPscrub module[0] should be the upstream default IMO.
Why not? Why wouldn't you know what people are doing on your website, within limits? Which section is getting most visits, which is the slowest, how many elements people click on in your lists, etc etc can have immense value in deciding what to do with your site, what needs work, what doesn't work at all, etc.
When a user downloads a web page from my server, I can temporarily record that an anonymous user made a request because I'm only logging what the software I'm running on my machine is doing. After that, the user should own that downloaded copy of the page and use it freely without surveillance from me or anyone else. The only restrictions in place are those set by the license (hopefully AGPLv3). Beyond those restrictions, the software doesn't belong to me anymore so I don't collect metrics.
It's true that I won't be able to get detailed metrics about what the user is doing, but it's also true that user's copy of my software will only be serving the user directly.
Web apps aren't different from other software; telemetry and user studies should be opt-in, transparent, and sent back only with informed consent. Only then can I consider running analytics.
Which also excludes real users, especially if the requirement is "allows JS from google analytics domains".
I always find it weird though that privacy advocates still go for "I have this hidden treasure trove of analytics about my site that you cannot access"
Analytics should always be 100% anonymous and aggregated. And if you've met those two bars, making them open and public should be a non-issue. If you can't make your analytics public, that's called spying. There should be no issue making your analytics public. Mine are[2]
I wonder how many extra dollars cumulatively DO makes from the $1 premium on getting a better Intel CPU or AMD CPU and/or NVMe up sell on the smallest droplets.
I pay the extra dollar as well. I doubt I actually see the difference for what I use it for. I feel no desire to pay a dollar less though, so they are making another 20% from me on that product than they would have.
I've had a lot of different droplets of different sizes over the last 6-7 years. Recently when I created a new one for a new project, I noticed that option and chose it. I'm not sure I would if buying a larger droplet, or a lot of small ones (I've done that for certain projects), but $1/mo is the point where it feels like nothing.
I think it's a little scummy that they go out of their way to hide it and make the more expensive options the default, though:
https://user-images.githubusercontent.com/3173176/115132649-...
DigitalOcean is going downhill fast after the layoffs, I'm not optimistic about them in the long term. I'd love to know what they burn on marketing, because their video ads on YouTube are crazy.
Eh, I wouldn't call that hiding it so much as making the current market default the one they default to. If you're looking for the lowest end droplet, then yeah, it makes sense you would want the lower end Intel without SSD, but if you're buying something more expensive at all, there's more of a choice to be made.
Also, I'm not going to fault them for putting forth the option that allows them to more easily decommission what are probably older hardware and disks for newer offerings, especially when the "hiding" is just a radio button above the droplets. It's not that hidden. I suspect eventually the newer gen Intel/AMD offerings will lower in price and replace the current default.
> I'd love to know what they burn on marketing, because their video ads on YouTube are crazy.
That's interesting. I remember seeing some like 6-12 months ago on youtube for the first time, but I haven't seen any for a long while. I must have just been flagged as not part of their target audience (or Google has determined I already use them).
It's my first website, and I think the total size of the backend is something like 40 LoC. I've never had any kind of analytics (or needed them), but I thought it would be neat to see what was going on. It took me all of 20 minutes to sign up, add the single Plausible <script></script> line to my index files and headers, and there it was. As someone with literally no other web experience, I can't imagine it having been easier. The cost of my website is now a whole $12/mo instead of $6, and I agree that it feels kinda nice to support a service that seems to place value on privacy and openness.
The idea that there should be no issue sharing them is absurd. Yeah just publish the minutes from your marketing strategy meetings as well while you’re at it. Invite your competitors to go over your books with you and your accountant and then arrange meet and greets with your most profitable customers as dessert.
In all seriousness, I'm OK with that as long as people are honest about it. Just say "We track our users to collect business intelligence." and I'm happy.
While I agree with the spirit of your comment, this is a very sweeping generalization. Many kinds of financial information are being shared regularly with the public and for publicly traded companies they're required by law. Many companies build their infrastructure on open source software, with proprietary glue being a small percent.
Knowing what you can or even should share (for marketing purposes) is crucial. Especially knowing that keeping information asymmetry forever is impossible.
This too is a sweeping generalization ;), a good one though. We're pretty fortunate to have a lot of our building blocks made using open source software. Boost libraries to openssl, most of the stack is open source including the programming language. Although, the proprietary glue is the important bit.
An architect's work is more than just bricks and mortar.
Basically whatever process you put in place to level the playing field, big businesses will find a way to game it to their advantage because they have the resources to do so and the deep pockets that allow them to take risks.
It's a way to level the playing field. The idea is that, if everyone has access to all the data, then there's more opportunity to make all business more efficient, whereas if some businesses hold information about the market and costumers for themselves, then they get to be sloppier in their implementation because they are already so far in the game.
Content marketing is part of the secret sauce these days.
Really? It’s spying now?
Entertainingly https://cpcchina.chinadaily.com.cn/ uses Google Analytics, while http://english.www.gov.cn/ (yes, no https) uses Alexa metrics. You can't even access the Chinese Communist party website without being "spied on" by US-based metrics companies. I don't think this proves anything but it's amusing to contemplate.
Instead of what solutions like Plausible are meant to do (only give a responsible well-defined party access to the data), now everyone can view your website users potentially PII.
Sure, someone could send out unique links to random `utm_campaign` IDs and see "did the person click it?" and I agree that's a bit nasty.
But even if you do convince someone to click _your_ link, what good does that do? We can see you yourself clicked that link from Germany on a mobile phone, and that's literally it. We can't get your IP address, or any PII - it's not stored.
You were able to verify what level of information my site is collecting *because* it is public. That is simply better than me spying and recording everything you do and you just having to "hope I never accidentally release it or sell it to someone" as is the case with modern tracking.
Yes, and _that's it_! Just by getting someone to click on a site they find trustworthy, you are now able to extract country, device, browser, OS, time of access, and to some degree other pages visited on the website. This in itself may not be PII, but together with previously known data points may be enough to assemble PII.
> Plausible literally provides a toggle to make it public in their UI: https://plausible.io/docs/visibility
Yeah, and I don't think that's a good idea that they do without any warning of the implications. If they make it easy to expose data publicly, there should at least be some measures to ensure anonymity like e.g. differential privacy based querying.
For a product that markets themselves as GDPR compliant and without cross-site tracking (which this technique enables to a certain degree), that's not really a great look.
> You were able to verify what level of information my site is collecting because it is public.
I mean kinda, but not really. I have no assurance the "all the data you are showing me" == "all the data you have about me". In the end I still have to trust you that you don't collect more data and that you won't sell that data to someone else.
-----
All in all, I really appreciate what you are trying to do here in terms of transparency, and I'd like to see that spread. However I think there is still a lot of room for tools to improve to provide a similar level of transparency while also insuring users privacy.
So you'd be using this as an overly complex IP geolocation lookup and UA header parser? All of this information is either already known to you when you trick someone to click on that link (because they are on your website so you get the same data) or can instead be obtained by tricking them into clicking on a link to your website instead of that one (if it's an email for example).
Or you provide the link via any social media website, or email, etc..
> or can instead be obtained by tricking them into clicking on a link to your website instead of that one
The target may feel much safer clicking a link to knowntrustedsite.abc than yourunknownphisysite.xyz
Big doubt honestly, at least for the vast majority of people. Just set up some blog with some contest that interests them (can be copied from other blogs) and they're not gonna suspect a thing. People might notice if you pretend to be their bank but are actually a phishing site, but they don't notice if you pretend to be a blog and are actually a blog that harvests their data just like any other website.
It was satisfying to remove the cookie notice, and to update the privacy policy with "we do not collect any personal data". That's much more in line with how I want to operate a website.
On my personal blog, I got rid of analytics altogether. I don't care to know who visits my website.
That's .. private? Why is it surprising that privacy advocates would say that this is private?
This script is for nginx web server, but I'm sure with some modification it will work on other web servers like apache.
EDIT: Later after some rewrite, I will also publish "stats.sh" a tiny script that generates daily/monthly/yearly static logs.
/usr/bin/awk '{print $7}' /home/logs/access_log.old | /usr/bin/sort | /usr/bin/uniq -c | /usr/bin/sort -r
1) If your PATH is well configured, you can omit the path to the binaries (i.e.: "/usr/bin/sort" becomes just "sort")2) Instead of using awk, you can just use cut.
# Delimit this line by the space character
# and get the 7th field
cut -d " " -f 7 /home/logs/access_log.old
3) Then, there's this: # This is not correct (lexicographic sort)
sort | uniq -c | sort -r
# This is correct (sort numerically)
sort | uniq -c | sort -rn
Why? because otherwise you get 1, 10, 2, 3, 4, 5, 7, 8, 9 instead of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10.Try it yourself:
seq 1 10 | sort # lexicographic sort
seq 1 10 | sort -n # numeric sort
4) If you are just interested in the top 10, you can add sort | uniq -c | sort -rn | head -n 10Cut is much more painful to use.
It should just be `sort -rn | uniq`, no point in sorting the list twice.
"uniq -c" adds a column with the count of unique elements.
For the purposes of returning the most frequent elements, you can only do the final sort once that column has been added.
As an aside, I find Google's model to be less bad than Facebook's attention capitalism model because it permits a diverse, unoptimised and decentralised web, amp notwithstanding.
Same principle applies to security: each of these trackers has basically full control over the page, and if one of them gets compromised and replaced by malicious code, it would be disastrous. The probability of any individual tracker getting compromised is low, but what you should be worried about is the probability of at least one being compromised, and that probability rises exponentially as you add more.
Let n be the number of trackers, let p be the probability that an individual tracker is compromised, and assume compromises are independent. Then: Prob(at least one comp.) = 1 - (1-p)^n ~~ np. In other words, linear in n. This approximation is fine for small n*p. For large n*p, the probability has to be <= 1 of course, so the real probability is sublinear.
When businesses invest in caring about this (and it can even be just one headcount for the biggest of businesses) there's opportunity to reign it in and maximize the value from a given platform, reduce it the need to add others. That helps with security, governance, and adoption of what's there.
Nevermind more forward looking solutions like server side which gets these injected scripts off your site completely and greatly improves security.
> Hey! We use cookies to track you through our website and use this data to optimize your experience! Thanks!
So apparently, their own software isn't good enough to optimize their own site. At least, not without cookies.
Best of both worlds.
https://matomo.org/docs/requirements/#tracking-100-million-p...
[0] https://www-bbkane-com.goatcounter.com/
[1] https://counter.dev/dashboard.html?user=bbkane&token=MEzACRA...
Summarize IPs takes in a list of IPs to create a report, like this[1], about location of IPs, if they are VPN/Proxy/Tor, and type of traffic. There's a limit of 1000 IPs here, but that can easily be circumvented by creating a free account.
Map IPs plots them on a world map, like this [2].
Both of them are free and can be used from command line. We will also filter out the text to find IPs, so you can just pass raw access logs.
[1]: https://ipinfo.io/summarize-ips/result?id=320d5003-a75f-42c3...
[2]: https://ipinfo.io/map/5b49e994-ec7a-416c-8140-b311364f6160
I’m just curious how can you use this solution with a static site generator like jekyll
I've not used it yet, just have it bookmarked for my next project.
[1] https://umami.is
<script defer src='https://static.cloudflareinsights.com/beacon.min.js' data-cf-beacon='{"token": "$SITE_TOKEN"}'></script>