Taking on Google
inthegood.co
inthegood.co
What the hell? How do you calculate this figure? That's roughly equivalent to the CO2 created by driving a gas-burning car 10 feet to the data center to ferry information about each request (figuring a car emits about 1.2 pounds of CO2 per mile traveled). That's an astounding claim, and there's no effort to even explain the idea behind it.
A typical server rack might use anywhere from maybe 5-50 kWh. Let's say Google has really beefy ones that consume 100 kWh per hour. That's 92 pounds of CO2 per hour. For the 12 seconds you mentioned, that's still only 0.011 kilograms of total CO2 used. And the claim is that they're BETTER by 4.5 kilograms.
They've gotta be talking about some other expense than the server. But what sort of expense? The cost to build a server? Something about general maintenance of the Internet? ISPs between clients and the server?
https://plausible.io/lightweight-web-analytics
The file size savings is 44.3 kB per visitor, which over 120,000 visits is 5 GB per year.
The Website Carbon Calculator uses a ratio of 1.8 kWh per GB of data transferred, and 475 g CO2 generated per kWh.
https://www.websitecarbon.com/
44.3 kB per visit * 120,000 visits per year * 1.8 kWh per GB * 475 g CO2 per kWh = 4.5 kg per year after fixing units.
"These numbers are all estimates but you can imagine if millions of website owners and Google Analytics users end up making a similar reduction in their website size too. The total reduction in the carbon footprint of the web would be immense."
This seems wildly high, even counting all the hops. For reference, 1.8 kWh is enough to move my car 7 miles and my e-bike over 100 miles.
https://www.websitecarbon.com/how-does-it-work/
"Energy intensity of web data
Energy is used at the data centre, telecoms networks and by the end user’s computer or mobile device. Of course, this varies for every website and every visitor and so we use an average figure. The figures used are for 2017 from the report On Global Electricity Usage of Communication Technology: Trends to 2030 by Anders Andrae and Tomas Edler, adjusted to remove manufacturing energy as this is not relevant to this calculator. We then divide the total amount of energy used by the total annual data transfer over the web as reported in the Nature article, How to stop data centres gobbling up the world’s electricity. This gives us a figure of 1.8kWh/GB."
So in power per GB transferred, it's counting all the power used by people's 60" internet-connected TV displays.
Which is, obviously, absurd to include if you're trying to measure the marginal effect of additional data. More data doesn't increase your screen's power consumption, obviously.
An accurate claim for Plausible would have to be based mainly on marginal increases of power by datacenter and communications networks.
I always find the discussion of marginal increases of energy tricky. If I buy a plane ticket on a half-empty flight, obviously that flight was going to take off anyway, so the marginal increase of my weight plus my luggage is fairly negligible in comparison, so I'm only to "blame" for a fraction of the fuel spent, right? But who else is there to blame except the passengers, without whom there would be (eventually) no flights? So shouldn't we all divide the blame evenly?
The first type is when marginal increase can lead to a "new unit", like planes you refer to -- or servers used by data centers. If a plane fits 100 people, then (simplifying) 1/100 of the time you'll result in a new plane being used, so it makes sense to divide the plane's total resources by passengers -- not just the fuel you used.
But the second type never results in a "new unit". In this scenario, using more resource-hungry analytics will never push someone to purchase a second cell phone to spread the load. So counting anything but marginal energy increase usage by the CPU directly is disingenuous.
So in the case of analytics software, their data center server/power resources fall into the first type. But the consumer device resources fall into the second type.
So in this case I don't think there's anything tricky at all about it.
Also as a note, the google analytics js is heavily cached and thus doesn't have to travel as far or at all. Also Google has onramps to their carbon neutral infrastructure everywhere, so theres also that.
If we removed 40k of CDN content per visit then the 1.8 kwh/GB would be 2.8 kwh/GB.
So, an absolutely negligible amount of CO2?
By virtually any metric? i.e., you, as an individual, exhale that much CO2 in a week.
Hold industrial processes responsible for CO2 emissions, not your website. (Unless you're bitcoin, I guess?)
Like everything else, the best approach is just to use a mix of regulation, renewable subsidies, and a carbon tax to make using fossil fuels cost prohibitive compared to renewables and the market will eliminate them on its own. The wider the cost difference becomes, the faster renewables will displace carbon energy. We're getting there slowly as wind and solar are now slightly cheaper than carbon fuels, but we should definitely be helping it along a lot faster if we're serious about avoiding the worst case climate scenarios.
So far, it seems like we aren't serious about it and our leadership is sleepwalking us towards increasing catastrophe.
Interesting you say that. There's no reason Plausible could not be used like AWStats. Parsing logs is just a different ingestion mechanism and we already provide self-hosting via Docker. On principle it wouldn't be too difficult to drain your logs into a Plausible instance or just run it on the same host along your web server.
We ran a test last summer and found the stats from our JS-based tracker much much much more usable: https://plausible.io/blog/server-log-analysis
So this is why we haven't put too much effort in log analysis. The stats we got from AWStats were mostly bot traffic with no good way to get rid of them.
Have you run AWStats and Plausible side-by-side? Do you not have ~90% bots in your logs?
But for most, GA is how Google ads knows how to calculate conversions. People who want to use Google ads (which are everywhere) have to use GA. If you are not using Google Ads, I dont think Google cares much about your site anyway.
https://plausible.io/privacy-focused-web-analytics
(I'm the co-founder)
If I'm using a service I won't mind that service knows my traffic pattern. (That's the point of having an account on the service)
Which website?
The Plausible register page uses hCaptcha (https://plausible.io/register).
LE: The best solution is still self-hosting, as hosted plausible is still a 3rd party entity that centralizes data (even though they probably don't use or share this data).
- For the large part the concern is what is done with the data.
- Data can be anonymized. (Although this is often hard to verify)
- You can hide the data in the client. For example imagine you want to know how many users use feature X. You can send an analytics report with 90% chance of a random value, and 10% chance sending the true boolean. You can't tell if any specific user has used the feature (because most likely it is a random value) but you can get a pretty good estimate what portion of your users use the feature.
My understanding is that Plausible is focused on the use an anonymization.
For me the default uBlock origin settings do block Plausible tracking, even on a website that used their own domain name to serve the script, but I assume it was because the name was "analytics.site.com".
https://plausible.io/plausible.io?period=day (39 current visitors)
They might well be the next Elastic/CockroachDB/MongoDB/etc. Or better yet, they might do the classic bait-and-switch later on: get developer buy in with a good story about openness, then once they'd gotten enough of a customer (aka dev) share, do the switch.
We were on the MIT first and got into a situation where a large corporation wanted to take our code and resell it to tens of thousands of their customers and they made it clear they didn't want to contribute anything back to our project whatsoever.
We are a two person team putting our own time and savings into this and it could have instantly killed the project and the chance of becoming sustainable.
We changed the license and that was a simple way to stop them without changing our principles/ideas. Could have gone proprietary too at that stage but we didn't.
Everything is clearly explained here https://plausible.io/blog/open-source-licenses
Absolutely nothing. That person doesn't know what they're talking about.
I am sorry to hear that you learned about the peril of a permissive license in the way you did, but I'm happy that you switched to strong copyleft. Arguments demanding permissive licensing instead of strong copyleft amount to saying "but then how will I stand on your neck?" You shouldn't have to put up with that.
Sometimes people learn something new that changes things. Sometimes situations change and so the strategy needs to change. Sometimes people realize, for whatever reason, they were wrong and so they take steps to correct it. Do some people sometimes flip flop for the purpose of misleading people or pandering? Of course. But I really don’t think that’s typically the motive. We should be supportive of people changing their minds, not suspicious.
At my app https://hanami.run I don't track user and cannot know if the same users visit our website :-(. I don't want to use cookie and want to get away with GDPR. At the same time, I love to see which visitors repeatly read my website/blog and where they drop so I can optimize my site.
Any recommended alternatives would be appreciated as well.
I imagine hCaptcha doesn't have enough trackers sprinkled around the web to use those as signals for this.
There will always be some individual variance, but when we've tested this people always solve hCaptcha faster than reCAPTCHA on average.
(disclosure: work there)
https://www.youtube.com/watch?v=tbvxFW4UJdU
It runs entirely in the background, and pretty much the only time you'll see a prompt is if you're using a VPN, Tor, or specifically block it.
Basically you are blocked if you care about privacy and refuse this tracking.
That's what I'm willing to call "hostile". I'd say, it's even worse than picking a few pictures, which is already hostile.
Or using a non Google browser or using an account that Google doesn't like (because they can't associate it with a real identity or whatever.)
The plausible landing page gives me zero cookies and only requests are to plausible.io and testing.plausible.io
Some examples: “click all the tractors” showed I did not complete the task because of a photo of construction equipment; “click all the crosswalks” because I didn’t select the photo of a thick white fence; “click all the traffic lights” because I didn’t select a photo of a parking meter. I just clicked the incorrect photo so I could move on but I can’t help but wonder if there’s any mechanism to catch those incorrect (manual, human) annotations on the training data Google is collecting.