To relate this to the article: what should these government agencies be using? Or should they not be looking for Javascript errors, A/B testing, etc. at all?
To relate this to the article: what should these government agencies be using? Or should they not be looking for Javascript errors, A/B testing, etc. at all?
I look at the data in Google Sheets.
On the page I want to track I paste a script tag that includes a few lines of JS from my counter site. That JS script hits a PHP script with the URL the user requested. I don't track ANY user details. No browser info, no IP address, no fingerprinting, etc. It would be trivial to track those things though. The PHP script logs the data to a CSV file (which I plan to change to an SQLite DB soon).
I have a Google Sheet setup where the first field of data is '=IMPORTDATA("https://example.com/data.csv")'. Google Sheets automatically fetches that data every time you open the sheet; no API integration required. Then I have a simple bar chart on the data.
We as a society haven't agreed precisely on what "privacy" means so it is effectively impossible to know is a particular service's handling of data provided to meats you definition unless you just don't hand it the data in the first place.
I mean, this is always going to be the case with modern computers. No one writes EVERYTHING themselves, so they have to trust someone else. You are trusting the microcode on the CPU, the system calls on your OS, your compiler/interpreter, your standard library.... I get that this is a different sort/level of trust than trusting a third party metrics system, but it isn't fundamentally different. It is all about trusting someone else's work.
if you've got php running already, it's straightforward to code up a bar chart from the weblogs you already have (bypassing, csv/sqlite and google sheets altogether). that is, after all, how google analytics started (as urchin).
I doubt that you care that much since the data isn't sensitive but just a heads up.
It's security-by-obscurity, maybe (as all public "secret token" URLs are) but it's better than what you're implying.
EDIT: You are right in theory though.
This is strictly as a learning exercise, no malicious intent on my part.
The javascript is located at SITE/counter.js
My first guess for the CSV was SITE/counter.csv
It worked.
Is "matomo" japanese? If yes then the definition is here https://www.nihongomaster.com/dictionary/entry/36221/matomo
Unluckily it cannot give you directly the search string used by people that ended up on your site because search engines don't forward it anymore because of the whole privacy movement that happened some years ago (I am still against it as I don't see why, as website owner selling e.g. clothes, I shouldn't know that a person landed on my website by searching e.g. "yellow pants" => this fake privacy just concentrates all power/knowledge in the search providers, but this is just my personal opinion), but here they sell a plugin ($/year) which apparently can do that: https://plugins.matomo.org/SearchEngineKeywordsPerformance
(I guess that it retrieves directly the search keywords from the search provider, but I did not read the docs nor I tried it out)
That doesn't really matter. But if you set up a content farm / honeypot, you shouldn't be able to tell that the search term that brought the person to you is "how to deal with my XXX infection"
The business argument for Google: they still have the information, and can use it in their analytics, and potential competitors or customers don't have it.
Setting up such a "content farm / honeypot" and making it reach the top results of Google/Bing/Yandex/etc... used to be simple but is nowadays probably successful in only very few cases (as search engines are nowadays more and more context-aware), and Google/Bing/Yandex/etc... can still see & use "how to deal with my XXX infection", but whoever doesn't use directly their services cannot.
What I mean is that, in my opinion, the privacy measures in this case centralized even more power in the hands of few companies with very little added/improved privacy.
In my case, running a small techy website, the search keywords were very useful because they allowed me to understand e.g. which keywords forwarded the users to my website by mistake or correctly, to then correct appropriately the contents of my articles to make it more clear what my articles were talking about, or to see that the users had a very specific problem that I did not take into consideration when I wrote a certain article, etc... . Now I cannot see those infos anymore without using Google Analytics which I don't want to use (or, by using the plugin mentioned above, which is good, but for which I would still have to pay $/year, which is bad as it increases fixed costs).
If they gave the information to you too, it likely goes to you, but also to the other 50 .js files you include from various sources of dubious trustworthiness which every site these days includes.
Furthermore, what you are saying is "this admittedly private information used to be available to all and it was useful for some, now it's only available to the entity the user specifically gave it to, and that's bad because the few who actually used it for good don't have it". But the whole idea of GDPR (and similar) laws is to put the control back with the user, which is a good thing.
I think some standard with which the user explicitly lets the website know "yes, the search engine query that brought me here is X and I allow you to have it" would be good, but I don't think dropping this info from the referer (sic) is bad.
Why not? How else are you going to 1) provide information on dealing with an XXX infection, or 2) recognize enough people are landing at your site looking for advice on their XXX infection that you should provide some answers?
Somewhere in the settings, you have the option to anonymise users, which is achieved by deleting/not saving the last triplet of IP addresses.
There are some other options, and some functions you should avoid:
- You should set data retention to the minimum of 14 months
- Do not use the User-ID functionality (tracking across devices/browsers)
- Also avoid Remarketing and Advertising Report features.
The downside to all this is that it is mostly invisible to users that you are somewhat protective of their privacy.
You should try the dlang forums to see what can be achieved if you don’t fetch 2mbytes of JS from 5 different origins: https://forum.dlang.org
But just thinking about my tax return can think of a few features that would have to be dropped, or would become more clunky, like some of the client-side validation, and the auto-saving of drafts.
ELK (Elasticsearch with Kibana). Pretty powerful and AWS recently forked free Elasticsearch.
Has anybody else experienced this?
Do you mean loading the UI which displays the graphs etc...?
That being said, ultimately it's a frontend on a MySQL database, so there's lots of ways it could theoretically be slow -- MySQL isn't configured/tuned properly, the database server isn't resourced appropriately for the amount of data it's hosting, etc. But this is going to be the case for any self-hosted solution.