Migrating from Google Analytics
thomashunter.name
thomashunter.name
The advantage of doing it this way is you preserve the event level data forever and you can write arbitrary SQL against it. You will need a BI tool to visualize; there are several excellent ones with free or cheap tiers for small companies.
That way, you will have more control and be able to ask any question you want. We have been working on a tool that basically connects to your Segment or Snowplow data and run ad-hoc analysis similar to Mixpanel and GA so that you don't need to adopt generic BI tools and write SQL every time you create a new report. I was going to create a Show HN post but since your comment is quite relevant to the topic, I wanted to share the product analytics tools that we have been working on, https://rakam.io. The announcement blog post is also here: https://blog.rakam.io/update-we-built-rakam-ui-from-scratch-...
P.S: I also genuinely appreciate the work the people have done at Countly, it's often not easy to use ETL tools to set up your own data pipeline and create your own metrics yourself so they're a great alternative if you don't want to get stuck with GA or third-party SaaS alternatives.
Depending on the data volume, you can use one the SQL based data warehouse solutions such as Redshift, Presto, BigQuery or Snowflake (even PG works if you have less than 100M events per month) and run ad-hoc queries on your raw customer event data easily. It's a bit tricky to run behavioral analytics queries such as funnel and retention but we provide non-technical friendly user interfaces that don't require you to write any SQL.
I would love to talk about your use-case, feel free to shoot an email to emre {at} rakam.io
While Lucene synax is very powerful, it is not SQL as pointed out by the OP. If you have a lot of spare developer time and people skilled in this, it will likely work well for a while (potentially a really long while).
Going with something like BigQuery or Redshift enables you to utilize other tech to supplement or accelerate skill sets, such as paying a SaaS for visualization/analysis tools (Looker, Mode, etc).
In particular, separation of storage and processing helps avoid costs when you are not using the data. With elastic, I believe you need the cluster whether you are using it or not. BigQuery will only charge you for compute you use and pennies for storage. Same thing for Athena/Redshift spectrum. Snowflake is similar, but I believe for enterprise contracts it a bit more complex with minimums and such.
Structuring data is another really important consideration - will you need to normalize and will you need to update labels/dimensions? That’s just the start.
Nope and nope. That's definitely the scary / unknown part!
> If you have a lot of spare developer time and people skilled in this, it will likely work well for a while (potentially a really long while).
We don't have a lot of spare developer time and people. However, the Elastic docs make our fairly straightforward use case - dump lots of events in and filter them by date and about 4 different properties - seem not too daunting.
The way it's sold on the Elastic site these days seems like a very "batteries included" kind of approach - am I interpreting their marketing a little too positively? Are we kidding ourselves? Would love to hear more about your thoughts!
> Structuring data is another really important consideration - will you need to normalize and will you need to update labels/dimensions? That’s just the start.
We don't expect much.
Thanks for mentioning the other technologies, too! The way we're structuring it allows us to either use Elastic, SQL, or something else entirely, so thankfully we're fairly agnostic on the data store. If Elastic turns out to be a time sink, we'll have no hesitation trying something else out.
Elastic is really good for data exploration. Kibana's discovery page makes it trivial to start diving into your events and gain insights. This is most valuable for large, complex applications where you don't necessarily control the analytics tagging (thus devs can create new tags willy-nilly). It's also easier for non-technical users to get into, since there don't need any sort of SQL experience to build reports.
Elastic is really bad at performing aggregations. You can do it, but they are incredibly memory-intensive, and using them puts the cluster at risk of failure due to out of memory exceptions. It doesn't take many concurrent users for this to be an issue.
Traditional data warehouses are still the best place to get counts though. Assuming your users can write SQL, they can get everything they would want and it will be reliable. Newer databases support JSON structures, so you can even host arbitrary-sized dynamic event data in them like you can with Elastic.
With Google the TOS states that you’re not allowed to store personal information in the first place at least.
events (id, user_id, ...)
users (id, first_name, last_name, email, ...)
If a user requests to be deleted from your database, or if you have a policy of deleting users when the cancel their subscriptions, you can accomplish that by deleting one row.This isn't going to make you GDPR compliant by itself but it's a start.
"any information relating to an identified or identifiable natural person (‘data subject’); an identifiable natural person is one who can be identified, directly or indirectly, in particular by reference to an identifier such as a name, an identification number, location data, an online identifier or to one or more factors specific to the physical, physiological, genetic, mental, economic, cultural or social identity of that natural person"
Pseudonymization (separating unique identifiers from the rest of the datapoints) can and should be used as a safeguard, but doesn't remove the need to protect the data, particularly if you keep a link between the two.
I didn’t say it was. The original conversation wasn’t about PII it was about anonymising. Fingerprints defeat anonymisation.
I reckon in practice with for example an Amazon sales log with names removed and my public Twitter feed you could de-anonymise me with a high degree of certainty.
You will not and will not assist or permit any third party to, pass information to Google that Google could use or recognize as personally identifiable informatiom
Seems like using Segment here is probably the least effort solution for future events, although it doesn't address ETLing past GA data into the warehouse. Curious if anybody else has had to deal with a similar situation.
So it’s best to start early, because it’s a slow process.
It’s even slower if you don’t instrument the right events in the beginning, because you’ll wait a while to gather enough event data to perform an analysis, find you want to track some new event, and then need to wait even longer to have enough data for analysis. It can be frustrating, and a surprising amount of work.
If you want to build out your BI infrastructure in the mean time, you could always ETL your e.g. server logs into the data warehouse as a stopgap. It’s definitely not the same as the customer-centric setup that you get with Segment, though, unless you have some advanced logging already set up.
Would live a recommendation or 2.
Saiku uses Mondrian (OLAP) tech to give a slice/dice drill down view of data. Unfortunately the docs for Saiku seem very disjointed these days - there's a docker image for the CE (Community Edition) that I recently setup okay. Quite a learning curve on setting up Mondrian schemas here, though, but worth it when it's done.
[1] https://redash.io [2] https://www.meteorite.bi/products/saiku
We're a small startup though, so last thing I wanted was even the slightest ops load from a new system.
My strange conclusion was to not capture analytics at all. My loose metric is now: which posts inspire people to email me directly?
I have decided I do not want to know what you read, how long you read it, where you came from... Only whether or not you felt something. And data will not tell me this.
More specifically, how often does all of this really need to be down to the individual person (or IP address) at all? Even if you know that piece of information at an ephemeral level, my own suspicion is that aggregate data should be sufficient for any non-creepy use case.
Perhaps one way to phrase the question, how would having personal-level details in the analytics change the actions you might perform based on available data?
That's why I moved Simple Analytics' (https://simpleanalytics.io) servers to Iceland where the law will forbid peeking in data before actually informing the owner of the server. I encrypted my server so if anything happens I can just turn it off and it will be nearly impossible to get any user data.
Could you please share which Iceland host you are using? Also, what is the process you're following to encrypt your server?
[1] https://1984.is
[2] https://blog.simpleanalytics.io
[3] https://en.wikipedia.org/wiki/Logical_Volume_Manager_%28Linu...
One reason could be Countly doesn't offer transparent pricing. People tends to switch off when they see call for quote.
This was the “killer feature” of the Google Analytics replacement I built for myself a few years back. I would grab referrers for social and scrape them to give a more useful report of the source of traffic.
Unfortunately this is challenging or impossible for many social platforms due to HTTPS everywhere, the prevalence of outbound link scrubbing, and app-driven embedded web views.
I still think it would be a killer feature to know who tweeted you out or which subreddit you are trending on but it’s just not feasible to do based on a website pixel alone.
I imagine originally its intent was for website owners to see what other websites linked to them but now it gets used to track the users
Clicky: https://clicky.com
Fathom: https://usefathom.com/
What's the Catch?
But I had an error installing Matomo and got not help in the official forum [4], if someone here can help me I will appreciate it, I'm not a developer (I'm also considering get rid of analytics tools for good, anyone?). Thanks.
[3] [redacted]
[4] https://forum.matomo.org/t/fatal-error-on-installation/29949
"I have a Hostgator shared hosting with PHP 5.5."
PHP 5.5 is EOL since mid of 2017. Can't you switch to newer version such as 7.2?
Matomo itself does not even display the required version according to [1]. Their FAQ talks about 5.3 - which is horribly ancient too.
I have to agree that one should be using PHP 7.2. It also gives a nice performance boost to Matomo. The required PHP version for Matomo is shown in [1] (5.5.9 or greater) Can you please send me a link to the FAQ page mentioning 5.3 (e.g. to lukas@matomo.org) so they can be updated?
Additionally, the error described by OP looks like autoloading is broken.
[1] https://support.hostgator.com/articles/php-configuration-plu...
It was *superior. Often times when people ran into a wall within GA they would flock to Urchin and its fairly reasonable price.
https://www.cardinalpath.com/what-is-google-urchin-the-diffe...
> How is the site doing compared to last week?
If this is really all the information you need, try out a server-side solution like GoAccess [0]. It is a little bit harder to set up than copy&pasting an HTML snippet, but it is well worth it (privacy, performance, etc).
For that reason, most analytics systems use come kind of client-side tracking code that runs within the client web browser. That way it works regardless of whatever caching happens.
* Sync your (snowplow, segment, ...) data to a data warehouse
* Try to clean up your data with SQL
* Keep running GA so you can compare your numbers to GAThe blacklist is on GitHub and is useful for other analytics software, such as GoAccess which is what I use.
Of course you can use the same list for every other software.
[1] https://github.com/matomo-org/referrer-spam-blacklist/blob/m...
That said, there are solutions where a cookie allows for greater privacy because it allows you to leave sensitive data on the client versus having to store it on the server. For example, https://usefathom.com/ is an open-source self-hosted analytics tool that uses a cookie to enhance privacy.
You can configure Matomo [1] to both not use any cookies [2] and to automatically delete just the raw data or all data that is older than x months. Log Analytics is also possible.
If you want something that is far more minimalistic, but also Open Source and self-hostable, you can take a look at [3]. (Not sure about how they use cookies)
(Disclaimer: I am part of the Matomo team)
[1] https://matomo.org/ [2] https://matomo.org/faq/general/faq_157/ [3] https://usefathom.com/
You can also configure to remove any data older than N days/months.
It is also GDPR compliant [2]
[1] https://resources.count.ly/docs/countly-sdk-for-web#section-...
eg, this is the Countly one (specifically run as root):
wget -qO- http://c.ly/install | bash
If someone manages to break into the c.ly redirection service, or the website/CDN/etc serving that, new users would likely be in for a bad time. And the problem could be very subtle, if it's done by someone clueful.Just for added er.. goodness (/s), the Countly instructions also need people to:
Disable SELinux on Red Hat or CentOS if it's enabled.
Countly may not work on a server where SELinux is enabled.
In order to disable SELinux, run "setenforce 0".
sighBut GA provide segments that marketers need. How do you get that kind of data without google of the fa pixel ?
I don't feel convinced to switch :(
That your pages are now lighter too is just a bonus.