1,034 karma · joined August 14, 2011
@alexcrdean https://github.com/snowplow/snowplow
Author of Event Streams in Action (Manning Publications), http://manning.com/dean/
My email is alex at my company domain.
Snowplow started at the same time as Segment (2012) but has evolved along a separate tech tree. Micro-service architecture, cloud native, using Kinesis or Cloud Pub-Sub as the data transit, enrichment framework plus a Confluent-style schema registry supporting very rich and versioned JSON Schema-based event payloads. We are built by and for data platform teams; our open-source behavioral data engine doesn't have a UI (our commercial Behavioral Data Platform does). Hosted trial here https://try.snowplowanalytics.com/
Definitely room for both product families in the market! I'm sure Jitsu will do great.
We built a kind of Maven Central for 3rd party webhooks' schemas[2], and got to work adding schemas for various popular webhooks, e.g. Sendgrid and Pingdom. But, we never met a single SaaS vendor who was interested in 'adopting' the JSON Schemas for their webhook, let alone publishing versioned JSON Schemas for their whole API. This meant that at Snowplow we have stayed on the hook for keeping these vendors' webhook schema definitions up to date in our registry.
It's a real tragedy of the commons - oceans of developer time wasted on tedious low-leverage work (updating API client code) so that each SaaS vendor can 'move fast and break things'. The ultimate irony is that if these vendors were to adopt, publish and respect versioned schema definitions internally, using something like OpenAPI, they would see huge productivity gains (think enhanced CI/CD testing, auto-gen of client code etc).
1. https://docs.snowplowanalytics.com/docs/getting-started-on-s... 2. https://github.com/snowplow/iglu-central/tree/master/schemas
This said, we are steadily working to refactor all our components to be more generic, and already we run a lot of the Snowplow components on k8s for our customers.
With this background, I am delighted to be able to Show HN our new Try Snowplow experience. Under lockdown last year, the team decided to go back to basics and build a new version of Snowplow called Try Snowplow; that version is now in GA.
Try Snowplow is a small version of the Snowplow technology that can be setup quickly and easily - normally under 15 mins. Like all versions of Snowplow, it runs in your own cloud environment.
The landing page is here: https://snowplowanalytics.com/get-started-try-snowplow/
And if you want to learn more about Try Snowplow before giving it a spin, there's a video here: https://www.youtube.com/watch?v=Aw5hdIjwVhY&ab_channel=Snowp...
If you are still puzzled why you want your own behavioral data in your own warehouse, check out our library of use cases here: https://snowplowanalytics.com/use-cases/
Any questions just shout! Happy to answer anything Snowplow-related, and thanks again for Hacker News' support of Snowplow over the years.
As well as the SaaS packages like Amplitude and Mixpanel, you also have great open-source tools and platforms for mobile and product analytics like PostHog (https://posthog.com/), Countly (https://count.ly/) and Snowplow (https://snowplowanalytics.com/).
Disclosure: Snowplow co-founder.
http://snowplowanalytics.com/blog/2014/05/13/introducing-sch...
SemVer doesn't work for data - for one thing, there is no concept of a "bug" in your data (so patches are meaningless).
We have hundreds of companies actively using SchemaVer via the Snowplow (https://github.com/snowplow/snowplow/) and Iglu (https://github.com/snowplow/iglu/) projects.
But you're right, maintaining even a blessed third-party library like boto is a massive undertaking: 437 contributors, 431 open issues, 192 open PRs. Competing against CloudFormation using anything other than the Java client library is crazy IMO.
I'm interested in adding SnowPlow support to this (https://github.com/snowplow/snowplow) - our tracking API is very similar to Google Analytics's.
We've just gone through the quite involved exercise of mapping SnowPlow to Google Tag Manager (http://snowplowanalytics.com/blog/2012/11/16/integrating-sno...) so I was a bit surprised in the code to see this mapping for GA events:
track : function (event, properties) {
window._gaq.push(['_trackEvent', 'All', event]);
}
I'm a bit confused by this - how would I use analytics.js to pass through all the valid data that I can log in a GA event or indeed SnowPlow event - https://github.com/snowplow/snowplow/wiki/javascript-tracker...I think you might be making the assumption that events consist of an event name plus arbitrary JSON envelope. This is a very MixPanelish view of the world - it doesn't really translate to Google Analytics, Piwik, SnowPlow, Omniture...