HNHacker News
TopNewBestAskShowJobs

mejakethomas

256 karma · joined July 24, 2012

data
submissionscomments
mejakethomas··on Big data is dead
Snowflake too.

Inefficient sql? Crank the virtual warehouse.

mejakethomas··on Big data is dead
Yes! This!!!

Volume != Quality

mejakethomas··on Big data is dead
So what I'm hearing is it's not the size of your data that matters, it's how you use it?
mejakethomas··on Why is Snowflake so expensive
Completely agree. Currently staring at 700k+ BigQuery costs annually and accomplished MUCH more with Snowflake at the same price.
mejakethomas··on Why is Snowflake so expensive
This. Snowflake introspection five years ago looks very, very different than today. Mostly due to enhancement requests.
mejakethomas··on Why is Snowflake so expensive
Totally agree with Redshift sentiments. It's been lovely seeing BigQuery and Redshift step their game up over the past 1.5yrs, because they really should have been doing certain things for many years prior.

Re: Firebolt, I don't consider it to be in the same class as Snowflake whatsoever (even though their advertising seems to indicate otherwise). Snowflake is like a very powerful swiss army knife. Firebolt is good for a very specific (dare I say niche?) workload but falls all over itself for the vast majority of data org needs.

mejakethomas··on Why is Snowflake so expensive
This, 100%.

It eats/consolidates formerly-disparate costs around the org. Because it's so good.

Which makes it look expensive.

mejakethomas··on Why is Snowflake so expensive
It's not expensive.

What it can do, successfully, with three engineers was previously impossible with dozens.

What IS expensive is not being careful with it.

mejakethomas··on Anyone else feel the constant urge to leave the field and become a plumber?
Ah yes, the common misconception that trades are simplistic!

Most trades (!including plumbing!) have VASTLY more stringent requirements than tech, are often more mentally stimulating, and yes, quite rewarding.

Disclaimer: I've worked in software for ~9 years and worked in trades before that. ASE master certified auto technician, father is a master electrician, uncles are contractors. I'll probably become a plumber.

mejakethomas··on Ask HN: Is there a way to efficiently subscribe to an SQL query for changes?
PipelineDB was awesome - built some v. neat things with it 4-5 years ago. Wouldn't use it now since the team went to Confluent. https://ksqldb.io/ functionality is getting close.

Flink is a solid option. Materialize shows a ton of promise.

mejakethomas··on Ask HN: Former software engineers, what are you doing now?
big fan of what you're doing @ elevated woodworking! gorgeous stuff
mejakethomas··on 268B Events with Snowplow Analytics and Snowflake
Hey cool! Thanks for sharing :)
mejakethomas··on Show HN: I replaced Google Analytics with simple log-based analytics
You'll still have to anonymize IP addresses, since those are classified as personal data in the EU.
mejakethomas··on Show HN: I replaced Google Analytics with simple log-based analytics
(data engineer here)

Nice post! It's always fun reading about people being creative and challenging the analytics status quo (aka GA). Besides the joy of doing it yourself, you've accomplished a couple other things worth mentioning:

1. You'll never be sampled. GA samples historical data pretty heavily, and you have to pay for 360 to retain unsampled event data (at a tune of $160k+ per year).

2. You have full access to all generated data.

I'd highly recommend using Snowplow's javascript tracker (https://github.com/snowplow/snowplow-javascript-tracker) in a very similar manner to what you've outlined here. You'll get a ton of extra functionality out of the box, which would add yet another level of insight. With snowplow, you get the following for free:

1. Sessionization, which is consistent with google analytics' definition - effectively a 30 minute window of activity.

2. User identification - the tracker drops a persistent cookie (just like GA), so you can see returning visitors.

3. Tools for splitting requests

4. A variety of event types, out of the box: https://github.com/snowplow/snowplow/wiki/2-Specific-event-t...

5. Ability to respect Do Not Track

6. Time on page, browser width/height, etc

7. Ability to make your event tracking 100% first-party

(Disclaimer: I don't work for them, but I've seen the system work very well a number of times.)

I'm running a similar setup on my blog, and it costs well under $1 per month: https://bostata.com/client-side-instrumentation-for-under-on.... I'm doing the same exact thing with Cloudfront log forwarding and have several lambdas that process the files in S3. From there, I visualize traffic stats with AWS Athena (but retain a ton of flexibility, since they are all structured log files).

mejakethomas··on Show HN: I replaced Google Analytics with simple log-based analytics
Totally agree! I'm running a minimalistic version of snowplow's collection/etl infra for under $1 per month, and it works great:

https://bostata.com/client-side-instrumentation-for-under-on...

mejakethomas··on Pluralsight will acquire GitPrime for $170M
I'm at pycon right now, and this is hilarious.
mejakethomas··on Roll Your Own Analytics
I totally agree
mejakethomas··on Roll Your Own Analytics
And also, huge props for the creativity! You hit a nerve here, and it's a topic being discussed at pretty much any/all companies.
mejakethomas··on Roll Your Own Analytics
Props for diving into such a hot topic though. This is a really big deal for companies of many sizes (probably _all_ sizes), and bringing creativity to this conversation is +1
mejakethomas··on Roll Your Own Analytics
Agreed!
mejakethomas··on Roll Your Own Analytics
Snowplow is awesome - it doesn't come with a UI but here's a sample of what data is included:

https://github.com/snowplow/snowplow/wiki/canonical-event-mo...

A pretty common move is to drop this data into redshift/snowflake and query it with Mode/Looker/Tableau/whatever. Athena is a viable option as well, until you get into higher data volumes and don't want to pay for each scan.

Context: I'm a tech lead (data engineering) @ a public company, have set this system up 15+ times @ numerous other companies, and could not live without it at this point. Current co's snowplow systems process 250M+ events per day peaking @ 300k+ reqs/min, on very cost-efficient infra.

mejakethomas··on Roll Your Own Analytics
I've found the documentation here to be very comprehensive if you want to start learning why/what/how:

https://github.com/snowplow/snowplow/wiki/javascript-tracker https://github.com/snowplow/snowplow/wiki/canonical-event-mo... https://developer.matomo.org/api-reference/tracking-api

mejakethomas··on Roll Your Own Analytics
(Data engineer here)

Nice article! I did something very similar to this for my blog but used Snowplow's javascript tracker (https://github.com/snowplow/snowplow-javascript-tracker), a cloudfront distribution with s3 log forwarding, a couple lambda functions (with s3 "put" triggers), S3 as the post-processed storage layer, and AWS athena as the query layer. The system costs under $1 per month, is very scalable, and is producing amazingly good/structured data with mid-level latency. I've written about it here:

https://bostata.com/post/client-side-instrumentation-for-und...

By using the snowplow javascript tracker, you get a ton of functionality out of the box when it comes to respecting "do not track", structured event formatting, additional browser contexts, etc. If you want to see how the blog site is functionally instrumented, filter network requests by "stm" (sent time) and you'll see what's being collected.

I've found (after setting similar systems for 15+ companies of varying scale) that where a system like this breaks down is when you want to warehouse event data and tie it to other critical business metrics (stripe, salesforce, database tables that underpin the application, etc). Another point it starts to break down is when you need low-latency data access. At that point it makes more and more sense to run data into a stream (kinesis/kafka/etc) and have "low latency" (couple hundred ms or less) and "high latency" (minutes/hours/etc) points of centralization.

Using multi-az/replicated stream-based infrastructure (like snowplow's scala stuff) has been completely transformational to numerous companies I've set it up at. A single source of truth when it comes to both low-latency and med/high-latency client side event data is absolutely massive. Secondly, being able to tie many sources of data together (via warehousing into redshift or snowflake) is eye-opening every single time. I've recently been running ~300k+ requests/minute through snowplow's stream-based infrastructure and it's rock-solid.

Again, nice post! It's awesome to see people doing similar things. :)

mejakethomas··on Ask HN: How do you continuously monitor web logs for hack attempts?
During my time at Wanderu we set up Filebeat on all app machines, forwarding json logs to Kafka, and then on to ELK and Pipelinedb for analysis/alerting/session killing/ etc.

It worked very well, as there are almost always application edge cases that should not be blocked immediately with Cloudflare or at a FW level.

ELK was used for exploration and/or setting alerts on thresholds that we had found/defined beforehand.

Pipelinedb was used to run continuous aggregates on a stream, augment (pop json log lines off a stream, enrich them with Maxmind, etc) logs in "realtime", and immediately surface malicious behavior. We'd then programmatically kill sessions or add Cloudflare rules upstream, based on the behavior that surfaced in a Pipelinedb continuous aggregate.

Wanderu: https://www.wanderu.com/ Filebeat: https://www.elastic.co/products/beats/filebeat PipelinedB: https://www.pipelinedb.com/

mejakethomas··on Show HN: Find Remote Work
When I applied, they said the queue was backed up.
mejakethomas··on How to succeed in school...and life.
No problem!
mejakethomas··on My College Experiment: Building a drone with financial aid money.
Cool! Thank you. I'll look into that!
mejakethomas··on My College Experiment: Building a drone with financial aid money.
Control of Mobile Robots - https://www.coursera.org/course/conrob
mejakethomas··on My College Experiment: Building a drone with financial aid money.
Yeah there actually is- Financial aid $$ for the drone stuff :) I figure if I'm going to get it, I might as well put it to good use.
mejakethomas··on The Education I wish I Had
Google Ed: online degrees accredited by Google without the $100k+ burden when you're done. Yes?