In my own testing PG performed very similarly to a 'real' warehouse. It's hard to measure because I didn't have the same datasets across several warehouses. Maybe in the future I'll try running something against a few to see.
908 karma · joined October 5, 2012
YC Badge: 0x31430d28FAf744B9012E8116F80d58D735180c79
In my own testing PG performed very similarly to a 'real' warehouse. It's hard to measure because I didn't have the same datasets across several warehouses. Maybe in the future I'll try running something against a few to see.
That said, given the performance implications, if someone wants to use PG as a warehouse upgrading to 12 is a no-brainer.
My uninformed assumption is if I do a group by over all rows in a table that they may not perform better.
I'll look into their continuous aggregates -- that could be one way to get around the cost of aggregating everything if it's done incrementally.
Narrator (YC S19) is a library of expert-written data analyses that anyone can run instantly on top of their data.
Our stack is Python and React.
We're looking for a senior data engineer and senior software engineer (preferably with frontend experience) to continue to build out our unique data platform.
Looking for someone with a data or technical background to write content for us. Mostly blog posts aimed for data engineers and analysts.
https://www.narrator.ai Narrator (YC S19) is a library of expert-written data analyses that anyone can run instantly on top of their data.
We're looking for a senior data engineer and senior software engineer to continue to build out our unique data platform.
Our stack is python, react
We're looking for a senior data engineer and senior software engineer to continue to build out our unique data platform.
Narrator (YC S19) is a library of expert-written data analyses that anyone can run instantly on top of their data.
Our stack is python, react
We're looking for a senior data engineer and senior software engineer to continue to build out our unique data platform.
Basically nothing can be done / no decisions made without a meeting. Why? Because X number of people feel they need buy in. If you don't get them onboard and give them a chance to voice opinions you'll be pushing uphill to get work done.
And meetings are actually a fairly effective way to do that -- you have a group's attention for a set amount of time. If you just sent a doc then you'd have to follow up, etc.
That sort of becomes the default, so there are meetings even when that sort of buy in isn't necessary, bc meetings are just how things get done.
Ultimately it was the right call from the company survival perspective -- everyone who was severely diluted (including early investors) at least still have something worth more than $0.
Assuming data is in a standard format then you can share your script for people to run themselves. Obviously this is fairly difficult in practice unless you can bundle everything into a client-side script on a website.
For reference Narrator [1] does this -- it puts data into a standard format so that analyses written for one company can be run for another. I'm not suggesting you build your stuff on that platform, but it's an interesting approach that does exist.
Seems pretty interesting. In terms of how it works it seems similar to how LaunchDarkly fetches its feature flags.
In practice if we configured all small bits of data in here it'll happen at app startup and be on the critical path. Do you have some sense about the latency there?
Tried again and now it doesn't load. Maybe I'll try later. Looks cool though
https://docs.narrator.ai/docs/activity-transformations describes these scripts and links to a few examples
If you (or anyone) uses our free tier and wants to upgrade to something between it and the lowest paid tier just send us a message at support@narrator.ai and we'll set up something for you.
The tech stack is Python for the backend scheduling and query engine hosted on AWS. For the frontend it's React. We have some internal data stores for managing our own state and a bit of caching - Postgres, S3, ElasticSearch. We use GraphQL a fair bit.
We actually have several e-commerce companies using our platform (with decently high volume). In practice we tend to see events like 'completed order' 'shipped order' 'product added to cart' 'order delivered'. I.e. they're all very discrete differentiated steps in the process.
There's a bit of an art between when to make a new activity and when to add it as metadata on an existing one. A completed order will more likely have 'discount code' as a feature than 'discount code applied' as an activity for example.
Your order completed event could have total amount along with tax, shipping, etc costs that add up to the total. It depends on the analyses you want to generate.
We do see things like an order submitted event with the total, num products purchased, discount code on it, and a separate 'purchased product' event with individual product price, sku, etc. Once can do things like MRR and another could let you identify best selling skus or product categories.
Happy to chat more offline if you want to dive into the specifics for your use case. We love digging into what sorts of analysis someone wants to do and figuring out which activities make sense https://calendly.com/ahmed-narrator/30min-1
It's one of the more common objections people have as they understand the model, but in practice we've found that it's not an issue. Our CEO loves asking people to describe their hard data questions and then redefine them in terms of the activity stream.
The metadata support tables are actually an exception -- we don't use them frequently in practice.
Thanks for engaging with us. If you ever want to dive deeper into this we're always happy to chat.
2. The report (we call it Dataset) you build with the Narrator UI is a table that you can aggregate different ways, plot, and export (including writing back to the warehouse as a materialized view)
3. done optionally as part of 2
4. Yes, we keep anything written back to the warehouse up to date. You can control the cadence.
Because of 4. we work well with BI tools like Looker. Once you have the data you want just point Looker to the right table in the warehouse.
I'll try to address your points in order.
Yes, we agree that 99% of the work is determining what the data means. Our structure doesn't magically make things better because of its structure. It's that once you have it analysis / aggregation on top of it becomes substantially less work because you don't have to constantly redo models to answer new questions.
Yeah, we've all been there -- production DBs aren't typically architected to store historical data. But we've found in practice that the data sources you most care about do have it. Page views, emails sent / received, completed orders, etc. all have timestamps. And for some things you don't need it. If you wanted to do a query with all customers who are VIPs, you wouldn't need a 'became VIP' activity. Adding is_VIP as a feature to the customer in the activity stream works too. Generally if you can do an analysis the more traditional way then you should already have the data to do it in Narrator too.
Sure, star schemas are the way of doing things and this is a new approach. But the efficiency gains realized by our own data scientists are enough to where they wouldn't go back -- it warrants the investment in learning it. Our challenge as a business will be how to convince others of that as well.
What we mean by single source of truth is that data is internally consistent - each term is defined once. In your scenario you'll have a single 'completed order' activity with the total order amount. If you want to add shipping cost that's fine -- add a 'added shipping to order' activity with the cost in it. Do the same with sales tax. Bob, Alice, and Jimmy can create reports with whatever activities they want. The crucial point is by making those reports they're not defining a new model. They're just combining activities. Since all tables are generated straight from the activity stream a future analyst won't use a materialized view based on Bob's data to build a new report -- they'll build it straight from the original activities.
That being said we would love to partner with a CDP like mixpanel and amplitude to have marketers and product people get quick insights using the data that is modeled and cleaned by the data team.
As far as flexibility beyond 11 columns: I'd love to know your use case.
We do support additional metadata on each activity with what we call enrichment tables.
Some events are going to need more metadata -- a page view would want to have the actual page, the five UTM parameters, referrer, etc, which is more than the 3 fields of metadata we store on the activity stream.
So we also support creating additional tables to add metadata to each activity. Each row requires a unique activity id and its timestamp and can an unlimited number of additional columns.
We'll then automatically join that table into the activity stream when queries need it.
You can check out our pricing page here https://www.narrator.ai/pricing
The initial consultancy approach helped us build out the product. Once we could show internally that it made us far faster to analyze data we were ready to launch.
- activity_id : a unique identifier for the row
- activity : the type of activity (eg 'page_view')
- timestamp : time the activity happened
- customer : the unique customer identifier
Metadata columns
Three columns for any info we'd like to add to an activity. Eg for a purchased product activity it could be product name.
- feature_1
- feature_2
- feature_3
- revenue_impact : the amount of money related to this activity. A completed order activity would have this
- link : a hyperlink related to the activity ('ticket submitted' might have a link to the ticket in Zendesk)
Additional customer identifier - source and source id are used when you're not entirely sure who the customer is. For example, a 'page view' activity wouldn't know the actual customer, but might have a unique identifier. So the source could be 'segment.io' and source_id could be their generated uuid- source
- source_id