HNHacker News
TopNewBestAskShowJobs

cedricd

908 karma · joined October 5, 2012

cofounder at Narrator (narratordata.com)

YC Badge: 0x31430d28FAf744B9012E8116F80d58D735180c79

submissionscomments
cedricd··on Using PostgreSQL as a Data Warehouse
Yeah, I think using Snowflake or BigQuery or something is ultimately the better move. But sometimes folks use what they know (what they're comfortable managing, tuning, deploying, whatever).

In my own testing PG performed very similarly to a 'real' warehouse. It's hard to measure because I didn't have the same datasets across several warehouses. Maybe in the future I'll try running something against a few to see.

cedricd··on Using PostgreSQL as a Data Warehouse
Maybe it doesn't get prioritized until it's important. I know PG upgrades are pretty straightforward, but sometimes people don't want to touch something running well.

That said, given the performance implications, if someone wants to use PG as a warehouse upgrading to 12 is a no-brainer.

cedricd··on Using PostgreSQL as a Data Warehouse
Would TimescaleDB be much faster for analytical queries that aren't necessarily segmented or filtered by time?

My uninformed assumption is if I do a group by over all rows in a table that they may not perform better.

I'll look into their continuous aggregates -- that could be one way to get around the cost of aggregating everything if it's done incrementally.

cedricd··on Using PostgreSQL as a Data Warehouse
Yes, that's a really great point. I should emphasize that more clearly in the blog :).
cedricd··on Using PostgreSQL as a Data Warehouse
We support multiple data warehouses on our platform. We recently had to do a bit of work to get Postgres running, so we wrote a high-level post about things to consider when running analytical workloads on PG instead of normal production workloads.
cedricd··on Ask HN: Who is hiring? (May 2021)
Narrator | New York, NY | Full-time | Remote | https://www.narrator.ai

Narrator (YC S19) is a library of expert-written data analyses that anyone can run instantly on top of their data.

Our stack is Python and React.

We're looking for a senior data engineer and senior software engineer (preferably with frontend experience) to continue to build out our unique data platform.

https://www.workatastartup.com/companies/12598

cedricd··on Ask HN: Freelancer? Seeking freelancer? (April 2021)
SEEKING FREELANCER | Narrator.ai | remote | part time

Looking for someone with a data or technical background to write content for us. Mostly blog posts aimed for data engineers and analysts.

https://www.narrator.ai Narrator (YC S19) is a library of expert-written data analyses that anyone can run instantly on top of their data.

cedricd··on Ask HN: Who is hiring? (April 2021)
Narrator | New York, NY | Full-time | Remote | https://www.narrator.ai Narrator (YC S19) is a library of expert-written data analyses that anyone can run instantly on top of their data. Our stack is python, react

We're looking for a senior data engineer and senior software engineer to continue to build out our unique data platform.

https://www.workatastartup.com/companies/12598

cedricd··on An ancient method that keeps Afghanistan's grapes fresh all winter
Not sure if they make them still today, but I saw those pots for sale by the hundreds at a market in Bagan about 10 years ago.
cedricd··on Ask HN: Who is hiring? (March 2021)
Narrator | New York, NY | Full-time | Remote | https://www.narrator.ai Narrator (YC S19) is a library of expert-written data analyses that anyone can run instantly on top of their data.

Our stack is python, react

We're looking for a senior data engineer and senior software engineer to continue to build out our unique data platform.

https://www.workatastartup.com/companies/12598

cedricd··on Machhapuchhare - The Himalayan Peak Off Limits to Climbers
I've been to Machhapuchhare base camp. You can see the mountain from further out on the way up, and when you turn a corner several days later right up close. It's easily one of the more beautiful peaks in the Himalayas.
cedricd··on Ask HN: Who is hiring? (February 2021)
Narrator | New York, NY | Full-time | Remote | https://www.narrator.ai

Narrator (YC S19) is a library of expert-written data analyses that anyone can run instantly on top of their data.

Our stack is python, react

We're looking for a senior data engineer and senior software engineer to continue to build out our unique data platform.

https://www.workatastartup.com/companies/12598

cedricd··on Ask HN: Which companies work like Gumroad?
I know this isn't a fun answer but in my experience it's just bad process and culture

Basically nothing can be done / no decisions made without a meeting. Why? Because X number of people feel they need buy in. If you don't get them onboard and give them a chance to voice opinions you'll be pushing uphill to get work done.

And meetings are actually a fairly effective way to do that -- you have a group's attention for a set amount of time. If you just sent a doc then you'd have to follow up, etc.

That sort of becomes the default, so there are meetings even when that sort of buy in isn't necessary, bc meetings are just how things get done.

cedricd··on Startup Stock Options – Why A Good Deal Has Gone Bad (2019)
I've seen it happen first hand. I had some stock in a company that was running out of money and having issues raising their next round. They ended up taking a deal with an investor that took a huge ownership stake and diluted the company significantly. To entice the founders not to bail they carved out some additional stock for them, but everyone else got hit by the full dilution.

Ultimately it was the right call from the company survival perspective -- everyone who was severely diluted (including early investors) at least still have something worth more than $0.

cedricd··on Ask HN: How to run analytics on data without access to the data?
There's another approach you can do -- make the analysis portable instead.

Assuming data is in a standard format then you can share your script for people to run themselves. Obviously this is fairly difficult in practice unless you can bundle everything into a client-side script on a website.

For reference Narrator [1] does this -- it puts data into a standard format so that analyses written for one company can be run for another. I'm not suggesting you build your stuff on that platform, but it's an interesting approach that does exist.

[1] https://www.narrator.ai

cedricd··on Show HN: Config.ly – Never hardcode your data again
So if I understand correctly, I can define bits of JSON / values in your interface. Then when my client app loads it'll hit your service backend to fetch them.

Seems pretty interesting. In terms of how it works it seems similar to how LaunchDarkly fetches its feature flags.

In practice if we configured all small bits of data in here it'll happen at app startup and be on the critical path. Do you have some sense about the latency there?

cedricd··on The Lonely Work of Moderating Hacker News (2019)
Thanks for all the work you do Daniel! I don't know how you manage it all.
cedricd··on The virtual device farm for rendering websites on multiple devices
Interesting. It breaks on my site. https://seleniumbase.io/devices/?url=www.narrator.ai

Tried again and now it doesn't load. Maybe I'll try later. Looks cool though

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
Yes, the customer would have to define their activities (e.g. 'page view', 'completed order', 'support ticket opened') and write sql snippets to define them.

https://docs.narrator.ai/docs/activity-transformations describes these scripts and links to a few examples

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
Thanks for the feedback. Yeah, we can definitely consider it. We haven't totally optimized our pricing yet.

If you (or anyone) uses our free tier and wants to upgrade to something between it and the lowest paid tier just send us a message at support@narrator.ai and we'll set up something for you.

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
We don't use an underlying storage engine. Everything runs directly on your warehouse. We build the activity stream by running queries against the warehouse and writing to a table that we create inside of it. When people build datasets with Narrator we compile everything down to sql and query the warehouse directly.

The tech stack is Python for the backend scheduling and query engine hosted on AWS. For the frontend it's React. We have some internal data stores for managing our own state and a bit of caching - Postgres, S3, ElasticSearch. We use GraphQL a fair bit.

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
It's not quite that granular. I was more responding to the idea that there can't be a source of truth for these activities.

We actually have several e-commerce companies using our platform (with decently high volume). In practice we tend to see events like 'completed order' 'shipped order' 'product added to cart' 'order delivered'. I.e. they're all very discrete differentiated steps in the process.

There's a bit of an art between when to make a new activity and when to add it as metadata on an existing one. A completed order will more likely have 'discount code' as a feature than 'discount code applied' as an activity for example.

Your order completed event could have total amount along with tax, shipping, etc costs that add up to the total. It depends on the analyses you want to generate.

We do see things like an order submitted event with the total, num products purchased, discount code on it, and a separate 'purchased product' event with individual product price, sku, etc. Once can do things like MRR and another could let you identify best selling skus or product categories.

Happy to chat more offline if you want to dive into the specifics for your use case. We love digging into what sorts of analysis someone wants to do and figuring out which activities make sense https://calendly.com/ahmed-narrator/30min-1

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
Yeah, that's true, but you'd be surprised by the number of things that can be modeled with an activity stream.

It's one of the more common objections people have as they understand the model, but in practice we've found that it's not an issue. Our CEO loves asking people to describe their hard data questions and then redefine them in terms of the activity stream.

The metadata support tables are actually an exception -- we don't use them frequently in practice.

Thanks for engaging with us. If you ever want to dive deeper into this we're always happy to chat.

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
Yes, that's basically it.

2. The report (we call it Dataset) you build with the Narrator UI is a table that you can aggregate different ways, plot, and export (including writing back to the warehouse as a materialized view)

3. done optionally as part of 2

4. Yes, we keep anything written back to the warehouse up to date. You can control the cadence.

Because of 4. we work well with BI tools like Looker. Once you have the data you want just point Looker to the right table in the warehouse.

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
Hmm. That's interesting. I'm not familiar with what Salesforce does. Do you have any more info about it? I'd love to learn more!
cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
Appreciate the skepticism, since yes, we're a totally different approach.

I'll try to address your points in order.

Yes, we agree that 99% of the work is determining what the data means. Our structure doesn't magically make things better because of its structure. It's that once you have it analysis / aggregation on top of it becomes substantially less work because you don't have to constantly redo models to answer new questions.

Yeah, we've all been there -- production DBs aren't typically architected to store historical data. But we've found in practice that the data sources you most care about do have it. Page views, emails sent / received, completed orders, etc. all have timestamps. And for some things you don't need it. If you wanted to do a query with all customers who are VIPs, you wouldn't need a 'became VIP' activity. Adding is_VIP as a feature to the customer in the activity stream works too. Generally if you can do an analysis the more traditional way then you should already have the data to do it in Narrator too.

Sure, star schemas are the way of doing things and this is a new approach. But the efficiency gains realized by our own data scientists are enough to where they wouldn't go back -- it warrants the investment in learning it. Our challenge as a business will be how to convince others of that as well.

What we mean by single source of truth is that data is internally consistent - each term is defined once. In your scenario you'll have a single 'completed order' activity with the total order amount. If you want to add shipping cost that's fine -- add a 'added shipping to order' activity with the cost in it. Do the same with sales tax. Bob, Alice, and Jimmy can create reports with whatever activities they want. The crucial point is by making those reports they're not defining a new model. They're just combining activities. Since all tables are generated straight from the activity stream a future analyst won't use a materialized view based on Bob's data to build a new report -- they'll build it straight from the original activities.

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
No, we mainly focus on data modeling on top of a warehouse and deep analyses. Our customers are often data analysts, data scientists, or engineers.

That being said we would love to partner with a CDP like mixpanel and amplitude to have marketers and product people get quick insights using the data that is modeled and cleaned by the data team.

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
Great to hear from someone who also built this themselves!

As far as flexibility beyond 11 columns: I'd love to know your use case.

We do support additional metadata on each activity with what we call enrichment tables.

Some events are going to need more metadata -- a page view would want to have the actual page, the five UTM parameters, referrer, etc, which is more than the 3 fields of metadata we store on the activity stream.

So we also support creating additional tables to add metadata to each activity. Each row requires a unique activity id and its timestamp and can an unlimited number of additional columns.

We'll then automatically join that table into the activity stream when queries need it.

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
Sorry for not making that clear :). We're a SaaS product.

You can check out our pricing page here https://www.narrator.ai/pricing

The initial consultancy approach helped us build out the product. Once we could show internally that it made us far faster to analyze data we were ready to launch.

cedricd··on Launch HN: Narrator (YC S19) – a data modeling platform built on a single table
Also replying since I wrote this up :)

- activity_id : a unique identifier for the row

- activity : the type of activity (eg 'page_view')

- timestamp : time the activity happened

- customer : the unique customer identifier

Metadata columns

  Three columns for any info we'd like to add to an activity. Eg for a purchased product activity it could be product name. 

 - feature_1
 - feature_2
 - feature_3

 - revenue_impact : the amount of money related to this activity. A completed order activity would have this
 - link : a hyperlink related to the activity ('ticket submitted' might have a link to the ticket in Zendesk)
Additional customer identifier - source and source id are used when you're not entirely sure who the customer is. For example, a 'page view' activity wouldn't know the actual customer, but might have a unique identifier. So the source could be 'segment.io' and source_id could be their generated uuid

- source

- source_id

← PreviousPage 3 of 6Next →