HNHacker News
TopNewBestAskShowJobs

cedricd

908 karma · joined October 5, 2012

cofounder at Narrator (narratordata.com)

YC Badge: 0x31430d28FAf744B9012E8116F80d58D735180c79

submissionscomments
cedricd··on We Can Finally Do Away with the Accursed Beep-Beep of Heavy Machines
I live in NYC and on many heavy machinery I hear a white noise-type sound when they back up -- just like the article describes.

It's still pretty noticeable, but the ambient noise is already pretty high in the city, so it does fade quickly. Not sure how it would be in a quieter environment, but it's much less annoying than a beep for sure.

cedricd··on PostgreSQL's Missing DateDiff Function
The point of the function is that it doesn't matter you've crossed the year boundary when counting days.

A naive implementation would accidentally get tripped up by the year -- extracting the 'day' part of a timestamp, as an integer, gives you the day from the start of the year. So on one side of the year boundary you have 365 and on the other you have 1. The way to do it correctly is to multiply the day in the year by the year itself so that a '1' on a later year is a bigger number.

And of course grouping by year isn't always what you want to do :)

cedricd··on PostgreSQL's Missing DateDiff Function
These are used pretty often when doing data analysis. It's a simple way to group things together.

For example, find all users who have have been active at least 10 days. This provides a simple way to do a complex operation (if a user signs up Dec 29, a naive implementation won't catch that their 10 days is in January of the next year).

cedricd··on PostgreSQL's Missing DateDiff Function
Yep. That's exactly why this function isn't trivial to write. Boundaries are not obvious and native pg functions (insofar as I'm aware of them) don't do these kinds of diffs.

Semantically the code has to diff days (for example) but be aware that you crossed a year boundary (or any other).

cedricd··on PostgreSQL's Missing DateDiff Function
Those postgres docs are usually so dry but this is great

> The first century starts at 0001-01-01 00:00:00 AD, although they did not know it at the time

cedricd··on PostgreSQL's Missing DateDiff Function
I've always wondered myself. Maybe it's because it's mostly useful for analytical workloads instead of operational ones.

Redshift, famously based off Postgres, chose to implement it

cedricd··on Windows 11 Officially Shuts Down Firefox’s Default Browser Workaround
I've run all of them over the years (progressively upgrading like a good computer user). It's not even subtle. The good / bad dynamic is drastic.

And each 'bad' always brings to it a horrible UI change. Vista brought the window manager and those weird transparent windows and was generally ugly and buggy. Win 7 cleaned up the UI and made it flatter and simpler.

8 brought a full-screen start menu (!!). 10 went back to a 7-esque vibe (mostly).

11 is where we are.

cedricd··on Plasmic – A headless page builder
Has anyone used some of these in production? Tina, Builder, Storybloks, React Bricks, or this one? Curious what the experience is like.
cedricd··on Tell HN: Thanks to thehodge and littlewarden.com, this site is up today
So does that mean that YC is now a paid subscriber to the service? ;).

Very classy callout in any case. I love the story of a startup getting good press for doing something nice. Also this sounds like a really good case study for them to put up.

cedricd··on Commonly used idioms in the tech industry
Yeah, this one bugs me too. I try to use 'requests' instead.
cedricd··on The Rise and Fall of the OLAP Cube (2020)
This isn't widely known, but activity schema[1] is aiming for this space (disclosure - I co-founded the company supporting it).

It's a modeling approach that separates data modeling from data querying -- meaning that once data is modeled, it can answer any number of questions. The base building blocks of the data model can be combined at query time without having to figure out explicit foreign key joins.

[1] https://www.narrator.ai/activity-schema/

cedricd··on PostgreSQL's Missing DateDiff Function
If I'm understanding correctly you think the boundary crossing thing is weird. Like why would it say that Jan 1st 2021 - Dec 31 2020 is 1 year, when it's more like 1/365 of a year.

I'm not necessarily the best person to defend it, but I think it has a couple nice properties.

1. it's an integer, so can be used in group_by, comparisons, bucketing, etc

2. it aligns to commonly understood boundaries, which helps with the above.

Another way to look at the boundary issue is that it only matters when things are close. If you get a 1 and don't like it, drop down a unit (year -> months or days) instead.

Comparing years the way I did above is obviously a bit of an edge case

cedricd··on PostgreSQL's Missing DateDiff Function
Author here. I've used pg a ton in the past in production systems as an operational database. Never really missed having datediff.

Once we started using it as a data warehouse I noticed that function was missing -- most other data warehouses have it. Thought it would be worth providing an implementation since date stuff is annoying to do.

cedricd··on How to track users for analytics in a privacy-first, cookie-less future
Yeah, this advice looks targeted to companies that benefit hugely from targeting their users.

If I'm reading correctly it's basically saying 'once a user has identified themselves to you, then you can go back and figure out the steps they took before that'

As a person, if a company knows what I did right before I bought their product (say in that session) I think I'm ok with that. If they follow me onto other websites or other devices then that feels a lot more invasive.

cedricd··on How to track users for analytics in a privacy-first, cookie-less future
The identifier on the urls isn't meant to identify the actual user I think.

If you look at the examples given they're more like identifiers to something else -- an order id or subscription id.

Wouldn't tracking something like an order (but not the user directly) be ok with GDPR?

cedricd··on Bitcoin Emits Less Than 2% of World’s Military-Industrial Carbon Emissions
I'm not sure that spinning that as a positive is a good idea. It's objectively a massive amount.
cedricd··on Shareable data analyses using templates
We've been running shareable / configurable data analyses in production for the past several years.

In the data analysis world this has always been seen as a bit of a pipe dream, because no one's data is structured in the same way. A simple calculation (like monthly recurring revenue) has to be rewritten for each company. We standardize data in such a way that an analysis built by a company can be shared with another.

cedricd··on Unsolicited Advice for Technology Writers (2014)
Thanks. I was thinking the same thing. The language is hard to read.

http://paulgraham.com/simply.html

cedricd··on Ask HN: Who is hiring? (June 2021)
Narrator | New York, NY | Full stack engineer | Full-time | Remote | https://www.narrator.ai

Narrator (YC S19) is an end-to-end data platform that models all data in a single time-series table. We're a totally new kind of data platform that makes data modeling and analysis extremely efficient.

Our stack is Python and React.

We're looking for a principal full stack software engineer to own some of our core systems.

Our user-facing app is built in React and allows customers to run a full end-to-end data system, including modeling, querying, and analyzing data.

Our backend system manages all api endpoints and powers the data infrastructure -- including interesting things like translating queries into SQL for multiple warehouses.

We're looking for someone highly technical to own one or both of these systems. We're still just getting started and have a lot to build.

https://www.workatastartup.com/companies/12598

cedricd··on Have you ever hurt yourself from your own code?
In other words: please go home and I'll call you.

'Head on home' is a real expression. Similar ones are 'head over there', 'head to ...', all of which just mean 'go to'.

cedricd··on What if remote work didn’t mean working from home?
Exactly. I wonder if this trend will create a meaningful turnaround for them. What a wild world.
cedricd··on Ethereum will use around 99.95% less energy post merge
That's a common take, but a well-designed carbon tax could be made revenue-neutral and non-regressive.

In other words, the amount it takes in can be given back to the average person -- in forms of tax rebates, investment in public transit, education, whatever. In an ideal world a carbon tax would have no adverse impact on lower and middle class people.

I should also add that in that world higher electricity cost due to carbon tax is a good thing -- it'll help carbon-neutral sources of energy to compete and replace the more dirty forms. Which is exactly what we want.

cedricd··on WeWork CEO Says Least Engaged Employees Enjoy Working from Home
This suggests a potentially disturbing trend -- companies or managers that will start implicitly punishing employees for working remotely.

My guess is that even at companies that officially support partial remote time employees will start to feel some pressure for taking advantage of it.

Adam Neumann once asked his executive assistant if she 'enjoyed her vacation' after coming back from maternity leave. Now they have a different CEO, obviously, but I can see the parallel -- any kind of situation that deviates from butts in seats at the office will be frowned upon.

cedricd··on Using PostgreSQL as a Data Warehouse
I think you're right but it's a bit out of scope. Hard to give generalizable advice around this I think.

What we personally do in practice is put everything into a single time-series table with 11 columns. [1]

1: https://www.activityschema.com/

cedricd··on Using PostgreSQL as a Data Warehouse
Yep. Fully agree. The point of the post wasn't to say that you should use PG as a data warehouse. Just that if it's what you've got available (for various reasons) that you can.
cedricd··on Using PostgreSQL as a Data Warehouse
Luckily Postgres' autovacuum works really well in normal workloads. If there's an even mix of inserts spread throughout time then it's probably best to just rely on it.

For data warehouses inserts can happen in bulk on a regular cadence. In that case it can help to vacuum right after. I'm not sure if it has a huge impact in practice.

cedricd··on Using PostgreSQL as a Data Warehouse
Also, what's your take? Do people use citus for analytical workloads as well as production at scale?

I'd assume yes, but I haven't personally used you guys. I'm just aware of you and broadly how you scale Postgres.

cedricd··on Using PostgreSQL as a Data Warehouse
Ahh! So sorry. Fixed it.
cedricd··on Using PostgreSQL as a Data Warehouse
I updated the blog post :)
cedricd··on Using PostgreSQL as a Data Warehouse
That's a great point.

This isn't really what you're saying, but citus [1] ships a distributed Postgres. A lot of the things they improve would help massively with analytical workloads actually.

1: https://www.citusdata.com/

← PreviousPage 2 of 6Next →