HNHacker News
TopNewBestAskShowJobs

georgewfraser

2,609 karma · joined October 28, 2012

fivetran.com

twitter.com/frasergeorgew

george at fivetran dot com

submissionscomments
georgewfraser··on Charts built for Chat
Yes this is a dbt product - will be in the dbt cli soon, we shipped it initially as a separate tool while it’s in preview.
georgewfraser··on AWS Acquires DuckLabs
AWS is a great home for DuckLabs. They just want people to use more compute and storage, so they have a degree of technology-neutrality. This is the key to allowing DuckDB to continue to grow in whatever directions are natural, rather than being warped by some kind of walled garden data platform strategy.
georgewfraser··on Does anyone run Postgres without PgBouncer?
What pgbouncer does is indeed core functionality. Compare Postgres to MySQL and sql server, where analogous standalone connection pools are rarely used. The fundamental reason pgbouncer needs to exist is Postgres’ utterly retrograde design. Other examples: xid wraparound, conflict with recovery, lack of undo space.
georgewfraser··on Claude Tag
Slack is such a simple product, and is so strategic as an interface for AI, Anthropic has to be considering building their own Slack. Hopefully this tag approach is an MVP and it proves the potential of workplace messaging as an interface for AI, but also the limitations of relying on the extension points they salesforce chooses to provide, and it turns into a full fledged slack competitor from anthropic soon.
georgewfraser··on Anthropic, please make a new Slack
I assure you I wrote it myself
georgewfraser··on Anthropic, please make a new Slack
Claude-in-Slack is a big enough feature to overcome the slack-connect network effect. Openness is absolutely key! I wrote this post because I hoped that if Anthropic is already planning to do this I might be able to influence them to make open-data part of the plan. But openness by itself isn't a big enough feature to get users.
georgewfraser··on Anthropic, please make a new Slack
Also true! The most important thing is that the NewSlacks commit to interoperability. I think Anthropic has a special opportunity to lead the way here, because they have a track record of standing by their principles to an extraordinary degree.
georgewfraser··on Anthropic, please make a new Slack
You can only access public channel data, you can't even access that at scale, and Claude needs to be more natively integrated in ways that Slack will never allow.
georgewfraser··on Timescale Is Now TigerData
I talked to the timescale CTO at pg conf a few years ago and asked him what timescale does differently than a standard columnar database that makes it better suited for time oriented data. He said a bunch of things and I said “but columnar databases do those things.” Then he got mad at me.

I guess it’s just another columnar dbms after all?

georgewfraser··on DuckLake is an integrated data lake and catalog format
They make a really good criticism of Iceberg: if we have a database anyway, why are we bothering to store metadata in files?

I don’t think DuckLake itself will succeed in getting adopted beyond DuckDB, but I would not be surprised if over time the catalog just absorbs the metadata, and the original Iceberg format fades into history as a transitional form.

georgewfraser··on Fivetran to acquire Census
This is exactly right. We even went so far as to build a proof of concept internally, and the technical challenges are just very different. The simplest way to explain it is that Fivetran connects a skinny pipe (APIs) to a fat pipe (databases) while Census connects a fat pipe to a skinny pipe.
georgewfraser··on Vertical Sharding Sucks
I am generally a huge vertical sharding skeptic but there are special cases where it is beneficial. If you have a simple query pattern on one table that represents a big fraction of your entire workload you can put it into its own instance and it becomes much easier to monitor. It’s easy to see why vertical sharding is sometimes the right answer by inverting the decision: should we put two unrelated large applications on the same instance? Obviously not, there is no benefit and ops becomes more difficult.
georgewfraser··on The Evolution of SRE at Google
Like so many things from Google engineering this will be toxic to your startup. SREs read stuff like this, they get main character syndrome and start redoing the technical designs of all the other teams, and not in a good way.

This phenomenon can occur in all “overlay” functions, for example the legal department will try to run the entire company if you don’t have a good leader who keeps the team in their lane.

georgewfraser··on Why is it so hard to buy things that work well? (2022)
In general, I absolutely agree with you. It’s basically an instance of “the customer is always right”: if a smart customer can’t get our product working, there is a problem with the product. But this post made a much bolder (and wrong) claim: “the product has a number of major design flaws that mean that it literally cannot work”.
georgewfraser··on Why is it so hard to buy things that work well? (2022)
I have some insight into this because this claim is about my company Fivetran:

“…relies on the data source being able to seek backwards on its changelog. But Postgres throws changelogs away once they're consumed, so the Postgres data source can't support this operation”

Dan’s understanding is incorrect, Postgres logical replication allows each consumer to maintain a bookmark in the WAL, and it will retain the WAL until you acknowledge receipt of a portion and advance the bookmark. Evidently, he tried our product briefly, had an issue or thought he had an issue, investigated the issue briefly and came to the conclusion that he understood the technology better than people who have spent years working on it.

Don’t get me wrong, it is absolutely possible for the experts to be wrong and one smart guy to be right. But at least part of what’s going on in this post is an arrogant guy who thinks he knows better than everyone, coming to snap conclusions that other people’s work is broken.

georgewfraser··on Show HN: Peerdb Streams – Simple, native Postgres change data capture
You can partition your BigQuery table however you like and Fivetran will leave it in place. I don’t think there’s any benefit to partitioning the staging table.
georgewfraser··on PostgreSQL reconsiders its process-based model
I wonder if it would be easier to create a C virtual machine that emulates all the OS interaction, then recompile Postgres and the extensions to run on this. Perhaps TruffleC would work?

https://dl.acm.org/doi/10.1145/2647508.2647528

georgewfraser··on Tesla finally breaks and offers round steering wheel on Model S/X
I have an S with a yoke and prefer it to a round steering wheel. 95% of the time it’s better: you can see the entire dashboard and your body can feel the angle of the wheel instinctively. 5% of the time it’s worse, during low speed maneuvers.
georgewfraser··on Database drivers: Naughty or nice?
The world would be a better place if database drivers were completely abandoned as a way for clients to connect to databases. A standard API, implemented by multiple vendors, is a vastly preferable solution. Arrow Flight is an example of this.

https://arrow.apache.org/blog/2019/10/13/introducing-arrow-f...

georgewfraser··on Show HN: Cozo – new Graph DB with Datalog, embedded like SQLite
Licensing under AGPL will make it hard for any startup to use Cozo. Lawyers always ask about AGPL in venture financing diligence and it is considered a red flag. You can argue that they are wrong, the linking exception and so on, but you’re basically shouting into the wind.
georgewfraser··on Why is Snowflake so expensive
The core claim of this article, that Snowflake doesn't implement optimizations that would reduce usage, is not true. Search optimized tables, partitioned tables, and per-second billing are all counterexamples.
georgewfraser··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
Yeah I think that is the key question: will data lakes become the dominant paradigm? There is certainly a lot of talk around them, though I see a ton of companies are still just going all in on a conventional data warehouse, but they tend not to talk about it because it’s not a new or interesting thing to do.
georgewfraser··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
The thing I wonder about with Presto and to a lesser extent Spark is, how many of their users adopted this tool because it was an easy migration path from Hive, and how many of those users will eventually re-platform to something else?
georgewfraser··on Alzheimer’s amyloid hypothesis ‘cabal’ thwarted progress toward a cure (2019)
There is no better way to get shut out as a scientist than to disprove the career-making findings of the important people in your field. One of the side effects of this is the “discussion does not match the results” paper. It’s routine to read papers where the discussion blatantly contradicts the papers own results. What is happening is the authors are disclaiming their own findings in order to get past hostile reviewers, betting that astute readers will notice the contradiction, ignore the discussion and draw their own conclusions from the results.
georgewfraser··on Ask HN: Is DBA still a good job?
I hope not. We need to hire a Postgres DBA at Fivetran, we have more or less a single Postgres database with all our state, and it’s become clear that we need a full time person to optimize it.
georgewfraser··on California tech billionaire launches Senate campaign to take on Tesla
Going to hazard a guess that he’s mad that Tesla isn’t a customer of his embedded-software-testing product https://www.ghs.com/
georgewfraser··on How we upgraded our 4TB Postgres database
This is such a huge problem. It's even worse than it looks: because users are slow to upgrade, changes to the database system take years to percolate down to the 99th percentile user. The decreases the incentive to do certain kinds of innovation. My opinion is that we need to fundamentally change how DBMS are engineered and deployed to support silent in-the-background minor version upgrades, and probably stop doing major version bumps that incorporate breaking changes.
georgewfraser··on How we upgraded our 4TB Postgres database
It is amazing how many large-scale applications run on a single or a few large RDBMS. It seems like a bad idea at first: surely a single point of failure must be bad for availability and scalability? But it turns out you can achieve excellent availability using simple replication and failover, and you can get huge database instances from the cloud providers. You can basically serve the entire world with a single supercomputer running Postgres and a small army of stateless app servers talking to it.
georgewfraser··on We’ve got a science opportunity overload: Launching the Wolfram Institute
Wolfram wolfram wolfram wolfram! Wolfram’s wolfram wolfram, wolfram wolfram. Wolfram wolfram? Wolfram wolfram wolfram wolfram.
georgewfraser··on Timescale raises $110M Series C
I would love to run a time-series benchmark against a good column store like Snowflake to see if purpose-built time series databases are actually faster. I have a sneaking suspicion the time scale databases are just reinventing the column store, and that an appropriate non-sabotaged benchmark would show this.
Page 1 of 11Next →