HNHacker News
TopNewBestAskShowJobs

bitsondatadev

73 karma · joined August 31, 2020

U.S. Marine turned software engineer turned developer advocate.

https://linktr.ee/bitsondatadev

Disclaimer: I am head of Developer Relations at Tabular and contributor to Apache Iceberg and Trino.

submissionscomments
bitsondatadev··on Ask HN: Tired of software career. What now?
Is there any market or industry that you would like to work in? Like the domain itself is interesting to work in?
bitsondatadev··on Elasticsearch SQL
Trino is a nice query engine solution if you want to not just run SQL on Elasticsearch but also want to be able to join data from Elasticsearch with data in other systems Trino supports. It also supports raw elasticsearch queries that are serialized back into Trino data types.

https://trino.io/docs/current/connector/elasticsearch.html

bitsondatadev··on [dead]
Check out Trino Summit 2022 to learn more about the federated query engine and "Federate 'em all". Talks include Lyft, Shopify, Astronomer, Starburst, and much more to be announced in the coming weeks!

Trino (formerly PrestoSQL) is a query engine that originally aimed to replace the Hive runtime and grew into a powerful federated query engine. It can connect to Snowflake, BigQuery, Hive, Oracle, Mongo, Iceberg, Elasticsearch, and many more. This enables powerful joins across your platform from one location.

bitsondatadev··on Apache Iceberg creator, Ryan Blue, on Trino and Iceberg
In this episode of Trino Community Broadcast, we interview Ryan Blue to discuss the latest innovations in Apache Iceberg, why choose Apache Iceberg, latest Trino innovations and much more.

Disclaimer: I am a Trino contributor.

bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
Yeah, though a lot of Fivetran customers are likely the type that would go all in on paying for a conventional data warehouse where people using open source stacks may be the ones that are using open ingestion alternatives.

We see a pretty even mix from the Trino/Starburst lens. Bigger companies like to mix and match.

bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
Also, you can keep track of all the BQ progress here: https://github.com/trinodb/trino/issues/6867
bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
For managing difrerent workloads, check out this blogs and this videos from Shopify, Salesforce, Goldman Sachs, and Electronic Arts, respectively:

- https://engineering.salesforce.com/how-to-etl-at-petabyte-sc... - https://shopify.engineering/faster-trino-query-execution-inf... - https://trino.io/episodes/33.html - https://www.youtube.com/watch?v=-5mlZGjt6H4

All use the Lyft "Presto but really Trino"-Gateway project to run different clusters to handle various workloads. They go into various details for how this is achieved.

https://github.com/lyft/presto-gateway

Regarding the Trino/Presto split. I recommend looking at this blog to better understand why these two communities aren't mergeing. TL;DR Presto is a Facebook-driven project that mainly considers running on the Facebook infrastructure. Trino is community-driven that works on running well with all clouds and common infasturcture in the Trino community which is why you see a higher velocity there.

https://trino.io/blog/2022/08/02/leaving-facebook-meta-best-... https://trino.io/blog/2020/12/27/announcing-trino.html

Soon we anticipate that Trino will become the common name in the community space but we'll always love the origins of the Trino project being Presto.

bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
Clickhouse is a realtime system where Trino is a batch-oriented system. There are tradeoffs for doing realtime vs batch.

Realtime is generally more expensive to run as you process every individual row as it comes, batch is when you can deal with minute latency and want to handle a lot of data in chunks.

Trino is also a query engine rather than a database and it connects to many different systems: https://trino.io/docs/current/connector.html

It also happens to connect to Clickhouse and it's very common that people will use Trino to query clickhouse realtime data and join it with data in big query, an object store data lake, or Snowflake: https://trino.io/docs/current/connector/clickhouse.html

bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
Thanks for the shoutout! :)

If you want to get started with Trino, here's a repo I created to do so: https://github.com/bitsondatadev/trino-getting-started

bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
If you want to try an SaaS Athena alternative that's backed by Trino you can check out Starburst Galaxy: https://www.starburst.io/platform/starburst-galaxy/

Full disclosure I work at Starburst.

bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
Check out this PR. I believe we may have tackled this one but you'd need to try it out on Trino: https://github.com/trinodb/trino/pull/1415
bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
I mean, the hive migration path is one thing. Now that Iceberg is taking over the old Hive model, data lakes are all the rage again.

The other thing I would say is that Trino and Presto are not one-trick ponies or just hive replacements. There's also the ability to query across multiple systems that is, to me, the feature that future proofs a lot of architectures. It inherently frees you up to fiddle with your data in different systems but keep the access to that system in one location.

bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
<3
bitsondatadev··on The SQL query engine Trino (formerly PrestoSQL) recaps a decade of innovation
btw, if you want to know the backstory on why Presto is now called Trino, here's the article:

https://trino.io/blog/2022/08/02/leaving-facebook-meta-best-...

bitsondatadev··on Near Real-Time Ingestion for Trino
Ah this makes sense. This is similar to what Elasticsearch shards do.
bitsondatadev··on Near Real-Time Ingestion for Trino
Why is this just near real-time? If we're talking about Flink + Kafka, isn't that considered really real-time?

Maybe near real-time is just considering ingestion and query latency?

bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
Yup, I would say Athena is the most commonly used version of Trino/Presto out there.

Also, if anyone is reading this wondering which version Athena comes from, it's a release that was common to both Presto and Trino before the fork. However, more recently we've worked with the Athena team on getting the newer Trino features that aren't in Presto today in Athena. So it's starting to become more based on Trino.

bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
It’s a query engine that can talk to multiple databases and run join queries across tables from multiple sources. It also runs as a faster alternative to Hive.
bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
Starburst employee here so take my view for what it’s worth. A few points:

Starburst spun out from some contributors working on Teradata’s Presto distro in 2017.

The Presto creators left Facebook in late 2018 and worked on building the new community and fork for most of 2019. They joined Starburst around a year later.

I asked these same questions in my last company and it seems like nothing was correlated between starburst and the new fork. After talking with Martin T. (One of the creators of Presto) about it and a few other long time contributors and my concerns went away.

I ultimately started to contribute to Trino and over the years I’ve only seen a community dedicated to making a successful query engine. You can see that by the activity and growth of the project.

Feel free to remain a skeptic but the best way to know is joint the community and make your own assumptions. :)

https://trino.io/slack.html

bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
That’s exactly the point of the post too. Sometimes business and open source align and other times they do not. Trino was created as the project and community the puts community and individuals that contribute as the owners.

It’s not bad or evil or wrong that Facebook does this…it just is.

bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
The latter tree-no like neutrino and the website is https://trino.io lol.
bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
Yeah, I think that was one of the reasons why Facebook enforcing the trademark ended up being a blessing in disguise. It at least made the forks clearer. Now Trino is gaining more momentum but it takes a while for the brand recognition to set in I suppose.
bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
Hah! It may have very well looked like that given the time constraints they were working with.
bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
Facebook has their own internal reviews for their own internal projects and forks of open source projects. I think the issue came down to Facebook adding anyone working on the project to have the ability to merge code to the open source project.

Typically if a company wants to contribute to open source, you have to create a PR and have it be reviewed and merged by the maintainers of the open source project on top of internal review. Facebook management decided to circumvent that process.

bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
I'm curious to know if this is anyone's first time hearing about Trino?
bitsondatadev··on Leaving Facebook/Meta was the best thing we could do for the community
Here's the renaming article that clarifies this...

https://trino.io/blog/2020/12/27/announcing-trino.html

Months after this consolidation, Facebook decided to create a competing community using The Linux Foundation®. As a first action, Facebook applied for a trademark on Presto®. This was a surprising, norm-breaking move because up until that point, the Presto® name had been used without constraints by commercial and non-commercial products for over 6 years. In September of 2019, Facebook established the Presto Foundation at The Linux Foundation®, and immediately began working to enforce this new trademark. We spent the better part of the last year trying to agree to terms with Facebook and The Linux Foundation that would not negatively impact the community, but unfortunately we were unable to do so. The end result is that we must now change the name in a short period of time, with little ability to minimize user disruption.

bitsondatadev··on What the Heck is a Data Mesh?
"Do you happen to have some reference which would help me understand how e.g. cross-domain joins between large datasets can be optimized with this approach?"

Check out the original Presto paper (To be clear Trino is formerly PrestoSQL). https://trino.io/Presto_SQL_on_Everything.pdf. Also check out the definitive guide https://trino.io/blog/2021/04/21/the-definitive-guide.html for an even deeper dive.

"(incl. the numerous COTS and legacy applications in any "normal" enterprise)"

Funny you bring this up. In my last company we had a couple legacy apps that we didn't even have the code for any more. Just the artifacts (the license for these services were expiring and we were just waiting to kill them off). For the year or so that we had, we were able to write a simple Trino connector to pull these values through the API, and represent them as a table to Trino. And that's the point to your last question. It's the fact that you can meet the legacy code, or RDBMS database, or NoSQL database, or whatever the heck all these different teams own and get access to it in one location without migrating data and maintaining pipelines.

"I have low hopes of getting the data mesh standards for governance ending up being adopted either."

My view is that data mesh should be opt-in and flexible. Each team can pick and choose what portion of their data is modeled and exposed. If a team models their data in a way that doesn't align with some central standard, then they just need to create a view or some mapping that exposes their internal setup to match the central standard. Per the data mesh principle of federated computational governance, each team has a seat as to what these standards are. There can be some strict standards, and some that are more open. Teams/Domains only need to opt-in or concern themselves with the standards that are being requested of their data by the consumers. All the rest they can opt out.

I personally think avoiding company-wide standards is likely the best way to approach adoption. You basically have a list of well documented standards that can easily be searched and understood by consumers (analysts, business folks, data scientists, etc..) and it's up to the domains/consumers to negotiate standards on more of an adhoc basis. Therefore participating in the data mesh isn't some giant meeting everyone needs to join to grow consensus around. It can be much more distributed and less invasive.

bitsondatadev··on What the Heck is a Data Mesh?
Agree to disagree then I guess.

"When it comes to performance, I was referring primarily to e.g. cross-domain queries. That continues to be challenging in data virtualization / federated query engines."

Have you tried Trino lately? Data virtualization like Denodo still relies on moving data back between engines to execute a query, Trino pushes the queries down to both systems and processes the rest in flight. The fact that you use both of those interchangeably makes me think you may not have tried it.

"Data mesh targets analytical workloads - and surely you do not suggest e.g. hooking Trino up directly to operational, OLTP databases?" Not operational data. Are you saying teams aren't allowed to store immutable data in PostgreSQL?

"Not only are the domain teams free to build the aforementioned pipelines and model data in any way they see fit"

You will organically see teams build data infrastructure with different tooling. This happens when you don't tell a team they have to use a DW to play the analytics game. I have never seen a company that has multiple teams that organically land on using the same tech. So naturally (that is without forcing the to use one substrate i.e. DW) you will have different teams (or domains) using different databases. This is why we have DW to begin with. They were literally created to be a copy of domain data that naturally lied in many operational databases.

I get that in an ideal world, there would be some magical one size fits all solution that everyone would just use. However, that system doesn't exist. It's certainly not the DW. DW can be one of those solutions, just not THE one.

bitsondatadev··on What the Heck is a Data Mesh?
You say, “one cannot argue performance of a data warehouse” but that’s precisely the issue with a DW. DW requires a lot of work to move data from the way that domains model their data on the service layer to how data is modeled in a central DW. You have to wait for data to become live to even begin running analysis on it. Setting up and worst of all maintaining pipelines is an expensive undertaking in both time and money.

It’s not to say the DW is bad and never the solution. The problem is making it the only solution and not providing domains the flexibility to model data the way they need it. You say it’s more complex to manage but that’s the idea behind data mesh, you don’t manage that part, the team with their domain knowledge and data solution does. They can make it as simple or complex as they want internally but if they follow the standards to play in your data mesh who cares? Not your problem. For example say a domain needs realtime data analytics and use something like Druid to store their data. That’s fine. If they want to play in the data mesh you’ve provided, they just need to follow the rules in their data model, but they don’t need to use a cloud DW to do that.

You can’t argue that avoiding the copying of terabytes of data a day from a domain to a DW is more performant than adhoc analysis (MB to GB) of that data. Why move or copy a dataset when you don’t need to? Why force domains to use any solution that’s not actually solving their domain problem?

bitsondatadev··on What the Heck is a Data Mesh?
Not central to the main ideas of this article, but if you want to have a data mesh that is self-service, why force folks to use a particular storage medium like a data warehouse? That still requires centralization of the data.

Why not instead have a tool like Trino (https://trino.io) that allows you to let different domains use whatever datastore they happen to use. You still would need to enforce schema, but this can be done in tools like schema registry as mentioned in the article along with a data cataloging tool.

These tools facilitate the distributed nature of the problem nicely and encourage healthy standards to be discussed and the formalized in schema definitions and catalogs that remove the ambiguity of discourse and documentation.

Nice example is laid out in this repo of how Trino can accomplish data mesh principles 1 and 3 (https://github.com/findinpath/trino_data_mesh).

Page 1 of 2Next →