HNHacker News
TopNewBestAskShowJobs

pauldix

2,378 karma · joined April 25, 2010

Cofounder of InfluxData, company behind InfluxDB (W13). @pauldix on GH and Twitter.
submissionscomments
pauldix··on Free Dolly: First truly open instruction-tuned LLM
Yeah, if it's getting astroturfed and pushed up artificially that's definitely cause to flag it. It's just a shame that happened because this almost certainly would have landed on the front page on its own.
pauldix··on Free Dolly: First truly open instruction-tuned LLM
Shame that this is flagged, I think this is a really exciting development and was hoping to see the discussion around it. Open sourcing the fine tuning training set is a great building block. Will be exciting to see if others continue to build on this. More open source datasets, models, and evaluation frameworks will accelerate the development and adoption of LLMs. It adds more hackers to the mix building the core, rather than just the stuff at the edges (i.e. apps).
pauldix··on Querying Parquet with Millisecond Latency
I would guess that it's mostly because of tooling and library availability. There has been very little cross language support for reading and writing Parquet files. Mostly it was concentrated in Java and C++. But since it fell under the Arrow umbrella, that situation has been improving.

My bet is that it will be much more widely adopted in the coming years. There are more tools and systems that accept and return Parquet files for bulk data (see most data warehouses). And there are libraries in most languages that are improving every day for working withe the format. Although, I'd obviously believe that as one of the creators of InfluxDB IOx :)

The format itself is, I think, a real pain to work with without libraries to hide its complexity. But the advantages are big enough that people are building those libraries to make it more broadly adoptable.

pauldix··on IOx: InfluxData’s New Storage Engine
DataFusion is great, we're happy to be contributing to it. Also excited to see so many people around the world picking it up and contributing as well. With our development efforts on IOx, it's like a strong tailwind. But we put a ton of effort into helping manage community efforts (thanks, alamb! our developer on IOx that is also on the Arrow PMC).
pauldix··on IOx: InfluxData’s New Storage Engine
Hi, post author and founder of InfluxDB here. We're supporting Flux (our scripting and query language), InfluxQL (our original SQL like language), and SQL (specifically the Postgres dialect as that's what DataFusion supports). The query engine is DataFusion, which is part of the Apache Arrow project. We contribute to it significantly. So that's what's built in natively. We support Flux and InfluxQL through separate Go processes that use an API to connect to the core DB. Although we're working on native InfluxQL support (it's a Rust based InfluxQL parser that will yield DataFusion logical query plans).

Right now we're focused on our cloud offering. We'll have official open source releases and documentation in the future.

pauldix··on Welcome Home, Garry Tan
Congrats Garry! So well deserved. I feel so lucky to have had your help as a group partner when we went through W13 as Errplane (hah, good thing we figured out something to pivot to).

YC changed my life too and I can't imagine it in better or more capable hands.

pauldix··on Voltron Data grabs $110M to build startup based on Apache Arrow project
More good news for Apache Arrow and any projects building around the standard. Excited to see more tools, clients, servers, and everything in between support Arrow, Arrow Flight, Arrow Flight SQL, and Parquet! Over the next few years I hope we see these standards take over data science and data warehouses. At least that's what we're building towards at InfluxDB.
pauldix··on Apache Arrow Flight SQL: Accelerating Database Access
Oh, and for anyone interested in pitching in on the Rust implementation, there's an issue logged here along with some discussion: https://github.com/apache/arrow-rs/issues/1323
pauldix··on Apache Arrow Flight SQL: Accelerating Database Access
This is great and makes a ton of sense as a refinement to their work on Flight previously. InfluxDB IOx already supports Flight, but we'll almost certainly be updating to support Flight SQL before we launch. We've been thinking about adding Postgres wire protocol support, but this would be even better if enough downstream clients and projects end up adding SQL Flight support.

Great work to the Arrow team!

pauldix··on Ask HN: What are some technically inspiring movies, TV shows or documentaries?
Sales: Boiler Room, Glengary Glen Ross

Inspiration: Free Solo (watching someone with singular focus that achieves excellence, it's absolutely incredible)

Tech: meh, not sure much does it here. Tons of entertaining stuff, but not sure what else. Sorkin is good I guess: Steve Jobs and The Social Network.

pauldix··on Apache Arrow Datafusion 5.0.0 release
The new core we're building for InfluxDB (named InfluxDB IOx) uses Datafusion for query execution. We have multiple team members contributing to this release and we're super excited to be involved with it.

I think it's a really exciting time for new OLAP systems because of Arrow, Rust, and the rise of object store + ephemeral compute for analytical and time series data.

pauldix··on Grafana, Loki, and Tempo will be relicensed to AGPLv3
The Grafana Enterprise point is the most salient here. The answer is no, they wouldn't have been able to create closed source plugins for a commercial product. Only AGPL code. As I've stated before: copyleft licenses are an evolutionary dead end: https://www.influxdata.com/blog/copyleft-and-community-licen....

P.S. Hey Todd ;)

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
There is absolutely no point benchmarking it at this stage. One of the many reasons we're not bothering to produce builds right now. We'll let you know when it's time to even have a look as I'm sure you'll want to continue to compare it with Victoria.
pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Thanks :)
pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Not confrontational at all. In short, it's taken me years to be able to do this. Last year we finally got our engineering organization to a point where I no longer have any direct reports. Instead, everyone rolls up to our VP of engineering. In most mid-sized startups (like ours), the VP of engineering is the person that manages most of the process and generally makes sure the trains run on time.

My view of the CTO role is that it's focused on more long term efforts or the product generally (in more technical companies). So that's what I've been doing for much of this year. With the caveat that I'm also the founder and on the board so there is still about 10 hours a week of executive meetings/management, reviewing of other people's writing & product efforts and working with our biggest customers and prospects.

All that being said, I'm not just working on this because it's what I want to do and it's what excites me (which it does). I also think it's my highest point of leverage within the company. Very few people have the same view and in depth experience in this problem domain. Most of the people I know of that do are either founding and running competitive offerings or they're working on similar projects at Google, AWS, or Azure and getting paid significantly more than we can afford to pay a single engineer.

I think this is one of our most important efforts right now and the best way I know to make it successful is to be working on it in depth. I'm not the best programmer on the team or the smartest, but I have some depth of experience that gives me a clear vision on what it should be able to do and what tradeoffs we should be making.

My role on this effort is basically as the product person and the tech lead. In this case the tech lead isn't the manager of people on the team. We try to have a multiple paths for advancement in engineering and one of them tech focused individual contributor.

Of course, I view this project as necessary, but not sufficient for our overall success. That's why most of our engineering team is working on our legacy products and the continued forward development of the overall platform.

pauldix··on InfluxDB 2.0 Is Out
This blog post has more detail and screenshots. It'll probably be more interesting for this crowd: https://www.influxdata.com/blog/influxdb-2-0-open-source-is-...
pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
For now the execution is in-memory only. Over time we want to be able to execute against Parquet files on disk.

However, for large scale analytics on huge data sets where you're scanning all of it, we'll likely push you to EMR or something like that. The nice part is that those big data systems can execute directly against the Parquet files in object storage that form the basis of our durability.

This is part of the bet we're making compatibility with a larger data processing ecosystem.

Almost all of our use cases fall within what you can compute in memory on a single system (which is pretty much up to 1TB).

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
I'm not too familiar with their stuff, but I think in terms of approach, they're very different. This project doesn't aim to product materialized views, which I think is more of what naiad is for? I'm not sure about differences in terms of what problems they're trying to sovle.
pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Yes, absolutely. There are the Arrow libraries in Rust that we're already contributing to. DataFusion, for example is an SQL execution engine.

We'll be publishing crates for parsing InfluxDB Line Protocol, Reading InfluxDB TSM files (for conversion into Parquet and other formats), and client libraries as well.

The entire project itself will also be published as a crate. So you can use any part of the server in your own project.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
So if my users can't have DDL access, that means that they can't define the schemas for the analytics that they want to do? It only works if as the developer of an application I have a fixed schema that my users interact with?
pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
This comparison is vs. InfluxDB. This thread is about a new project called InfluxDB IOx. It's under development and we're not producing builds, so any kind of operational comparison would be very premature.

Its architecture is dramatically different than InfluxDB. You can do a comparison of the design goals. Read the post this thread refers to. I think you'll find it has very different goals than Postgres, an OLTP database, and Timescale, which is built on top of it.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
My post is about InfluxDB IOx, which is the project this thread is about. You're correct about InfluxDB having HA and clustering under a closed source enterprise license. If you read the post, I even mention this as a shortcoming of the project. One which we're hoping to rectify with InfluxDB IOx.

So some parts of Timescale are under actual Apache 2 and some parts are under a proprietary source available license. I'm not sure what the LOC of which is which, or how it's actually organized in your repo. I'll leave it up to your potential users to try to figure out which and disentangle what parts are actually open.

As I recall, AWS very publicly forked Elastic because of this very same type of confusion. The difference is that if AWS were going to fork your project, they'd just fork Postgres, which is the real open source software that you're benefitting from.

If I were building an developer focused analytics, monitoring, or data analysis product, I wouldn't do it on top of Timescale because some parts of your codebase most definitely prevent that through your license. But that's me.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
It's under a community license, which has restrictions. The limitations on derivative works and value added products or services are the ones that will create the most problem for people trying to build a business on it: https://www.timescale.com/legal/licenses

For users within large organizations, they're likely not able to use the software without approval from their legal department because it doesn't fall under any open source license.

Like I said, whether you care is really case dependent.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Timescale is built on top of Postgres, which is a row oriented database. They've built a kind of columnar layer on top of it, which is quite interesting. Because it's Postgres you get their full SQL support.

Meanwhile, InfluxDB IOx has a very different set of goals than Postgres. It's not an OLTP (transactional) DB and never will be. It's firmly targeted at OLAP and real-time OLAP workloads.

That means we can do things like optimize for running on ephemeral storage with object storage as the persistence layer. It'll have fine grained control over replication, how data is partitioned in a cluster, and where data is indexed, queried, queued for writes and more. Push and pull replication, bulk transfer, and persistence with Parquet. This last bit means you get integration with other data processing and data warehousing tools with minimal effort.

It'll also support Arrow Flight which will give it great integration into the data science ecosystems in Python and R.

Right now, InfluxDB IOx is really too early to do any real comparison on actual operation. We're putting this out now so that people can see what we're doing, comment on it, and maybe even contribute. We think it's an interesting approach where no single item is completely novel, but the composition of everything together makes it an entirely unique offering in open source.

Edit: one other thing I forgot to mention. InfluxDB IOx is open source, Timescale isn't. For some that matters, for many it doesn't. Depends on your use case.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Right now it's just a project in development, so nothing yet. But it's supposed to be for time series data. This could be metrics, events, or any kind of semi-structured data that fits into tables where you want to ask questions about it over time.

It also defines rules for replication (push and pull), subscriptions to subsets of the data, and processing data as it arrives via a scripting engine.

Some of these features will arrive before others. Right now we want to make it work well for the data InfluxDB is currently good at (metrics, sensor data) and also work well for high cardinality data like events and tracing data.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Ah yes, of course, hi Nevi :). Thank you again for all your work on the Rust implementation. We're obviously big fans.

Gandiva bindings is definitely something we should look into, but I'm guessing there's much lower hanging fruit within DataFusion in terms of optimizing, particularly for our use case.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
We'll be producing builds early next year. Those won't be anything we're recommending for production. Our goal is to have an early alpha in our own cloud environment by the end of Q1. I stress alpha. But we'll also have a bunch of tooling around it (which we've already built for the other parts of our cloud product) to backup data out of band, monitor it, shadow test it against production workloads, etc.

We're also building on top of the work of a bunch of others that have built Arrow and libraries within the Rust ecosystem.

When the open source GAs, I don't really know. But we're doing this out in the open so people can see, comment, and maybe even contribute. Who knows, maybe after a few years you'll be a convert ;)

pauldix··on InfluxDB IOx – The New Core of InfluxDB Built with Rust
So far working with Rust has been great. We haven't had to build any special build tooling. The build times are greater than Go, but we're able to keep things manageable by breaking the project into separate Crates (all still within the same repo and directory tree), which can take advantage of incremental compilation.

Rust has the memory control we're looking for with the safety of a higher level language. For programming that requires concurrency, as most server software including this project does, it's fantastic as the compiler will prevent data races. The way Rust has you deal with errors also makes it easier to write correct code at compile time rather than finding problems at runtime when they can be harder to track down.

Most of what I wrote here a few years ago still feels true: https://www.influxdata.com/blog/rust-can-be-difficult-to-lea...

I also talk about using Rust in the announcement blog post here: https://www.influxdata.com/blog/announcing-influxdb-iox/

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
It absolutely is and we've been contributing back. A even bigger amount of this is based on Wes McKinney's work on Arrow. Andy is great and he's been helpful as we've been working with DataFusion.
pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
They're heavily into Arrow. A few years ago they contributed Gandiva, an LLVM expression compiler for super fast processing. https://arrow.apache.org/blog/2018/12/05/gandiva-donation/

It's one of the reasons I like being all in on Arrow. Why do everything ourselves when a ton of other smart people are working on this too?

← PreviousPage 3 of 12Next →