2,378 karma · joined April 25, 2010
My bet is that it will be much more widely adopted in the coming years. There are more tools and systems that accept and return Parquet files for bulk data (see most data warehouses). And there are libraries in most languages that are improving every day for working withe the format. Although, I'd obviously believe that as one of the creators of InfluxDB IOx :)
The format itself is, I think, a real pain to work with without libraries to hide its complexity. But the advantages are big enough that people are building those libraries to make it more broadly adoptable.
Right now we're focused on our cloud offering. We'll have official open source releases and documentation in the future.
YC changed my life too and I can't imagine it in better or more capable hands.
Great work to the Arrow team!
Inspiration: Free Solo (watching someone with singular focus that achieves excellence, it's absolutely incredible)
Tech: meh, not sure much does it here. Tons of entertaining stuff, but not sure what else. Sorkin is good I guess: Steve Jobs and The Social Network.
I think it's a really exciting time for new OLAP systems because of Arrow, Rust, and the rise of object store + ephemeral compute for analytical and time series data.
P.S. Hey Todd ;)
My view of the CTO role is that it's focused on more long term efforts or the product generally (in more technical companies). So that's what I've been doing for much of this year. With the caveat that I'm also the founder and on the board so there is still about 10 hours a week of executive meetings/management, reviewing of other people's writing & product efforts and working with our biggest customers and prospects.
All that being said, I'm not just working on this because it's what I want to do and it's what excites me (which it does). I also think it's my highest point of leverage within the company. Very few people have the same view and in depth experience in this problem domain. Most of the people I know of that do are either founding and running competitive offerings or they're working on similar projects at Google, AWS, or Azure and getting paid significantly more than we can afford to pay a single engineer.
I think this is one of our most important efforts right now and the best way I know to make it successful is to be working on it in depth. I'm not the best programmer on the team or the smartest, but I have some depth of experience that gives me a clear vision on what it should be able to do and what tradeoffs we should be making.
My role on this effort is basically as the product person and the tech lead. In this case the tech lead isn't the manager of people on the team. We try to have a multiple paths for advancement in engineering and one of them tech focused individual contributor.
Of course, I view this project as necessary, but not sufficient for our overall success. That's why most of our engineering team is working on our legacy products and the continued forward development of the overall platform.
However, for large scale analytics on huge data sets where you're scanning all of it, we'll likely push you to EMR or something like that. The nice part is that those big data systems can execute directly against the Parquet files in object storage that form the basis of our durability.
This is part of the bet we're making compatibility with a larger data processing ecosystem.
Almost all of our use cases fall within what you can compute in memory on a single system (which is pretty much up to 1TB).
We'll be publishing crates for parsing InfluxDB Line Protocol, Reading InfluxDB TSM files (for conversion into Parquet and other formats), and client libraries as well.
The entire project itself will also be published as a crate. So you can use any part of the server in your own project.
Its architecture is dramatically different than InfluxDB. You can do a comparison of the design goals. Read the post this thread refers to. I think you'll find it has very different goals than Postgres, an OLTP database, and Timescale, which is built on top of it.
So some parts of Timescale are under actual Apache 2 and some parts are under a proprietary source available license. I'm not sure what the LOC of which is which, or how it's actually organized in your repo. I'll leave it up to your potential users to try to figure out which and disentangle what parts are actually open.
As I recall, AWS very publicly forked Elastic because of this very same type of confusion. The difference is that if AWS were going to fork your project, they'd just fork Postgres, which is the real open source software that you're benefitting from.
If I were building an developer focused analytics, monitoring, or data analysis product, I wouldn't do it on top of Timescale because some parts of your codebase most definitely prevent that through your license. But that's me.
For users within large organizations, they're likely not able to use the software without approval from their legal department because it doesn't fall under any open source license.
Like I said, whether you care is really case dependent.
Meanwhile, InfluxDB IOx has a very different set of goals than Postgres. It's not an OLTP (transactional) DB and never will be. It's firmly targeted at OLAP and real-time OLAP workloads.
That means we can do things like optimize for running on ephemeral storage with object storage as the persistence layer. It'll have fine grained control over replication, how data is partitioned in a cluster, and where data is indexed, queried, queued for writes and more. Push and pull replication, bulk transfer, and persistence with Parquet. This last bit means you get integration with other data processing and data warehousing tools with minimal effort.
It'll also support Arrow Flight which will give it great integration into the data science ecosystems in Python and R.
Right now, InfluxDB IOx is really too early to do any real comparison on actual operation. We're putting this out now so that people can see what we're doing, comment on it, and maybe even contribute. We think it's an interesting approach where no single item is completely novel, but the composition of everything together makes it an entirely unique offering in open source.
Edit: one other thing I forgot to mention. InfluxDB IOx is open source, Timescale isn't. For some that matters, for many it doesn't. Depends on your use case.
It also defines rules for replication (push and pull), subscriptions to subsets of the data, and processing data as it arrives via a scripting engine.
Some of these features will arrive before others. Right now we want to make it work well for the data InfluxDB is currently good at (metrics, sensor data) and also work well for high cardinality data like events and tracing data.
Gandiva bindings is definitely something we should look into, but I'm guessing there's much lower hanging fruit within DataFusion in terms of optimizing, particularly for our use case.
We're also building on top of the work of a bunch of others that have built Arrow and libraries within the Rust ecosystem.
When the open source GAs, I don't really know. But we're doing this out in the open so people can see, comment, and maybe even contribute. Who knows, maybe after a few years you'll be a convert ;)
Rust has the memory control we're looking for with the safety of a higher level language. For programming that requires concurrency, as most server software including this project does, it's fantastic as the compiler will prevent data races. The way Rust has you deal with errors also makes it easier to write correct code at compile time rather than finding problems at runtime when they can be harder to track down.
Most of what I wrote here a few years ago still feels true: https://www.influxdata.com/blog/rust-can-be-difficult-to-lea...
I also talk about using Rust in the announcement blog post here: https://www.influxdata.com/blog/announcing-influxdb-iox/
It's one of the reasons I like being all in on Arrow. Why do everything ourselves when a ton of other smart people are working on this too?