HNHacker News
TopNewBestAskShowJobs

pauldix

2,378 karma · joined April 25, 2010

Cofounder of InfluxData, company behind InfluxDB (W13). @pauldix on GH and Twitter.
submissionscomments
pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
I think the best new projects are created by a small focused team. Adding too many people too early actually slows things down. But, of course, I'm biased.

The thing about this getting to production next year is that we're doing it in our cloud, which is a services based system where we can bring all sorts of operational tooling to bear. Out of band backups, usage of cloud native services, shadow serving, red/green deploys, and all sorts of things. Basically, it's easier to deploy a service to production once you've built a suite of operational tools to make it possible to do reliably while testing under production workloads that don't actually face the customer.

As for us rewriting the core of the database, that's true. But I think you're unrealistic about what the data systems look like in closed source SaaS providers as they advance through orders of magnitude of scale. Hint: they rewrite and redo their systems.

As for Grafna, MetricsTank was their first, Cortex wasn't developed there, and Loki and Tempo look like interesting projects.

None of those things has the exact same goal as InfluxDB. And InfluxDB isn't meant to be open source DataDog. That's not our thing. We want to be a platform for building time series applications across many use cases, some of which might be modern observability. It also doesn't preclude you from pairing InfluxDB with those other tools.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Users of InfluxDB 1.x can upgrade to 2.0 today. It's an in-place upgrade and just requires a quick command to make it work. Further, InfluxDB 2.0 also has the API for InfluxDB 1.x. We've been putting out InfluxDB 1.x releases while we've been developing 2.0. One of the reasons we waited so long to finalize GA for 2.0 was because we had to make sure there was a clean migration path and there is.

For our cloud 1 customers, they'll be able to upgrade to our cloud 2 offering, but in the meantime, their existing installations get the same 24x7 coverage and service we've been providing for years.

As for how this will be deployed, it will be a seamless transition for our cloud customers when we do so. Data, monitoring and analytics companies replace their backend data planes multiple times over the course of their lifetime.

For our Enterprise customers, we'll provide an upgrade path, but not until this product is mature enough to ship an on-premise binary that won't get a chance to get upgraded but for a few times a year.

The only difference here is that we're doing it as open source. They always do theirs behind closed doors. I'm sure most of our users and many of our customers prefer our open source approach.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Yeah, Parquet is awesome. One of the things we really want to do here is to push DataFusion (the Rust based SQL execution engine) to work on Parquet files, but to push down predicates and other things and operate on the data while it's compressed.

You pay such a high overhead marshalling that data into an Arrow RecordBatch. Best thing ever is to work with the Parquet file and not even decompress the chunks that you don't need. Of course, this assumes that you're writing summary statistics as part of the metadata, which we plan to do.

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Execution is vectorized, but that's Arrow really. We'd like this to be useful for general OLAP workloads, but the focus for the next year is definitely going to be on our bread and butter time series stuff.

That being said, Arrow Flight will be a first class RPC mechanism, which makes it quite nice for data science stuff as you can get data in to a dataframe in Pandas or R in a few lines of code and with almost zero serialization/deserialization overhead.

This isn't meant to be a transactional system. More like a data processing system for data in object storage. I'm curious what your need is there for OLAP workloads, can you tell me more?

pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Yes, this will likely be the case, but it's not a specific goal. We're focusing our efforts right now on the InfluxDB time series use cases we see most. More general data warehousing work isn't our focus, but we expect to pick up quite a bit of that along the way as we develop and as the underlying Arrow tools develop.
pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
Those two systems are designed to work with Prometheus style metrics, which are very specific. You have metrics, labels, float values, and millisecond epochs.

I'm not totally sure how they index things, but I would guess that it's by time series with an inverted index style mapping for the metric and label data to underlying time series. This means they'll have the same problems with working with high cardinality data that I outlined in the blog post.

InfluxDB aims to hit a broader audience over just metrics. We think the table model is great, particularly for event time series, which we want to be best in class for. A columnar database is better suited for analytics queries, and given the right structure is every bit as good for metrics queries.

pauldix··on InfluxDB IOx – The New Core of InfluxDB Built with Rust
InfluxDB creator and lead for InfluxDB IOx here. Happy to answer any questions people might have. In short, it's an in-memory columnar database with object storage for persistence. It's a bunch of other things as well, but that's the main thrust of the project.
pauldix··on InfluxDB 2.0
HN thread for that is here: https://news.ycombinator.com/item?id=25049253
pauldix··on InfluxDB is betting on Rust and Apache Arrow for next-gen data store
InfluxDB creator here. I've actually been working on this one myself and am excited to answer any questions. The project is InfluxDB IOx (short for iron oxide, pronounced eye-ox). Among other things, it's an in-memory columnar database with object storage as the persistence layer.
pauldix··on Apache Arrow, Parquet, Flight and their ecosystem are a game changer for OLAP
We use Arrow under the hood in our new scripting and query language, Flux. We're looking into adding Flight RPC to InfluxDB and there are some new internal projects that I'm looking at to use Parquet + Flight + Arrow all the way through. But this article was meant to just be about those projects and less about what we're doing. We're early stages on a bunch of this work, but hopefully we'll have more to say later in the year.
pauldix··on Actix project postmortem
Sad to see this. I'm using Actix in a new project and have been watching the project for a while. Guess I'll be refactoring to use Hyper. If I were farther along I'd probably try to take ownership, but it doesn't make sense given that it's only around 40 LOC in my project that depend on Actix at this point so switching out to another framework is more straightforward than taking ownership of a codebase I don't know.

I understand fafhrd91's (the maintainer) frustration, but it would have been much better to just abandon the project and throw it up for some other volunteer to come in and take it over. Then the project can live on and he can bask in any success it has in the years down the road. Instead, it's a complete mess where all the good work and good will has been undone in a single move.

I've abandoned multiple OSS projects over the years and let other maintainers come in and take them over. Now over 10 years later I can still say I created that project that still gets used because other people have seen value in continuing to contribute and maintain it.

pauldix··on We tried to hustle our way into YC after we got rejected
I don't think of YC as an accelerator, but maybe that's where things are headed since we (InfluxDB) went through in W13?

Originally it was enough to have a few people, an idea, and a prototype. I think that was the model for the most successful YC companies currently out there (Dropbox, Airbnb, Stripe, PagerDuty). Some didn't land on their actual idea on product until after the batch (Twitch).

There are certainly companies that come in with a baked product and the start of some real users/customers who then use their time in YC to juice their numbers and raise big rounds at crazy valuations right at the close of the batch. However, I think what YC offers that is unique (and a real strength) is that they back completely unknown founders very early in their process of building a product and a company and give them the connections and advice to help build something big.

The YC series A program strikes me as more of an accelerator.

pauldix··on A New Way of Voting That Makes Zealotry Expensive
Totally agree, that's where my mind went to immediately. Features and/or proposals on design in larger projects.
pauldix··on A New Way of Voting That Makes Zealotry Expensive
This actually looks like it could be applied very effectively to open source communities on a single project when setting priorities for future feature development. For example, you have 20 features in the backlog and you're going into a two month sprint in which you can only focus on 3. Might be an interesting way to open up voting to everyone. On larger projects this would be useful even if it's only the developers working on it that are voting.
pauldix··on On “Open” Distros, Open Source, and Building a Company
I don't think open core is being called into question. By that rationale, anything that is closed source that gets too popular is also called into question. Nothing is stopping AWS or anyone else from copying your API, features or anything else regardless of if your product is open, closed, or some combination in between.

I think if anything is called into question, it's the development of infrastructure software in a startup. Does it make economic sense anymore? As in, is it possible to profit directly off it at this point? Right now it obviously is, but it seems to be trending away from that.

pauldix··on On “Open” Distros, Open Source, and Building a Company
This is a really great response and I'm curious how AWS will respond. It's particularly damning because AWS was arguing from a position of moral authority on the move, which I think was disingenuous. They wanted to create a fork and act like they weren't. Shay just called them out on it.

I also wrote up my own thoughts, but it's obviously not as relevant as this. I'm a little more blunt about what AWS is doing. I even throw in a comparison to Microsoft abusing their monopoly power in the 90's by distributing Explorer for free to kill off Netscape: https://www.influxdata.com/blog/aws-intends-for-their-new-pr...

Longer term, the real question is how Elastic will respond to this commercially. Will they create more closed source? Go higher up the stack into specific applications? It's going to be an interesting and tumultuous couple of years in open source infrastructure software.

pauldix··on Open Distro for Elasticsearch
I view that as a significant negative impact to OSS. I have a strong preference for liberal licenses like MIT or Apache 2.0. Infectious licenses like copy-left put restrictions on it that make it less appealing as open source. I wrote a bit on it here: https://www.influxdata.com/blog/copyleft-and-community-licen...
pauldix··on Open Distro for Elasticsearch
Discussion from just over a year ago: https://news.ycombinator.com/item?id=16487440

They took the closed XPack code and put it into the open source Elastic repo, but under a commercial license. It muddied the waters for anyone looking to use or contribute to the open source.

pauldix··on Open Distro for Elasticsearch
Adding another thought to this, which is that I doubt the code behind this is a direct response to Elastic's recent license changes. This is a ton of work and I'm guessing AWS has been working on the code side of this since well before last summer. By all accounts their hosted Elastic offering is wildly successful so they've probably been hard at work to add features that Elastic formerly had as closed and commercial (before the license change).

I think their open sourcing of this work is a direct response to the recent license changes. If it weren't for those, they might have just kept this work closed and in use only for the hosted product (like they do with many other projects).

pauldix··on Open Distro for Elasticsearch
Oh I assumed that someone would do it. I just didn't expect it to be AWS. Elastic is popular enough and the core open source license permissive enough that I expected some sort of fork eventually. The license of the original project is an important factor in this. For example, I wouldn't expect a fork of MongoDB because of the limitations of the AGPL license make it less appealing. With that case I'd expect what AWS already did, which was to create an API compatible system. Also notable is that they're not open sourcing that.

With Elastic, this looks like it's half marketing on AWS' part. They want to start repairing their image in the OSS world so they say they're doing this for the community.

pauldix··on Open Distro for Elasticsearch
This is a very surprising move from AWS. In the past they haven't seemed to be willing to contribute to the OSS that they pick up and host. Even though they claim they'll contribute changes upstream, I doubt that Elastic would accept changes that are competitive with their commercial offerings. So you effectively get a fork.

I think this is a win for the Elastic community as a whole, but presents a real problem for Elastic the company. And that begs the question of what happens to the Elastic community if there's a real fork. And what happens if Elastic the company is very negatively impacted? Do we see a fracture of the community?

Looking even farther out, at this point does it make sense for any startup looking to create infrastructure software to open source it? If their project becomes successful, they'll get eaten up by the hosting providers. It makes smaller scale open source more commercially viable because you won't attract the attention of the providers that would come in a take your business.

Will be interesting to at least watch what early stage VC investment in open source companies looks like over the next 12-24 months. Will the open source pitch work for series A investments or will investors shy away?

pauldix··on InfluxDB 2.0 Alpha and the Road Ahead
Flux is almost a subset of Javascript. There are only two exceptions. Our pipe forward syntax, which I've seen proposed for Javascript, but who knows if that's going to happen. And the other is our named parameters with optional defaults. You can get close to this in JS by passing an object literal combined with destructuring.

However, there are other things we'll probably be adding to the language over time. Also, if you're only allowing a subset of JS, is it actually JS anymore? I'd imagine that some JS programmers could get frustrated when things they think should work (because they're valid JS) don't because they're not in the subset.

But say we went with that. Then we'd need to adopt an existing JS engine (most likely written in C++) and then modify it based on our needs. And then integrate that into our broader Go codebase. All of that seems a bit less than ideal.

Ultimately, I really do feel that if you're someone that knows Javascript, you can learn the elements of the Flux language in less than an hour. The bigger learning curve is the library of functions and the API, which would exist regardless of what language we chose.

I think the proof will come over time based on what kinds of things we enable in the language and the platform. In the near term I expect a large number of totally reasonable people to question the choice of a new language. After all, that's probably the rational response. But as we improve it, add to it, refine it, and improve the developer and user experience, I expect to win more converts. Developers pick up new tools because they enable them to get their jobs done faster. Ease of use, speed of development, and productivity are our guiding lights.

pauldix··on InfluxDB 2.0 Alpha and the Road Ahead
Our current cloud offering has much smaller initial pricing. Our Cloud 2.0 offering that we'll launch later this year will have usage based pricing as well as a free tier.
pauldix··on InfluxDB 2.0 Alpha and the Road Ahead
Lua was certainly an option and it's one I was even considering as an embedded scripting language back in the fall of 2014. I gave a talk in London where I asked for a show of hands about Lua vs. Javascript as the choice and the majority raised for JS.

One of the reasons we didn't go with that is because we didn't want a separate query vs. scripting experience like you have with SQL engines that embed programming languages. We wanted something that felt and looked seamless.

pauldix··on InfluxDB 2.0 Alpha and the Road Ahead
There isn't a specific feature of Flux that requires a new language, although part of the language is a query planner and optimizer that is tied to certain functions in it. Technically any Turing complete language is equivalent, but in practice that is seldom true. Performance is one vector, but so is expressiveness, aesthetics, ease of use, etc.

We chose not to use Lua because it's not a widely known language. Instead, we chose to create a language that looks and feels much like Javascript, which I think is the most widely used language today. However, we wanted to limit the scope of it and to add new syntax over time to make frequent tasks easier to express.

Any time you create a new API or library, you create new surface area for a developer to know. A language is no different than this. Some languages are also much easier to learn than others. For example, I found Go a language that I could pick up the basics in very quickly, while Rust is something that I'm still learning over months of effort. I like both languages, but their learning curves are very different.

I think Flux is quite easy to pick up for many programmers, although I'll have to see questions and have countless interactions with people trying to learn it to prove that out and make adaptations over time.

Finally, one goal with including the UI in InfluxDB 2.0 is that as we improve it, most users won't have to learn Flux at all. They'll be able to accomplish what hey want by just clicking around the interface. We're not there yet, but it's what we aspire to. And we want to have control over the language to build it in conjunction with UI tooling to automatically manipulate it.

pauldix··on InfluxDB 2.0 Alpha and the Road Ahead
Absolutely. We won't address that in the initial release of 2.0, but there will be ways to get it done. The eventual solution will probably revolve around using the tasks system to downsample into buckets of different retention. Then using a function in Flux at query time to look at the metadata of buckets in the query and the time range and selecting the precision based on that. We should be able to show examples of how to do this in Flux later this year.
pauldix··on InfluxDB 2.0 Alpha and the Road Ahead
InfluxDB 1.x was definitely not designed for that. With Flux in InfluxDB 2.0 you will be able to do things like store reference data in other places and join that with time series data in InfluxDB at query time. You can also sort, group and filter by any measurement, tag, field, or value. However, there are no user defined secondary indexes so the scope of this will be a bit more limited based on how things are stored. I'd have to know more about the specific kinds of queries to figure if it's something that would make sense within InfluxDB 2.
pauldix··on InfluxDB 2.0 Alpha and the Road Ahead
Those are indeed blockers for a full 2.0 release. There is backup and restore in the 1.7 release, but it's not nearly as powerful as it needs to be. We aim to correct that before any full GA release of 2.0.
pauldix··on InfluxDB 2.0 Alpha and the Road Ahead
This release of 2.0 is not a commercial offering. It's completely free and open source under an MIT license without restrictions of any kind. For example, you could use it as the basis of your own commercial offering or use code from it and that is all fair game.

For the query you mention, you can do that in 1.x using a combination of group by time, fill and subqueries. In 2.0 you can do this in Flux using similar operators. More complex possibilities for interpolating missing data is also in the works.

Solving more complex query processing like what you mention is a specific reason why we started developing Flux. We couldn't figure how to address some of the more advanced query feature requests in the old language so we started with something new.

In addition to giving Flux more functions, flow control and more power, we also will be adding more out of the box functions and syntactic sugar to make some of these more advanced queries possible without being overly verbose and complex. Our order of operations for developing the language:

1. Make it powerful 2. Make it easy 3. Make it fast

So it'll take us time to get to all of our goals, but I think we've laid down a pretty good foundation.

pauldix··on InfluxDB 2.0 Alpha and the Road Ahead
Thanks for the TICKscript feedback. That’s all stuff that we want to have addressed in Flux. This alpha release doesn’t have that yet but printf, a test runner built into the influx CLI, and test inputs and outputs are all on the near term roadmap
← PreviousPage 4 of 12Next →