Polars: Company Formation Announcement
pola.rs
pola.rs
For some history, there has been a bit of contention between the official arrow-rs implementation and the arrow2 implementation created by the polars team which includes some extra features that they find important. I think the current status is that everyone agrees that having two crates that implement the same standard is not ideal, and they are working to port any necessary features to the arrow-rs crate and plan on eventually switching to it and deprecating arrow2, but it will take some time to get there.
If companies making money have to / want to pay for support and some extra metrics, it seems very reasonable to me. You can always fork the open source project and add the enterprise features yourself of course!
I don't think your conclusion is axiomatic. There are plenty of flourishing open source projects, bug and small, that got handed over from the original author to new maintainers - sometimes multiple times. All without VC funding.
The popular open source projects make great targets for low-multiple acqui-hires, derisking the investment. They also have huge established branding and generally obvious opportunities for displacing existing players. In a non-zero interest rate environment, those factors make established open source projects a more appealing bet.
Since they mentioned the company, undercutting and eating the market share of DataBricks is enough to appeal to some investors.
Compare:
Funding databases is some of the most appealing for sw infra VC b/c as business fundamentals like monetization (pay for hosting, data, etc), growth, retention, are some of the most successful & low-risk
Funding the sw compute tier is a peg down but appealing for similar reasons. Basically same-but-weaker than DBs on the above dimensions, but still worth it as customers struggle w/ compute at scale (technical + business), so still works. Think early databricks vs snowflake, and how databricks grew to owning more than compute to data lakehouse, dashboards, etc: started as pure compute and now closer to snowflake.
Python is popular for a lot of these compute tier stacks. Orchestrators, AI, ETL, etc. The technical, social, & economic reasons are all interesting & relevant for why.
That's similar to anyscale (ray), coiled & saturncloud (dask), and early databricks (spark). Managing infra for that kind of thing is annoying. These companies don't OSS their cloud stack.
Wishing them luck! A lot of arrow-core compute tier & db co's emerging, so cool to see the many years paying off.
That's what I had missed - thanks.
I'm very interested in improvement suggestions or comments you may have (just reply to this comment).
For classification and visualization I prefer tags (not necessarily mutually exclusive) rather than categories (too rigid and therefore arbitrary and therefore hard to keep consistent), but then the database management becomes more complicated (need a transactions-tags table or equivalent).
I want to automate the classification as much as possible, so I set up merchant-tags, which are automatically applied to all matching transactions. Again, this means managing a merchants table. It turned out that both the number of repeat merchants is high (so lots of patterns to maintain; these are stored in the DB as part of the merchant record, rather than in the code) and the number of novel merchants is high (so still spend a lot of time on the manual classification).
Part of the motivation for the above complexity was a vision of some kind of dynamic, auto-generated, multi-level breakdown of cost categories. For example, I might want to see a "date night" tag in both food>dining>date night, and entertainment>date night.
Anyway, all the above complexity was fragile, and too much work on top of the manual classification. A few years ago something in the DB broke and I just didn't get around to fixing it. I probably need to restart with something simpler.
Also, I'm considering trying this: https://lunchmoney.app/
+1 on restarting with something similar. If you can get away without a db it might be easier
> We are aiming to deliver a Rust-based compute platform that will efficiently run Polars at any scale.
> We believe that the Polars API can be used for both local and cloud/distributed environments. Our API is designed to work well on multiple cores, this design also makes it well poised for a distributed environment. We also believe that a Rust based columnar OLAP engine (Polars), is perfectly suited for efficient distributed computing.
I suspect they will sell "cloud-scale distributed computation" systems. Perhaps something like snowflake?
This has me thinking:
Every time I think to myself no way there is a business for X I need to remember there quite frankly could be a business for most things, as long as the problem domain being solved for saves labor vs cost, enables new use cases, or business expansion etc.
I've sat too long on far too many things that I'm like oh no one would pay for this when in fact, I bet at least one of my projects could be revenue generating.
However we have a huge amount (and growing) of Pandas code - is there an easy way to convert that in small pieces to Polars code?
I feel like its the data exploration that locks me into Pandas, and I kinda want out.
We are hiring ... We are looking for +- 4 CET.
This rules out people in North America, right?Will you also offer paid support?
You can also email us info@polars.tech to get more info now.