TimescaleDB raises $40M
blog.timescale.com
blog.timescale.com
* They are still trying to figure out their monetization strategy. Initially, they betted on their on-premise Enterprise version, then abandoned it. Now they are pushing their cloud version.
* Even though most of their code licensed under Apache license, some code is under their proprietary license.
* I'm sure one can get some ideas about their development directions from their issue tracker and source code, but they don't have any public product roadmap.
* Even though the product itself is technically very stable, the version compatibility leaves a lot to be desired. There are removed features and broken APIs from version to version.
* Their commercial support terms for on-premise instances don't seem to be well defined, not publicly at least.
1. Monetization strategy
This funding round is actually a sign that our business model is working really well.
To quote Redpoint Ventures, who led this funding round:
"The [Timescale] team capitalized on their significant community momentum last year, with their cloud business being one of the fastest-growing database businesses we have seen in the past 20+ years." [0]
2. LicensingMost companies (including open-source companies) actually have both open-source and proprietary software, but the proprietary software is often hidden inside private repos. The difference with Timescale is that we have made the source code for our proprietary software available (on Github), even allowing users to modify it (eg "right to repair), and made all of our software free (ie no paid software features). [1]
3. Public product roadmap
We aim to be transparent re: product roadmap via Github, blog posts, etc, but I appreciate the feedback that we could be more transparent. Thanks!
4. Version compatibility / broken APIs
Could you say more? AFIAK the only time we "broke" (ie changed) some APIs is with TimescaleDB 2.0, and when we did so we explained why we did that (mostly to improve user experience based on feedback). We take this topic very seriously and even the decision to do so in 2.0 was not something we did lightly (and it was also made after a lot of discussion with users). More about this decision here in our docs: [2]
5. Commercial support for on-premise
We offer free support for on-premise instances via Slack (where you can often find our engineers, support team, CTO, and myself). [3] However, if you would like a higher level of support for on premise (e.g., commercial SLAs), please reach out to us directly (e.g., via the form on that same page). [3]
Hope this helps!
[0] https://medium.com/redpoint-ventures/building-a-next-generat...
[1] https://blog.timescale.com/blog/building-open-source-busines...
[2] https://docs.timescale.com/timescaledb/latest/overview/relea...
I believe that our "major version" upgrade from 1.x to 2.0 was the first time we changed/broke any APIs, but that involved a long beta/RC process, much documented about the changes [1], and upgrades that also meant to seamlessly migrate.
For example, upgrading from 1.x to 2.0 was still just running `ALTER EXTENSION timescaledb UPGRADE`. The main difference was if you were, for example, using some of our informational views in your applications, those had a change a bit. Or if you were querying internal catalogs in your app (although that is never recommended =)
Even after 2.0 was launched, we did backport bug fixes to some follow-on 1.x releases, and continued to support users running 1.x on our cloud platform.
[1] https://docs.timescale.com/timescaledb/latest/overview/relea...
As for 4 - yes, I meant 2.0 and also removal of adaptive chunking in the earlier versions.
There is always a trade-off between getting features to users quickly to experiment and incrementally improve, versus doing it always very conservatively.
When we launched adaptive chunking (introduced in 0.11, deprecated in 1.2), we explicitly marked it as beta and default off, to hopefully reflect that. [1]
The approach we are now taking with Timescale Analytics [2] is to have an explicit distinction between experimental features (which will be part of a distinct "experimental" schema in the database, and must be expressly turned on with appropriate warnings) and stable features. Hopefully this can help find a good balance between stability and velocity, but feedback welcome!
[1] https://github.com/timescale/timescaledb/releases/tag/0.11.0
[2] https://github.com/timescale/timescale-analytics/tree/main/e...
I was going to write a lot here but I'll keep it short.
What I discovered is that everything I want to do can best be done in PostgreSQL. It's one thing to do data analysis in a Python notebook and another in an environment that works dynamically on a server. My first guess was to do the heavy lifting in Python with Numpy, Pandas, and machine learning and have the node server -- instead of Django and if I'm learning a new web framework it will be Phoenix -- execute the Python scripts through stdin / stdout. Since started I learned that I don't need machine learning and that I can do the calculations inside PostgreSQL sometimes orders of magnitude faster than in Python.
I'm using TimescaleDB which provides the postgresSQL time_bucket() function and with chunks should scale very well. First I tried to integrate it with Prisma in node, however, that proved to be far too difficult and convoluted. I reverted back to using TypeORM in node and it was extremely easy to run all the boilerplate code to initialize the TimescaleDB plugin inside of migrations which would probably be just as easy in another framework like Phoenix with Ecto. Sometimes I use SQL queries in a string literal and other times I use the query builder for more dynamic interaction with the database and to leverage some of TypeORM's other features beyond only being a connection manager.
What I discovered which interestingly someone yesterday shared a popular link to a blog post on the subject[0], for most of time series analysis, Pandas isn't required and perhaps not the fastest solution. Grokking window function was a little difficult until I found this lecture on YouTube, Postgres Window Magic[1]. Leveraging and understand window function in SQL is probably the most important skill to have.
I don't need Python and Pandas for time series analysis. I can using TimescaleDB and some increased knowledge of using PostgreSQL do time series analysis using all the same infrastructure I've been using for the past several years.
remeber that you are in a unique position where you know the ML application and specialized pSQL to implement it.
The market is paying big bucks for people that have either of those skills. If you are making less than 300k/y (at the very least), move out now ;)
Also, what does "knowing ML" really even mean?
So I personally like your company, but I find these sorts of marketing speak responses obnoxious. Your communications strategy here is causing brand harm, not benefit.
1) You have an open source core, with proprietary components.
2) Open source adopters get a crippled product.
3) You have a custom license for the proprietary components, which is designed to allow people to make some use of those, but is poorly-written ambiguous (preventing many types of commercial use), non-open-source compatible (preventing integration into open source projects), and requires a lawyer to review (preventing integration by smaller projects).
This feels like your Achilles' Heel.
Troll Tech tried to go down this line for years, with their QPL license. And they even had sane messaging, where whenever I read your messaging, it feels weaselly, and it changes week-to-week. Still, they didn't really take off until they went with a licensing system customers could trust and understand.
The standard dual-license model would be AGPL and commercial (or GPL+commercial).
* Most open source developers won't mind (or even notice) licenses, so long as their open source and have the nice OSI and FSF logos.
* Most commercial companies won't mind paying $$$.
Commercial customers treat you like Microsoft. Open source developers treat you like community members. Hybrid customers are okay too; if I'm working on a piece of BSD code, I can use the AGPL license on your code, while commercial users of my code can buy a commercial license from you.
And if you insist on the crazy custom license, figure out the messaging. This was better than what I read before. "Proprietary with a public repo" makes more sense than previous messaging which sounded like open source but wasn't. At that point, at least the license overdelivers rather than underdelivers. I still trust that at some point, as an adopter who can't or won't use those components, they'll become increasingly mandatory if you ever fall on hard times. The problem is still that it makes it sound like you have open source and proprietary products. You don't. You have a product with open source and proprietary components, a confused freemium model, and not something I'd ever use without consulting a good lawyer, who in turn would tell me to stay away.
There are many other good models. You could go in the other direction and close up a bit too.
I see posts like yours all the time - attacking companies who believe in open source but know that the existing situation can and will lead to abuse by huge players in the market. It sounds like they have a very happy medium.
I find posts like this very frustrating because you almost never see the same kind of standards for proprietary software. People get less pushback for closed source than they do for "90% open, with restrictions on the last 10%".
As for AGPL dual license, have you tried selling AGPL software before? Even dual licensed, you've just added a massive roadblock - and it's no better than a custom license anyways, since it implies one.
Not to mention you've completely broken any path from "I'm a free user" to "I'm a paying user" - someone has to decide upfront to pay. Given that this is the most significant "win" for open source software (try before you buy, easy to inspect and get started with) it's kind of ridiculous how often I see it suggested.
Let's say I work for a company willing to pay for software. I have a hackweek where I want to try out TimescaleDB, in a world where it's AGPL or, upon payment, dual licensed. That's now dead in the water - I can't use it, AGPL is banned at every company I've worked at and we're not going to pay for me to try it out for a hackweek project.
I disagree with both the tone and content of your whole comment, but this bit in particular stood out. I'm a very happy TimescaleDB user, and I really like the company and community around it too. Also, support on Slack is amazing and they are always responsive on GitHub. And of course, I have read the license - to me, it's easy to read, and the intent is very clear. I see no ambiguity.
"that does not expose or give access to, directly or indirectly (e.g., via a wrapper), the Timescale Data Definition Interfaces or the Timescale Data Manipulation Interfaces to any person or entity other than You or Your employees and Contractors working on Your behalf"
If I have a system, and it has an AJAX API, at what point am I violating this? Virtually every system I build provides customers with access to data via some API which uses "SELECT, INSERT, UPDATE, and DELETE" "via a wrapper." I have no intent of competing with TimescaleDB for hosting, but it's hard to argue that data interface don't provide some form of access to Timescale Data Manipulation Interfaces "via a wrapper."
"the customer is prohibited, either contractually or technically, from defining, redefining, or modifying the database schema or other structural aspects of database objects, such as through use of the Timescale Data Definition Interfaces, in a Timescale Database utilized by such Value Added Products or Services"
I won't even begin to get into this one.
If you build on this kind of legal language, you're taking on a legal liability the size of a moon crater.
The second quote is about providing access to customers (Section 2.1.b). Note it certainly allows SELECT, INSERT, UPDATE, DELETEs (those are DML operations), it prohibits you from allowing customers to do things like `CREATE TABLE` (those are DDL operations).
https://www.timescale.com/legal/licenses#section-2-1-grant
This is our approach to define what it means to provide "TimescaleDB-as-a-Service" from a more technical perspective, that hopefully a developer can grok, as opposed to just stating something about "you can't compete", which is open to broader interpretation.
I think you've done a very poor job with doing something which a court will grog the same way.
I think that's where the astronomical potential liability comes in with respect to using your product.
A lot of 2.1.a versus 2.1.b will hinge on details of how a court will read ambiguous language like "not primarily database storage or operations products". I assume you wanted to say "not primarily database storage or database operations products." However, it could just as easily read "not primarily operations or database storage products." At that point, "operations" has broadly different meanings (e.g. business operations?).
And aside from that, if I'm making a medical database product, is that primarily "database storage?" Probably.
The problem with ambiguous legal language is that:
1) You, or a vulture successor, can plausibly sue anyone who does just about anything.
2) If we assume you vulture successor has a 20% chance of winning $20 million, the outcomes is a $4 million settlement.
Which is why good lawyers avoid it. The whole document is just bad legal language.
But even if it was GOOD legal language, it wouldn't matter. The difference between a custom-form license and a standard OSI license is that competent customers need to spend a few grand on legal fees before they use yours.
I understand what you're trying to do, but every other organization that went that way eventually went with a standard license. You'd be better off doing likewise. Or if you really can't, you're better off working to make a community-recognized standard form license which is used by enough products that it has a standard, common, recognized legal understanding.
I know the risks of the AGPL, and where it will or won't hurt me. I don't know the risks of your license, except that they're obviously huge.
For others, I can share at least that it was drafted by some of the most experienced copyright & IP counsel there is, including with significant open-source licensing experience.
But anyway, we're providing it as free software, so if you don't feel comfortable with it, you are certainly free to use our Apache-2 version. Cheers!
"Never Take Legal Advice from Opposing Council."
Terms-of-service and Facebook's employment agreement were generally drafted by experienced counsel. That means they do a good job of protecting the person on the opposite side of the table, not of protecting me.
And the term isn't "free software." It's "freemium software."
It's exactly this sort of comment which makes me distrust TimescaleDB.
I don't know about civil law jurisdictions, but in the US, this license is a liability bomb.
> * Even though most of their code licensed under Apache license, some code is under their proprietary license.
I don't really think this is a perfectly fair characterization. Their proprietary license is essentially "don't host a cloud database and charge for it" to stop Amazon from building TimescaleDB right into RDS, or similar.
I think it's a totally fair license without too much to worry about if they go out of business.
> * Even though the product itself is technically very stable, the version compatibility leaves a lot to be desired. There are removed features and broken APIs from version to version.
This would be my biggest worry. Upgrading Postgres is already stressful enough, having to deal with broken APIs from version to version would leave me pretty upset, though I've not heard of anyone complain about this before, so I'm not sure how much of a problem this is in practice.
> I don't really think this is a perfectly fair characterization. Their proprietary license is essentially "don't host a cloud database and charge for it" to stop Amazon from building TimescaleDB right into RDS, or similar.
"yeah, go ahead, infringe that copyright and host a internal/for-direct-clients only database, because i guess that would be OK from their license, even though it is not explicitly allowed"
pardon the sarcasm, but i literally heard that from our lawyers today regarding another project's license, as something (quoting again) "no lawyer would ever say to their clients".
At least that's what I think, I'd want to hear kemitchell's review of the most recent iteration of their license, I think it incorporates much of what he's discussed as the correct legal direction for open-except-for-clouds licenses which strikes the right balance between user protections and safe guards against cloud providers.
Just in case you're not a native English speaker: The verb "to bet" is irregular and the past tense is simply "they bet" rather than "betted" (which would be far more logical).
Modifying business models to optimize for success is a good thing, not a negative.
The performance is a great feature but its also just an intuitive, familiar (pretty much just SQL) tool that makes life easier.
* InfluxDB has way better documentation on functions. For example, look up moving average by time (not points) on TimescaleDB vs InfluxDB. We use these more complex queries and have no problem on Influx. Going further, the number of functions built in is impressive with the same ability to define new functions.
* InfluxDB containers are totally self contained which is great for simple architectures. As a process, InfluxDB is a single executable thanks to Go.
* This is extremely subjective, but I find Flux easier to comprehend as a separate query vs. the use of SQL to do higher complexity functions; however, I am sure this is due to my lack of experience and know how to write said queries in SQL.
The benchmarks are interesting, showing TimescaleDB to be the clear winner in most scenarios.
For me that's nice, but it's a bigger deal to me personally that I already have Postgres and SQL experience that translates directly to TimescaleDB, I don't have to learn a new tool and query language. Development is complex enough and I have to learn too many things as it is. The older I get the less enthusiastic I am about adding something new to the stack.
Tangentially related to that: their mongo benchmark numbers always looked odd to me. Given that I've used mongo for 10+ years for high throughput time series data without major issues, I decided to do my own benchmarks. In my testing, mongo outperformed timescale significantly both in write throughput and query performance.
This is likely in part due to the fact that I'm using well-understood internal data from real production systems, and as such my ability to be able to build performant indexes / query strategies in the database that I know best introduces a performance bias.
I always take benchmarks with a grain of salt, for this reason. And I try to lean into the tech I understand best.
Always strive to do the best and fairest benchmarks we can, and for that reason, all our benchmarks are fully open-source for both repeatability and improvements/contributions:
https://github.com/timescale/tsbs/blob/master/docs/mongo.md
We also really did spend a lot of time investigating approaches with MongoDB, so you'll see our benchmarks actually evaluate two _different_ ways to use time-series data with MongoDB (culled & optimized from suggestions in MongoDB forums). But always welcome to feedback:
https://blog.timescale.com/blog/how-to-store-time-series-dat...
Thanks!
I've reviewed all these resources multiple times in the past, which is what prompted me to do my own benchmarks (in which mongo outperforms both multinode and single node configurations).
Some issues I noticed:
- youre using gopkg.in/mgo.v2 which is a mongo driver that hasn't had a release in 6 years. Not sure of the general performance impact here, but my tests use mongo 4.2 with a modern node.js driver. So thats one difference.
- your indexing strategy for mongo is easily changed to be able to get much better performance than the naive compound approach you took in the code (measurement > tags.hostname > timestamp).
- you didnt test the horizontal scaling path at all, this is where mongo arguably shines
I'm glad you all open source this stuff because it helps engineering leaders make better decisions, so thank you for that. But your data does not align with my own: either our production metrics or through structured load testing.
That's probably not something most companies would do for benchmarking, but we take ours seriously :-)
Currently I'm stuck on figuring out how to get data into TimescaleDB. My company makes heavy use of Telegraf, which is a natural fit for InfluxDB, but not so much for TimescaleDB. The original pull request for the Telegraf plugin for Postgres/TimescaleDB was closed because the author was non-responsive: https://github.com/influxdata/telegraf/pull/3428
I can even write data to it using simple TCP or UDP tools like netcat or curl. And for some cases I have simple scripts that do exactly that. TimescaleDB, on the other hand, requires some sort of Postgres client.
What do you, or other people, use for writing data into TimescaleDB?
I still wonder what other people are using to feed information into TimescaleDB. I'm wondering if I should switch to a different approach, such as using Telegraf but routing the data to something else that will push data into TimescaleDB.
Don't want this to come across as overly defensive, but was under PR review for 3+ years by Influx with little progress (first submitted in November 2017) and during that period I think we did something like 2 significant rewrites. Became a bit of a moving target against telegraf that became harder to prioritize.
However; our problem space is not high cardinality data; it more closely aligns to the first performance comparison with 10 devices and 10 metrics. The ease of getting high performance with pre implemented functions is great for us. Reliability is obviously a concern, and I can agree that if data is sacred, then choosing something built on Postgres is going to be a better thought.
Again, this is just our problem space; small scale deployments on many machines with no preexisting RDMS, low cardinality data, etc. I think it’d be a different story if we were huge, but for us, InfluxDB provides some seriously handy feature and is worth consideration if your problem is similar.
One project we launched earlier this year "Timescale Analytics" actually seeks to address exactly this, e.g., bring more useful features and easier programmability to SQL [1] and you can see (or add) to the discussion on github [2].
Also informed by some of the super helpful functions we've seen in PromQL. And by the way, if you are interested in PromQL, we have 100% compatibility with PromQL through Promscale [3], which provides an observability platform for Prometheus data (built on TimescaleDB).
[1] https://blog.timescale.com/blog/time-series-analytics-for-po... [2] https://github.com/timescale/timescale-analytics/discussions [3] https://www.timescale.com/promscale
Most developers are already familiar with Postgres or at least SQL.
The tooling around Postgres is basically universal.
There's huge value in an option that is literally just "install this Postgres extension and everything works and gets out of your way".
We use TimescaleDB for a handful of products. In several cases we literally just updated a DSN to point a product at TimescaleDB instead of an existing database and the project Just Worked(TM) except hundreds of times faster.
And some of those that we developed on TimescaleDB natively, it was more or less the same thing... give a team TimescaleDB and they're basically productive immediately. There's no learning and integrating new libraries and query languages, no time from ops finding new and exciting problems to solve in hosting and scaling the DB, etc.
We get all this with all the functionality and strong guarantees that Postgres provides.
Example: your temperature sensor is faulty and produces values like -100. You can't delete this data by using "delete from measurement where temperature < -50". You have to get all timestamps, then delete those timestamps one by one.
Btw, ClickHouse is under Apache 2 license, which makes it much easier to use in big companies.
[0] https://blog.cloudflare.com/http-analytics-for-6m-requests-p...
Good old HN with its healthy skepticism :)
IMO HN's classic "skepticism" is usually just engineering nerd insecurity projected outwards, with enough techno-jargon to maintain plausible deniability. Folks feel threatened by a great idea so it's safer to find some way to tear it down. Not to dismiss all feedback as projected insecurity of course.
1. Universally Negative - Either it's cryptocurrency-related, or it depends on source of negativity:
A. "I read the site and I don't know what this is" - Genuinely bad explanation of an idea that doesn't seem particularly technically interesting or challenging.
B. Criticism of superficial aspects (e.g. website, related topics) - Genuinely bad explanation of an idea that DOES seem particularly technically interesting or challenging. _(Commenters don't get the message, but are worried they'll appear ignorant if they say it.)_
C. "Nobody needs this" "Why is this a thing" - Either bad or HN is nowhere near the target audience.
D. "This is not the right way to do it" "You can just do X" - Either bad or revolutionary (and new enough that the idea hasn't clicked with anyone.)
2. Polarization - A. If positive people are REALLY positive about it - potentially a disruptive technology, potentially ahead of its time.
B. If negative people say it's actually much harder to solve - the idea is great in principle but the only reason it hasn't already been solved is it's not possible or very difficult in practice.
3. Universal Adulation - It will transparently never make any money, it is some kind of attempt at decentralization that will never get adoption beyond hardcore nerds.Timescale when it first launched was little more than an automatic-sharding extension for Postgres with some convenience functions for handling time data. It was competing with Postgres itself which added native partitions, other sharding extensions like Citus, and an entire class of column-oriented relational databases that have become much more capable.
Timescale today is very different and has added a lot of the missing functionality to make it a very attractive database option, especially the columnstore/compression feature mentioned in that first HN comment.
But I just don't see anything that makes creating an entire database design for one specific index type worthwhile...
I index many tables on my site by num_upvotes so I can find the top ranked items to show. Does this mean that I need an UpvoteDB? I don't think so.
A previous time I argued this point, it was mentioned that you rarely need to update or delete old rows. This allows you to tailor the storage solution better. However, this basically means a compressed column store, which again, doesn't really have much to do with time.
Logically they're the same thing, but engineering is about details, details in this case that could easily be a 2x to 20x budget difference given an appropriate project
A column store can take 100 years worth of samples occurring every 10ms that yield a constant result and using technology we actually have, represent those ~87 million data points on disk and in CPU using somewhere under 10 bytes.
Also, some of these time series databases have very specific use cases and you have to also think about the client tools associated with the database. Many of these databases sit in power plants, factories, etc. and they stream data to tools that are built to visualize or analyze the last few minutes of data and then trigger alerts based on patterns. Also, these database are very "device" aware and integrates with other systems that represent their data in a timeseries fashion already (like a sensor). A lot of customers who needed this type of database care only about this index because their concern is record keeping and monitoring. Not necessarily number crunching (this is changing though).
There are drawbacks to storing your data this way. If your primary index is time, it can be hard to merge that with some based on a coordinate system. So doing certain types of analysis is really difficult unless you replicate your data into some other database with a different index.
> However, this basically means a compressed column store, which again, doesn't really have much to do with time.
It does though: which data do you compress? The old data. Why not let the database figure that out for you, so you don't specifically have to tell it.
Other features include:
- Continuous Aggregates: a materialized view aggregating data over time is doable, but why not let the database materialize it for you, and automatically fall back to an un-materialized query for the newest data?
- Retention: deleting (or downsampling) old data is easy to do on your own, but why not let the database do it for you according to a policy?
HN is addicted to bikeshedding. It's among the top 3 comments on almost every "Show HN" or new product launch.
Founders should have a thick skin when it comes to criticism on HN, because we don't know either.
TimescaleDB's biggest feat here is of course pulling the engineering magic rabbit out of the hat by chipping away at it for 4+ years, and effectively answering the skepticism by delivering on their promise.
Note though, Amazon Redshift is built on Postgres, and (allegedly) so is Amazon Timestream.
[0] https://blog.timescale.com/blog/building-columnar-compressio...
[1] https://blog.timescale.com/blog/time-series-compression-algo...
To me building upon PostgreSQL was however, a good idea. Long term all databases gain relational features, through a rather painful process of realization that RDBMS actually did some things right. They'll skip that pain and focus on new features.
Microsoft does something similar by offering a Graph DB on top of MS SQL.
The point is that a lot of "bad ideas" have something major going for them.
So they can bring their domain expertise in process manufacturing and elsewhere, and then build on a modern, powerful platform. We're excited to see this!
https://ecostruxure-building-help.se.com/bms/Topics/show.cas...
But feel free to hit me up at mike@timescale or mike on slack.timescale.com.
Always nice when we get to just swap it out for something like Canary or even Ignition although folks are always trying to trash the Ignition Historian when it works well for most use cases people need to solve.
Is TimescaleDB suitable to store logs? If yes, how to architect the tables?
This is bad. Just be precise and don’t try to conflate open source with proprietary stuff. It’s ok to not be OSS, lots of good software is non-OSS.
It’s only bad to say you’re one thing and be another.
But I can attest that the single-server version is rock solid, just like PostgreSQL that it's based on. And it's free and source visible. The pace of innovation has also been really high, it just keeps getting better with every release.
A good reference point for this to check out the `multinode` label in their github issues. https://github.com/timescale/timescaledb/labels/multinode
One of the big items that stood out to me is the inability to be able to migrate data from an existing table when creating a distributed hypertable. There were also some significant query performance reports as well.
These all may improve with time of course, so watching the dev cycles will give you a good sense of that I think.
msaharia[a]iitd.ac.in
The main difficulty was getting access to the data (and ensuring it was valid, as with all ML projects), lucking we managed to get that from several sources (councils, water companies etc). The flow data was in TimescaleDB, pandas dataframes so we could use varying levels of frequency of data and we used HDF5 iirc as well (detail is hazy, it was a few years back now). We did demo it to the Met Office here in the UK too. They were interested but already had their own thing cooking up so the project never really got out of prototype, but it was making accurate predictions. I think there were some other areas that might turn out to be flaky over time using this method (such as rapid changes to catchment areas) but that could maybe be factored in someway with more thought on the model and verification on a larger set/timeframes. Feel free to hit me up if you want any more detail on the tech side, but don't ask me about stats, maths fu I ain't :)
joel [at] smashthesystems.com
Best of luck from a happy customer.
"With real-time aggregation, when you query a continuous aggregate view, rather than just getting the pre-computed aggregate from the materialized table, the query will transparently combine this pre-computed aggregate with raw data from the hypertable that’s yet to be materialized. And, by combining raw and materialized data in this way, you get accurate and up-to-date results, while still enjoying the speedups that come from pre-computing a large portion of the result."
https://blog.timescale.com/blog/achieving-the-best-of-both-w...
(You would then just need to query / point to the Continuous Aggregate to get the new and aggregated data in the same query)
The contrasting queries can be seen here: http://www.timestored.com/b/kdb-qsql-query-vs-sql/
I do agree they could use competition as the overall offering is weak but the core database is very strong.
NB: I'm one of the co-founder of questdb [1] https://www.questdb.io