TimescaleDB vs. Amazon Timestream
blog.timescale.com
blog.timescale.com
I would eventually like to see comparison for a more analysis heavy workload vs Druid though.
Simple metrics workloads like what Timestream and Influx do aren't really of interest when comparing to Timescale ability to do JOINs and use other SQL and PostgreSQL specific functionality.
I suspect there are many enterprises that are already so locked into AWS that it would take them a decade to move to a different provider.
One example which was influential in nature - Dyanmodb. The Dyanmodb paper inspired a number of NoSQL databases, including Apache Cassandra. If they have went with the popular NoSQL database at that time, we might have missed some good contribution.
https://github.com/timescale/timescaledb-kubernetes
Or just use Timescale's managed cloud on AWS:
I am not complaining AWS is doing vendor lock in makes complete sense, I am just commenting that is their main direction for their offering. And their vendor lock is aimed at the enterprises where they can really run wild on it and get lots of money in return. Smaller orginisations will generally never lock themselves in so much because they won't build so many things that are so hard coupled since they just don't have the manpower.
You're right, they're definitely doing things for the sole purpose of locking you in.
Serving is ok on Sagemaker, i will admit. The new ones: feature store, pipelines, and data wrangler - we will see, all rushed (some more than others).
1. Queries will need to touch a lot of data - will cost a lot, and there is no ability to optimize the queries in any way (no EXPLAIN, no indexes, no downsampling)
2. No integration with data exploration and visualization tools
3. No ability to have non-timeseries data, or correlate any data in two different tables (no JOINs)
(Disclaimer: I work at TimescaleDB)
Authentication to the database is typically governed to different security mechanisms, including certificate-based auth, password via SSL, etc.
I would like to see comparisons between Timescale and Druid.
I already know Timestream isn't up to the task lol.
[1] https://valyala.medium.com/measuring-vertical-scalability-fo...
You basically reinvent all the stuff - from databases to all avaliable tooling
normal companies dont do that, so there' an opportunity to get deep into some specific branch of applied informatics.
Also, they hear about the problems some customers have with an existing product and try to (re)invent a better solution. (MySQL->Aurora)
So this give them a bit of lock-in. Come for the EC2, stay for the minimally viable queues, machine learning, containers, etc.
The business model is solely focused on competing with and taking out other businesses, not actually providing any value. AWS is like a pack of sharks.
"The viability of our company, Timescale, is 100% dependent on the quality of TimescaleDB. If we build a sub-par product, we cease to exist. Amazon Timestream is just another of the 200+ services that Amazon is developing. Regardless of the quality of Amazon Timestream, that team will still be supported by the rest of Amazon’s business – and if the product gets shut down, that team will find homes elsewhere within the larger company."
Source: https://pages.awscloud.com/awsmp-h2-fin-time-series-database...
As for “making something for the sake of having it”, I believe the key reason for building their own applications is that they can build it on the same multi-az incremental ledger/block store that underlies most services (first built for aurora), or a variant thereof. They have some pretty stellar underlying tech built to be reliable at cloud scale, something you won’t get from running a service on your own EC2 fleet. The new stuff built on top of it just isn’t all that mature though. That’s just my take though. If interested, the deep dive talk from reinvent last year on aurora is pretty cool.
Disclaimer: I work at Amazon, nothing to do with AWS or timestream. In fact, I’ve been excited for time stream for a while but found its initial performance lackluster for what I needed.
But all of the Deep Dive talks are good for technical details: https://www.youtube.com/results?search_query=aws+reinvent+20...
It's not just this. In many cases, the services launched are essentially MVPs that were only created to serve a handful of customers' needs. If a huge whale of a customer says they want hosted Jupyter notebooks on AWS (just a theoretical example I picked because I saw someone criticize Sagemaker, but I don't have any specific knowledge of Sagemaker), then AWS will create that for them, but, AWS doesn't actually do "one off" services for a specific customer, so they will also launch those hosted notebooks as a public service.
The result is of course that this service was built with one customer's use case in mind, so it might not serve other customers' use cases very well, which leads to some negative perception. But ideally AWS then goes and incorporates feedback into the service, turning it into a "tier 1" service rather than an MVP. It just takes time.
If I were a SAAS competing with an AWS service, I would want to know how long before a given AWS service might get good enough to be an alternative for my customers. If you run with the idea that AWS offerings get better over time and eventually cross some threshold of "not good" to "good", maybe a crawl of the AWS announcements[0] in combination with a metric could provide a predictor of an AWS service's time to "good"...
Kinesis was a solution to their in-house woes with metering and billing which absolutely drowned in so much data from internal services with myriad metering and billing items and rules that they had to do something about it themselves: https://gigaom.com/2014/03/20/why-amazon-built-its-data-stre...
BTW, Kinesis Data Firehose is a pretty good product.
Amazon designed a system that is perfectly horizontally scalable. The only issue is that they now need 33,000 machines to achieve the throughput of TimescaleDB running on a single node.
I would call the results this new school of software development produces horrifying, but I already get lynched every time I tell a startup that their horrifyingly complex horizontally 'scalable' AWS setup that costs them $10k a month could be replaced by a handful of ducktape running on a rPI.
The best part is how the scaling never really works and just adding a load balancer and a second rPI would've worked better.
"We didn't think we'd get that many users. I mean 10k! Whew! It just wasn't designed for this!"
"Oh. That sucks. My rPI is serving 2 million monthly visitors right now and I'm still waiting for CPU usage to dip into the double digits."
I'm not saying people should design their stuff to handle millions of users on underpowered hardware, but I am suggesting they not lose their sense of perspective while building their whatever.
Because if you do, you may end up thinking you did well when your DB maxes out at 500 inserts/second - Another DB running on a Nintendo DS might just outperform it.
Single-node VictoriaMetrics provides almost perfect vertical scalability though. Its' performance scales almost linearly from rPI to a monster with hundreds of CPU cores and terabytes of RAM. See https://valyala.medium.com/measuring-vertical-scalability-fo... and https://valyala.medium.com/billy-how-victoriametrics-deals-w... .
Might as well add on a bit here -- if you're into this sort of thing you might enjoy Timescale thrashing other purpose-built databases (which have since also improved so YMMV):
- Timescale vs Influx[0]
- Timescale vs Mongo[1]
- Timescale vs Cassandra[2]
These blog posts are all from the timescale side, but the numbers are undeniable and their methodology is open source -- not sure there's more you can ask. Cassandra probably scales better at FB/Twitter scale, but with the recent work done to improve Timescale's scaling features[3] and the permissive license (unless you're an Amazon) I'm not sure even that is a real benefit.
I'd love to see a response from Amazon -- if HN is so lucky to have a dev who worked on AWS Timestream that would be awesome.
[0]: https://blog.timescale.com/blog/timescaledb-vs-influxdb-for-...
[1]: https://blog.timescale.com/blog/how-to-store-time-series-dat...
[2]: https://blog.timescale.com/blog/time-series-data-cassandra-v...
[3]: https://blog.timescale.com/blog/timescaledb-2-0-a-multi-node...
As an aside, we tried to solicit help/feedback on both Twitter and Redit groups. One AWS employee reached out and offered to pass it to the team (I was upfront about what we were doing) - but we never heard back.
However none of the databases tested above is anywhere close to the efficiency of a column store database that can do vectorized execution over batches of rows. ClickHouse is a good example of one such database. For queries that have to shift through large amounts of data, either filtering or aggregating it, the performance difference is easily >10x. I was seeing aggregation performance above 2B rows/s and that was I/O throughput bound.
> However none of the databases tested above is anywhere close to the efficiency of a column store database that can do vectorized execution over batches of rows. ClickHouse is a good example of one such database. For queries that have to shift through large amounts of data, either filtering or aggregating it, the performance difference is easily >10x. I was seeing aggregation performance above 2B rows/s and that was I/O throughput bound.
Agreed -- OLTP (in the case of Timescale) and purpose built timeseries-focused (but not necessarily analytics focused) DBs hold nothing to a proper OLAP database.
Did you write about this anywhere? would love to read it. I've never had a real need for the kind of stuff that Clickhouse does, but it looks to be the best in class for F/OSS OLAP DBs. Have you ever tried Druid?
I'm still not 100% sure what SciDB is for but it looks like none of the databases we're mentioning are directly comparable:
- SciDB for scientific computing (??)
- VoltDB for in memory + SQL
- Druid for OLAP queries
- TimescaleDB for row-based timeseries storage
Seems like all apples to oranges to me. I personally like TimescaleDB because it runs on Postgres and I can usually find a way to do all those other things in postgres relatively efficiently (memory requires some squinting while using UNLOGGED tables).
I highly recommend MemSQL (now called Singlestore) for a really polished distributed relational database that combines OLTP with OLAP columnstores.
I haven't published the benchmarking results anywhere yet, but I will probably do some conference talks on it once that is a thing again.
Didn't look into Druid in detail, but did try out Hive. Both of them look more suitable for cases where there is significant engineering effort in developing the data ingest and structuring pipeline. I wouldn't recommend either to a small team. With a measly triple digit TB database size both seemed overkill.
I think zedstore[0] might be something that could help here. I've mentioned it in the past but one of the best things about postgres is it's extensibility, and if timescale rides that wave (and maybe contacts zedstore to get this integration started early) it could be awesome.
[0]: https://blogs.vmware.com/opensource/2020/07/14/zedstore-comp...
> Didn't look into Druid in detail, but did try out Hive. Both of them look more suitable for cases where there is significant engineering effort in developing the data ingest and structuring pipeline. I wouldn't recommend either to a small team. With a measly triple digit TB database size both seemed overkill.
Thanks for this -- I haven't tried it yet at all but will try to remember this. ClickHouse was already first on my list for hobbyist->enterprise scalability but this cements it.
[1] https://github.com/lomik/graphite-clickhouse
[2] https://valyala.medium.com/how-victoriametrics-makes-instant...
[3] https://valyala.medium.com/measuring-vertical-scalability-fo...
- https://valyala.medium.com/measuring-vertical-scalability-fo...
- https://valyala.medium.com/when-size-matters-benchmarking-vi...
- https://valyala.medium.com/high-cardinality-tsdb-benchmarks-...
My bet for Timestream is that is built on top of DynamoDB. There are many potential indicators (1KB writes, throttling) and some clear ones(pricing follows the exact proportion up to the 4th decimal if you compare across regions, for example), it supports "unlimited" scalability, batching, etc. That reads are eventually consistent may be because they are computed over a GSI. It would be interesting if true as it is cheaper for writes than DynamoDB (on demand, which is the model Timestream has).
Plus there is a reasonable amount of mindshare and possibly market opportunity into offering time series on a serverless database like DynamoDB.
If this were true, it would also mean that Timestream, as hinted in the post, is more performant when accessed more in parallel (as DynamoDB itself is by design a massively parallel multi-tenant infrastructure).
The one thing that threw me was querying. The limitations felt more Athena-like than PartiQL. And the billing based on scans felt almost like Redshift Spectrum. I mean, S3 is infinitely scalable, right?
It's just impossible to say right now, eh?
I wouldn't say is impossible to say. Maybe impossible with 100% certainty, obviously, but for me it's quite clear ;)
This largely accords with my experience with DynamoDB (a few years ago) as well - sql-class data management tooling was just missing.
Timescaledb achieved 6000x higher inserts, 5-175x faster queries, 150x-220x cheaper
Benchmarks open-source, methodology in post.
I'd like to see those numbers 2-3 years from now.
Unless I'm missing something this is not an apples to apples benchmark. TimescaleDB is running as a single node without any replication whereas Amazon Timestream is replicated[0] to three AWS Availability Zones for durability. I've only skimmed the TSBS[1] repo and the start/stop scripts for TimescaleDB. Can someone confirm this?
0 - https://aws.amazon.com/blogs/aws/store-and-access-time-serie...
EDIT: Also your response contradicts what is in the blog post:
> 1 remote client machine, 1 database server, both in the same cloud datacenter
> Disk Size: 4.8TB of disk in a raid0 configuration (EXT4 filesystem)
(On some benchmarking equipment on Digital Ocean, it's advertised as "multiple racks", but managed to reduce blast radius.)
> 1 remote client machine, 1 database server, both in the same cloud datacenter
> Disk Size: 4.8TB of disk in a raid0 configuration (EXT4 filesystem)
Both those statements lead me to believe it's a single server with locally attached SSDs in a RAID0. Which is it?
I know benchmarking is hard, and it's difficult to test certain aspects of Amazon Timestream due to it being a managed service, but I really think these details need to be firmed up to make sure you are comparing apples to apples. TimescaleDB seems like a cool product.
Another suggestion I'd make is to run the Amazon Timestream clients across multiple AZs if you aren't already. The blog post doesn't mention whether all the t3 instances are in the same AZ or not.
They were all run in the same AZ. This brings up a good point, however, that we discussed internally when the first results came back. If there are tricks like this that might improve performance, it's not (currently) spelled out in the documentation so there's no way to know that. And we reached out for help in various forums with no response.
It's worth noting that since we performed this analysis, Amazon did release their own tooling for a similar benchmark and created a post[0]. While neither it, nor the tooling documentation[1] specifically spell out how many threads or instances they ran to achieve their results, it's hard to draw an apples-to-apples comparison. It does reveal that they used (had to use??) an m5.24xlarge instance (96 vCPU,384GB) to run their tests. As discussed in the article, one much smaller t3 instance was able to add >1 million metrics/second into TimescaleDB running in Timescale Forge.
[0] https://aws.amazon.com/blogs/database/deriving-real-time-ins... [1] https://github.com/awslabs/amazon-timestream-tools/tree/mast...
Just want to make sure the right numbers are being used! Thanks!
(Post author @Timescale)
For example my data looks like (row_id, user_id, time, data), where row_id is a unique ID, and time is a non-unique time field. Timescale will refuse to create a hypertable like this because it causes partition issues.
Otherwise, you are correct in that we require your partitioning keys to be at least _part_ of unique constraints; otherwise, we'd need to build global indexes across all your chunks (which would inhibit scalability)...this would only be worse with multi-node =)
Partitioning on the row_id by making it an ordinal instead of a UUID could work, however I feel I would be missing out on TimescaleDB's advantages for querying based on the `time` filter?
Consider that my main queries are:
* DELETE FROM t WHERE row_id = xyz AND user_id = xyz
* INSERT INTO t (...)
* SELECT FROM t WHERE user_id = xyz AND time > xyz and time < xyz
That first graph tells you everything you need to know. Time to insert a billion events:
TimescaleDB: 5 min
AWS Timestream: ~2 weeks!
I think the team at AWS that built it should be reassigned and contractors or an A team brought in to try and salvage it. Because they've built themselves a very expensive lemon and they're trying to sell it to the world as a Cadillac. The fact that they did not detect this themselves before or since releasing it upon the world demonstrates to me that they're incompetent and have to be replaced.
I was using MySQL at the time, and did a performance comparison of how long it took each database to dump and load a snapshot of a 100GB database.
MySQL took 3 days.
PostgreSQL took 30 minutes.
I posted on a MySQL forum asking why the drastic difference, and they tried to hand-wave it away by saying I had a lot of indexes. Well yeah, databases have indexes. I shouldn't expect my MySQL database transfers to be quick if it has an index? Features shouldn't cripple the application for routine tasks.
[In Postgres] if we have a table with a dozen indexes defined on it, an update to a field that is only covered by a single index must be propagated into all 12 indexes to reflect the ctid for the new row.[1]
[1] https://eng.uber.com/postgres-to-mysql-migration/That's a knee-jerk assessment. You don't know the first thing about Timestream's development and we are only a few years into the development process of what I believe is a novel architecture for a timeseries database. Give it another couple of years and then we can take a look at how Timestream is doing.
One of Amazon's advantages is the ability to plan and execute on a longer time scale than their competitors. They can take approaches that take longer to bear fruit but are better in the long term.
Timestream is a scalable timeseries database with separated compute and storage built for extremely high volumes (or that's my guess given the architecture offloads cold data to magnetic storage). Whereas Timescale seems to be timeseries functionality added to Postgres - which means they are probably going to hit scaling (perf/cost) issues once they need to work at petabyte scale. And my guess would be that their coupling to Postgres is going to make it painful to build for that scale. They also claim to offer separated compute and storage, but based on the pricing model it seems to be that you can change your CPU+Storage configuration, not an ephemeral design where you only pay for compute when it's needed - very different from Timestream's ephemeral pricing model. This is a pure guess, but I would imagine that Timescale being built on top of Postgres is going to make a truly serverless SaaS (i.e. pay for only what you use) very difficult to build.
Time will tell, but this is sort of if MySQL benchmarked writes against Snowflake 5 years ago and then thinking that everyone at Snowflake should be fired.
That's like making a car with square wheels. The only logical thing for management to do to a team who delivered that is disband them - because something is horribly wrong.
Now to be completely fair they do get good ingest performance if you open thousands of connections to send the data over. So they can probably fix it. I still wouldn't touch it with a stick though, that's not the only issue they have at present.
For the record, in reviewing HN conversation tonight I saw this and realized it was an incorrect quote of the article. Totally honest mistake I'm sure, but I wanted to set the record straight.
We spent a little over a week working at Timestream benchmarking. Trying different approaches to batching metrics for ingest, threading differently, running multiple EC2 instance, etc. to improve performance.
Once we felt like we were getting the best we could, we started our final ingest which we let run for nearly 40 hours (~2 days).
The other value is absolutely correct, however. We were able to ingest 1 billion metrics into Timescale in 5 minutes.
TL;DR; - TimescaleDB was tested with a cloud setup in Digital Ocean, separate client and server. We actually ran a second, unpublished, test into Timescale Forge from the same client(s) that we tested Timestream with. We did this second TimescaleDB test just to see if something was wrong with our EC2 instances or setup in other way. So, completely separate service offering with no VPC or anything - just a raw PSQL connection from client to server. The Timescale Forge instance was 1/4 the specs of the DO server from the published results (again, the intent wasn't to replicate the TimescaleDB tests all over again), and it still easily achieved 1.2 million/sec ingest from the same client computer where Timestream only achieved ~525 metrics/sec.
They have a lot of money and a bit of time.
It is very embarrassing that Timestream was delayed for so long only to lose so badly.
It is the THING.
Thanks the timescale team for making this happening!
One year ago I moved our old system to TimescaleDB, and I have been really happy since then, it's a really amazing product. Kudos to the development team!
Thank you for the links, I don't think we have the resources to implement something similar, but they are good food for thought.
Timescale works by creating a 'hypertable', which is an aggregate of a lot of smaller 'chunk' tables. These chunk tables are automatically split by date or incrementing id. This means that for queries that specify IDs or a date range within a certain range, you only have to query results within a few chunks, instead of looking through all the contents of the entire 'hypertable.' [1]
Timescale also offers some other things like compression which can save you up to ~96% disk space while also improving query performance in some cases. [2][3]
It also has something they call 'continuous aggregates' [4], which are similar to postgresql's materialized views, but do not require manual refreshing - they instead update periodically through an automatic background job. There is also a feature which builds on this called 'realtime aggregates' that allows you to combine the data within a continuous aggregate with the raw data in the tables that has yet to be materialized.
There are a lot more things besides that, but I think that's a decent overview of the major features it brings to the table. From a dev perspective these things all make the data and the database easier to work with (especially targeting timeseries data). There is an api reference [5] that has some of the other commands timescale adds, if you want to see some of the other things it can help you do.
[1] https://docs.timescale.com/latest/using-timescaledb/hypertab... [2] https://docs.timescale.com/latest/using-timescaledb/compress... [3] https://blog.timescale.com/blog/building-columnar-compressio... [4] https://docs.timescale.com/latest/using-timescaledb/continuo... [5] https://docs.timescale.com/latest/api
The two main things most developers will benefit from is how we manage the automatic partitioning of your incoming data (hypertables), something which is non-trivial to do yourself even though other tools exist for it. And because we do it with a time-based focus, we can be really efficient and smart about it.
Second, we've improved the query planner in PostgreSQL around the parts that relate to querying time-based, partitioned data, and provided special time-based functions. These improvements help you efficiently query data that time-series applications most often need. A quick example is something like "LAST()", which retrieves the most recent value for a given time-range. There are ways in SQL to do something similar (LATERAL JOINs or CTEs for instance), but they're usually slower and bulkier to maintain. When dealing with time-series data, getting the most recent value for an object is usually what you're doing the most often.
When you add those two foundational features, everything else that @drpebcak mentioned become amazing value-adds that you just can't get elsewhere.
Back in 2015, I'd architected and deployed a system for a AAA game that handled 24B events/day on launch without breaking a sweat, and supported 200ms round-trip ingestion-to-aggregation SLAs with no windowing (the protocol and ingestion layer did most of the heavy lifting: sequentially ordered _guarantees_ on events even when loadbalanced/connection migration meant no need for windowed batch ordering)... but the scenario for which it was designed was cut and we ended up using it for just 15m slices. :eyeroll:
Still, it was used by a dozen+ games, including a few more AAA titles, and still in use today, and portions of the tech have been cannibalized into other products. I still get the occasional inquiry about memory fencing or memory boundaries on Console X for the 5-15μs event generation API (improperly aligned memory could cause interlocked increment corruption!).
Annnyways:
I had an opportunity to chat with one of the founders at Snowflake in 2017? 2018? for a few hours. I tried to convey how imperative I felt true-realtime time series engines would be critical moving forward, an the reception was rather lukewarm. If they had been as excited as I, it'd have been one of the few opportunities to pull me away from my dream job.
I still feel the world will need this architecture, as we start moving towards more ML/AI driven decision making, and that the company which can get traction will be in a pivotal position moving forward.
Sometimes I wonder about feeling pressured to shift into Data & Applied Science to stay at that org (there just didn't seem to be vertical opportunities in the dev track). I excel in this job too, and I love what I work on... but dang sometimes I feel that the architect career path had even bigger impact potential. It was a fun couple decades. :P
The documentation on their website [2] is also very good.
[1] https://github.com/vincev/tsdbperf [2] https://docs.timescale.com/latest/main
While the comparison seems terrible for Timestream, as a customer who does not want to manage my databases I would love a similar GCP option, if the product had better tradeoffs.
It's also interesting that Timescale attributes this AWS product to their own licensing. They had some much discussed [0] developments on that front, and if it did in fact force AWS to build their own implementation, that seems like a win for opensource, but not so much for serverless users.
Slightly off topic:
I have a side project that generates ~10 daily metrics for ~350 entities. Around 150k readings currently (the number of metrics gathered per entity per day has increased since I started).
I'm pretty interested in these timeseries databases, but I think my use case is still on the side of using MySQL/Postgres out of the box. Not to mention that I can get by with archiving data > 1 year old.
Does anyone have any quick checklist or metrics to make decisions like this? When does it make sense to evolve from a vanilla RDB to a timeseries one?
Before coming to work at Timescale a few months ago, I spent 18 years managing products at two companies, both of which were time-based applications (utility billing/energy data & IIoT). In both cases, the app started small and everything seemed fine. We could usually get around performance issues with other hacks. But in both cases there was a tipping point because the original database (one relational, no NoSQL) just wasn't designed for the challenge as the app scaled.
So whether you manage it yourself in a smaller environment for now or try something like Timescale Cloud, I can (almost) guarantee that you'll thank yourself in a year or two. ;-)
I'm not a database expert in any way, but I've had to evaluate database options for both existing and new applications, and my "methodology" is as follows:
- Understand the workload
Is the database workload read-intensive, write-intensive, or both? Does the application require strong consistency? Does it require replication? What kind of replication? Can writes be batched, or are processed ad-hoc? What kind of read operations will you do, and what is the desired response times for those operations (ex. is 2s tolerable? 500ms? 20ms?)
- Scouting
Select a set of available options to benchmark using whatever criteria you may see fit, but following the previous step constraints.
- Benchmark
Create a somewhat simple model that fits your workload and models your real case, using the language you'll be using to build your application; Create benchmark code for the most critical/complex operations; Run it at least 5 times on similar hardware of what you expect production to be (this is important; eg. storage (both memory and disk) latency varies wildly between "my laptop" and a cloud provider) and record the results; Be sure to keep an eye on memory consumption, cpu and disk usage when running the benchmarks;
- Decide :)
Usual criteria:
* Price;
* Features (of course);
* Easiness of deploy or cloud provider availablity;
* Documentation quality;
* Driver quality for your application language (it may happen you already excluded a candidate on the previous step due to this);
* Query language familiarity;
* Performance;
* Adoption (translated to: can I expect a new developer to have the basic skills on this tech?);
10 daily metrics for 350 entities result in `10 * 350=3.5K` reading per day. This translates to `3.5K * 365=1.3M` readings per year. Such workload can be easily handled by any DBMS out there. There is no need to search for specialized time series database for this workload.
The need in specialized TSDB solutions arises when ingestion rate reaches a million of readings per second and the number of readings stored in the database exceeds hundreds of billions and trillions. Such a workload cannot be handled by general-purpose database, but it is easily handled by specialized databases such as VictoriaMetrics. Fun fact that it easily handles 10 trillions (e.g. 10K billions) of readings in a single-node setup - see https://victoriametrics.github.io/CaseStudies.html#wixcom
4) comparison with VictoriaMetrics
Performance is often transient and it isn't the only criteria used when choosing a vendor. Some enterprise customers prefer a single vendor, i.e., "one throat to choke". Others have very specific constraints: greenfield decision making is a luxury.
My impression is that Amazon responds to customer requests to solve specific problems and they make a decent effort to continuously improve the performance and other aspects of their services. At one point people said "no one gets fired for choosing IBM" and this switched to Microsoft and, rightly or wrongly, Amazon AWS now wears this crown.
I get that, but I would have thought it would make more sense to approach TimescaleDB for a licensing deal. That way, right from day one they get a mature product with incredible features and performance. And they'd win developer mindshare by supporting OSS.
In their mind, long term, that leaves money on the table. If they entered a licensing deal, they would have exposure risk if they ever wanted to replace it with an in-house product (and, given their corporate culture, an actual legal exposure, intentional or not, is a possibility).
See this HN discussion about Timescale's "cloud protection" licensing [0].
Of course, Timescale Cloud is available across 20+ GCP regions =)
I read that you can't easily migrate 9.6 to another version of PG via streaming replication until 10 :(
It also doesn't mention which license limitations the "community" edition has, I only found a blog post talking about them https://blog.timescale.com/blog/building-open-source-busines....
I would have expected one place which compares all available options.
- Community features: All features labeled with "community" on here: https://docs.timescale.com/latest/api. Only restriction is that you can't offer them as part of a Timescale-DBaaS.
- Open-source features: Everything else, licensed under Apache 2
- Enterprise features: No longer exist, we made all of them free/community earlier this year: https://blog.timescale.com/blog/building-open-source-busines...
Hope this helps.
Will pass along the feedback to the team :-)
65 rows in 80.31s Execute:80.16s Network:144ms Total:80.31s
Is this the fastest? Or having a <5ms would have a huge impact?
I've also heard gcp can be a real mess.
Is Azure any good yet?
We also prefer/buy services from companies who build OSS like in case of ELK, we, rather than hosting and doing it by ourselves. We are more than happy to pay bit more for hard work their team has done.
For me, my list is: IAM, STS, Route53, EC2 (and all of its children), S3, and SQS.
(By the way, hello fellow T-bird.)
Seriously, If you need to do this at scale, use Druid. It's much more efficient for time-series data.
There is a branch for PostgreSQL 13 support available for beta testing if you want to build it yourself. It's important to us and we'll focus on completing that integration soon!
I'm having a hard time understanding the cost comparison without details of the above. Are they saying that hosting your own cluster of timescaledb nodes within EC2 still comes in cheaper than timestream? This seems impossible, depending on the instance types of course.
AWS use to be great when it first came out in 2006ish, but now it is just because a pain to work with.
EDIT: especially considering all the other options are available in VPS now days.
There are two main points here.
1. To complete this benchmark workload, it took less than an hour in two different environments (Digital Ocean self-managed & Timescale Forge fully managed) to ingest 1 billion metrics and run all 30K queries. It took us a week of work (testing, modifying code to try and make Timestream better) to get 40% of the metrics into Timestream and then query it. 1 hour vs 7 days.
2. If you look at the bill/costs, the main driver was querying. We (attempted) to run the same 30K queries on less than half the data (410 million metrics) in Timestream and somehow scanned 21TB of data. I have no idea why and there's nothing we could do to change it.
As a developer, that's going to be your biggest unknown. If you're ingesting millions or billions of metrics a day and querying it with a real application, you could really get hit with crazy query costs.
With a more traditional server architecture, at least you know your day-to-day costs and can set a known capacity to achieve the performance you need (or scale in understandable ways when you need it)
"Amazon has a history of offering services that take advantage of the R&D efforts of others: for example, Amazon Elasticsearch Service, Amazon Managed Streaming for Apache Kafka,[...]"
And at the end of the article they promote how they themselves rely on other people hard work:
"TimescaleDB uses a dramatically different design principle: build on PostgreSQL. As noted previously, this allows TimescaleDB to inherit over 25 years of dedicated engineering effort that the entire PostgreSQL community has done to build a rock-solid database that supports millions of applications worldwide."
In the end, it sounds like they are doing exactly what Amazon is doing with open-source.
They criticize AWS for making money on Elasticsearch for example, AWS is "taking advantage of the R&D efforts of others". So Amazon is making money on a "serverless" / cloud experience. At the same time, it is known that Amazon is contributing back to Elasticsearch [1]. To that regard, I find their business model really similar to the one AWS is relying upon.
[1] https://www.techrepublic.com/article/aws-contributes-to-elas...
Some people would want k8s, some docker swarm or whatever, some an aws config, others ansible, etc etc etc.
If I want a managed service, I _want_ to pay for that. The price includes people responding to pages and fixing problems, as well as fiddling with configs.
And I'd much rather be paying that money to a small (relatively) open-source company than a behemoth like AWS.
Is this not what S3 is? You are free to use Elasticsearch and handle everything yourself, but if you want a managed service, you can use AWS. They are "attacking" the S3 offering, I still struggle to see any difference with them hosting Postgres.
> And I'd much rather be paying that money to a small (relatively) open-source company than a behemoth like AWS.
I am 100% with you on this, I do not want to defend AWS nor do I want to promote their products. The vendor lock-in situation you are in when using AWS is pretty bad and quite scary... An yes, I agree that the way they monetize open-source software is questionable.
Also, in this case, Timescale actually has a pretty forgiving license[0] as long as you are not a add-nothing-aaS-provider, perhaps more than it should be, which I've asked about before[1]. Even before that change was made, running just the community edition as a add-nothing-aaS-provider would have been an improvement on the status quo, given how soundly it thrashed some other solutions in the past (ex. Influx[2]) and what you can do it (promscale[3]).
I know it can't be all roses, nothing is, but I don't think they've put too many feet wrong so far.
[EDIT] - I should note that on the scale of "contributing" to Postgres, the scale heavily tips in favor of 2ndQuadrant, EDB, and Citus as obviously they have the most committers and core team members. All those companies are to be commended of course, they're making postgres work as businesses and keeping it free while also improving it.
[0]: https://news.ycombinator.com/item?id=24579905
[1]: https://news.ycombinator.com/item?id=24585564
[2]: https://blog.timescale.com/blog/timescaledb-vs-influxdb-for-...
Again, you're free to use it for any project, all Community features, wherever you want. The only thing you can't do is run a DBaaS for TimescaleDB. Seems like a fair tradeoff, right?
Amazon primarily runs and monetizes closed-source, SaaS-only managed services.
TimescaleDB instead is implemented in the open as an extension to PostgreSQL, and enriches and benefits the broader PostgreSQL community by unlocking a new use case (time series). The Postgres extension framework exists very much for this purpose, for projects like TimescaleDB to contribute back without needing to “pollute” mainline with domain-specific features. Most of TimescaleDB’s code (and all development for the first few years) is Apache 2, and all features are free for anybody to self-manage.
(Disclaimer: I work at Timescale)
The article makes a lot of points really well, and if I had 8 highly qualified people who had nothing else to do on my team working for free (or on salaries that were a rounding error in my actual and opportunity cost budget), I’d certainly get them to learn and use Timescale. But I don’t right now, so I’m going to use the option that gives me a time series DB with three or four clicks, gets the job done, and charges me for deployment and maintenance amortised across thousands of other customers.
Having these arguments purely on cost of service terms is a very slippery slope. Yes, it’s less dollars paid in hardware if I put my life on hold and learn how to configure this. Or pay someone else to do it. Yes, it’s cheaper if I order parts on newegg, build a server, drive to a colo and install it in the rack, and then drive there again each time something breaks. And yes, it’s even cheaper if I run it off a Raspberry Pi duct taped under my desk. No one is disputing these things.
AWS and other cloud providers sell peace of mind, acquisition and procurement speed, professional maintenance and remediation, deep integration, faster development and testing, and infrastructure management - with servers thrown in for free. Timescale sells time series database software. This is not a valid comparison.
https://www.timescale.com/timescale-signup
2. Beyond operations, your team probably already knows how to use much of TimescaleDB, if they know SQL and PostgreSQL. AWS Timestream, on the other hand, introduces a bunch more of strange gotcha's, as evidenced in the blog post (even from the weird SQL hoops that you need to jump through.)
Not sure about "comparably provisioned re: IOPS" given that you don't know about that at all with AWS Timestream, but our blog post reports the performance/cost of a suitably provisioned Timescale Forge instance (8vCPU / 1TB storage).
100GB ingest then query benchmarks took far less than 1 hour with Timescale Forge @ $2.18/hour.
The "consumption based pricing" for AWS Timestream for same benchmark took $336.40.
Post-reading:
"faster queries via continuous aggregates". So is this it? I couldn't find how tables / materialized views were created in the source though [1].
TimescaleDB is probably a very good product (and pg-compatible!), but producing such articles hiding the usage of a magic feature is sort of dishonest. Why not make an article directly on the power of the feature? It's hurting their brand reputation a bit.
I'm sorry you feel like we were trying to be dishonest in the post. On the contrary, we put a lot of effort (and 7,000+ words) into trying to explain everything that we did - just as we've done with other benchmarks which others have linked to.
The TimescaleDB test did not use continuous aggregates for these test, only raw time-series data stored in hypertables.
For each database, we (and other contributors) do our best to use features in all cases that take advantage of the DB. For TimescaleDB, a function like LAST() happens to be really powerful for most workloads and is really, really fast. That's not cheating, it's using the software properly! :-)
The SQL that we generate for each database can be examined here (https://github.com/timescale/tsbs/tree/master/cmd/tsbs_gener...) and as an open source project, anyone is free to contribute!
I've been reading about your columnar compression pipeline [1], and it sort of makes sense if the comparison is against a regular row-oriented DB. AWS Timestream must really be doing something wrong here, or serving an entirely different use case.
5-175x faster queries and 150x-220x cheaper I do get it. But 6000x higher inserts does not make sense to me. It is insane, and literally unbelievable to me.
Storage savings are at 96% for "IT metrics (DevOps dataset from TSBS)", so it should be closer to 25x higher insert rate. Where is the missing 240x? Is this from some distributed replication overhead? Is this from local vs remote insertion? Is this from bulk inserts vs per row?
Anyway I wanted to thank you for your kind efforts in writing the blog post and providing answers here; and for the patience that you show to the audience here, me included.
[1] https://blog.timescale.com/blog/building-columnar-compressio...
In the end, if you read the article (and not just the headlines - not saying you are, but it's easy to see 6000x and latch on to it), the comparison is absolutely focused on this one, pretty straight forward use case (although we normally run 5 different scenarios):
From one client, given a specific kind of workload (100 hosts, 10 CPU metrics every 10 seconds for 30 days = ~1 billion metrics) - how fast could we save the data. Most other time series databases at least perform marginally well with the same setup... load data with one client.
But Timestream just doesn't seem setup to work that way. Some of the responses today imply that we need really large clients with thousands of threads to get those speeds. And that might work if we kept going and spent more time and significantly more money. We just haven't ever had to do that before.
If your use case better aligns with what Timestream offers, then it might be a great product for you. Given some of the many other concerns we discovered along the way, it doesn't yet seem like the time to jump in.
All the best!
https://aws.amazon.com/blogs/aws/store-and-access-time-serie...
"This is made possible by the way Timestream is managing data: recent data is kept in memory and historical data is moved to cost-optimized storage based on a retention policy you define. All data is always automatically replicated across multiple availability zones (AZ) in the same AWS region. New data is written to the memory store, where data is replicated across three AZs before returning success of the operation. Data replication is quorum based such that the loss of nodes, or an entire AZ, does not disrupt durability or availability. In addition, data in the memory store is continuously backed up to Amazon Simple Storage Service (S3) as an extra precaution."
Feels like apples to oranges comparison, as the consistency models are really different. Then again, I woul definitely optimize for performance on most timeseries use cases. Different products with different features baked in.