My team ingests petabytes of data each day in S3 that is then queriable in Athena, and it supports all the same features that are mentioned here, such as interfacing with other types of datasources like an existing PostgreSQL database.
My team ingests petabytes of data each day in S3 that is then queriable in Athena, and it supports all the same features that are mentioned here, such as interfacing with other types of datasources like an existing PostgreSQL database.
The potential downside is unpredictable query performance. For example, suppose you have a query that calculates daily statistics from your time series data. The query takes 1 second to execute when you run it for yesterday's data, but 1 minute to run for a day six months in the past, because you've inadvertently shifted your work onto the slower storage layer.
Our experience is that isn't a common use case for Athena, while it is the primary use case for TimescaleDB.
So performance matters in the sense that querying data is our core offer, but we're talking about multiple seconds here, not milliseconds.
I think the distinction makes sense then, thanks.
I haven't heard great things about Athena, but I'd be curious to hear more about what works.
Then in terms of cost it's around $0.5 per query maybe? The average is a very bad number because most queries will be something like $0.0001, but then some will be hundreds of dollars. But that's the only number I have off the top of my head.
It shouldn't be hard to use another block level storage.