Timescale Announces New Database Cloud
blog.timescale.com
blog.timescale.com
Apache Arrow, K8s, ML analytics have given rise to another DB War.
The end of NoSQL was the realization that SQL had a good reason for existing in most cases. Now we have massively distributed SQL in many flavours. I wonder what the hard lessons will be this time?
I'll wager small data companies will be spending good money on vastly overpowered engines... I wonder what else tho?
[0] https://blog.timescale.com/blog/building-open-source-busines...
As a result, NoSQL never really catch on in the analytics/BI crowds, SQL is always king there, if you discount Excel :)
I think the lesson might be- be careful who and when you take funding from. All these managed services smack of VCs looking for ARR growth and a big exit...
Supabase as an add on to Timescale Cloud or Timescale as a supported alternate provider for Supabase would be compelling for get-it-done devs and teams out there
Of course grafana too, but that goes beyond just PostgreSQL.
If we can't manage we'll eventually do something else for timeseries data - eg, we could do a Foreign Data Wrapper to Clickhouse tables
By the way I can see there are some already but I do not use them https://github.com/Infinidat/infi.clickhouse_fdw/ https://github.com/adjust/clickhouse_fdw
It would be beneficial to couple transactional data (Postgres) with OLAP data (CH). There are a few other things coming in PG which might mitigate the need for CH too (pluggable storage, zheap), so we're not going to rush this part.
If not, it may be worthwhile to migrate to a host that Does support timescaledb, if not timescales managed product itself
By this i assume you want columnar access along time dimension.
There are a bunch of columnar options out there ( timescale being one). you can operate hybrid row + column access.
https://www.citusdata.com/blog/2021/03/06/citus-10-columnar-...
https://swarm64.com/post/postgresql-columnstore-index-intro/
Can we think of timescale as OLTP for timeseries data with gurrantees from postgres and clickhouse as OLAP for timeseries data?
Continuous Aggregates is a neat feature though.
But for time-series workloads, we've found that the results are quite close, with TimescaleDB outperforming Clickhouse for many different types of query workloads. We'll be sharing our results soon.
I wonder how the storage works: https://docs.timescale.com/cloud/latest/scaling-a-service/#p...
Seems like to scale it separately it would need to be EBS not local instance storage? I wonder if magnetic or SSD? That does constrain the performance, especially for queries.
To tie GP2/3 back into the serverless vs. DBaaS concepts we are looking at auto-scale for IOPS/Throughput performance while also allowing more direct access such that a customer could control performance via APIs to manage on your own.
(timescaler here)
So for most users, they can "set it and forget it" (easy, scalable): Launch a default cloud database, which starts out small, autoscaling on by default. As they start inserting data, the system automatically detects when it starts to approach current capacity and automatically increases (without any downtime) to more capacity & IOPS.
But for the power user, they have greater control. What that means is that you can manually resize, and as mentioned above, we'll likely also give you independent control over IOPS (independent of capacity). That's the flexibility and control we think developers also want...or at least know is there is they need it.
(Timescaler & post author)
Some kinds of database workloads would run more efficiently on the local NVMe storage. But that does come with lots of operational considerations.
That said, I'd be curious to hear about other folks experiences.
Disclaimer: I work at Timescale.
[1] https://hasura.io/blog/using-timescaledb-with-hasura-graphql...
CoralCDN ran from about 2004 - 2015. Eventually, its need was pretty much negated by the rise of free CDN services (e.g., Cloudflare) or just lots of SaaS services that supported user-generated content. For example, in late 2004, many of the amateur videos of a large Indian Ocean earthquake & tsunami were shared using CoralCDN, but that soon went to YouTube. Podcasters like "This Week in Tech" were using CoralCDN, but those went to freemium podcasting services. And so on.
What eventually took CoralCDN "down" was that the academic platform on which the ~1000 servers ran, PlanetLab (https://planetlab.cs.princeton.edu/) was end-of-life'd.
But taking a step back, the original thinking behind CoralCDN was as a peer-to-peer CDN, but there were a lot of web security issues that actually made that difficult. If folks are interested, I talk about some of these issues in this 2009 retrospective [0], also also outline a browser-based P2P CDN in this workshop paper that could address (and actually make it P2P but secure)[1]. But still, I think the economics of CDNs (and transit costs) have just changed, such that most of the p2p architectures just don't make sense today like they did in early 2000s.
But thanks for the kind words!
[0] https://www.cs.princeton.edu/~mfreed/docs/coral-nsdi10.pdf
[1] https://www.cs.princeton.edu/~mfreed/docs/firecoral-iptps09....
I'm a fan of Timescale but definitely not a fan of "serverless" as a phrase.
That phrase is just abstracting "other people's computers" one additional degree, in a borderline meaningless way.
There's still a server. How is that serverless? It's just a server managed in a more indirect way.
Please convince me I am wrong here.
I'm pretty sure everyone is pushing it so they can make money from SaS. But if this allows them to give it away for free, then I'm all for it ( just won't be using it ).
But for better or worse, it seems like the industry has adopted it.
So, we're trying to explain actually how our vision is different.
In that we don't want to hide developers completely from their services behind these black-box abstractions. But provide something that's similarly easy and scalable (and automated), but allow developers greater control, flexibility, and understanding when they want it.
I also used to hate the phrase. But now I get it. And it's not about marketing, but about the idea that the developers _doesn't shouldn't have to worry about the server_.
For example, when you run a DB on a classic DBaaS, you have to worry about CPU, memory, storage provisioning. But when you use a classic SaaS service (e.g., Stripe, Twilio), you don't think about servers at all - but rather just how much you are consuming.
The Serverless model for databases is aiming to apply a SaaS like experience but for DBaaS. Where you don't need to think about provisioning, but about consumption.
What we (try to) do in this article is to push beyond that serverless paradigm. We believe there are real drawbacks to the "serverless black box" architecture that the industry is building, and what (we believe) developers need (including ourselves) is a more a "transparent box."
Hope this helps. But "serverless" also feels a little like "horseless carriages". I suspect in 5 years we'll have a better term that describes this concept for what it _is_, not what it _isn't_.
From the blog post: "So today’s serverless data platforms are not familiar or flexible. But further, black boxes are never truly easy and worry free: you never know if there are any skeletons lurking in the proverbial closet, just waiting to cause your service to fall over."
(Timescale employee here)
Call it managed DB as a service and you're more honest in your product offering. You're suffering marketing-speak in lieu of honesty for a technical product - bad plan, imo.
Your target market will understand.
Ignore the HackerNews naysayers who are stuck in the past.
Language changes. "Serverless" clearly doesn't mean "computers aren't involved", but "the computers are someone else's responsibility". It's a marketing term, just like "cloud", and it's unlikely to go away.
As with "hackers vs. crackers", "Linux vs. GNU/Linux", "copyright infringement vs. piracy", and other similar scenarios, the ship has sailed here.
I also like cloud services like Timescale. I just want a database of a certain size. How it runs? I don't care at all as long as it works.
The problem is that while you don't want to know you end-up having to. As soon as you hit the limitations you have to workaround it. As soon as what you are doing becomes expensive for reasons you ignored before, you have to go back and rewrite. The only thing is that because it's a Blackbox, you also have to guess and poke it until it does the thing you want.
You just have to "know" that these things happen, and because it's all serverless, it's essentially invisible and out of your control.
You still face constraints, and the platform design is informed by the hardware it runs on.
Am I correct? If so what is the plan for existing customers of those services? Especially since forge didnt support other clouds than AWS last time I checked.
We have two cloud products, Timescale Cloud (which is what this post is discussing), and Managed Service for TimescaleDB (MST), which is what you are also referencing.
Also, as we say in the post:
Some of you may remember that we launched the first “Timescale Cloud” 2.5 years ago, as the world’s first fully-managed time-series database-as-a-service on AWS, GCP, Azure. That product is alive and well, and fully supported as before, but is now called “Managed Service for TimescaleDB”.
We're investing in and maintaining both. They are just different products, depending on what you are looking for.So here's an "RFP" - any suggestions for a better term to describe this consumption-based experience, where you don't need to worry about servers (and ideally, you don't pay for what you don't use)?
To me, serverless is fine. Yes, it’s a buzzword, but it makes sense.