Introducing ReadySet
blog.readyset.io
blog.readyset.io
With that said, it looks like cool tech and I read Jon's Rust for Rustaceans which serves as a stamp of quality for this even if I haven't tried it yet!
I've been following the space since a bit of time, and I must say it's exciting. To me this is the future of apps where the Truth lives server-side, and everything reacts from there; With partial state evaluation lowering resource consumption to a minimum.
Kafka Streams and Apache Flink seem to be focused on real-time analytics, and I wish they'd get there to stimulate the space.
Are you affiliated with ReadySet?
Some context: https://twitter.com/jonhoo/status/1511401461669720068
Basically, I co-founded the company around the time I graduated, but had had my fill of database research after six years of PhD. So I joined AWS to work on Rust while Alana (the CEO) took on leading ReadySet.
Unfortunately I'm not privy to whatever improvements ReadySet has made in the past two years, so I can't comment on differences between ReadySet and Materialize. Perhaps Jon can, though!
Our official docs also have an aptly-titled: "what's the catch?" section: https://docs.readyset.io/concepts/overview#whats-the-catch
Which mean read might return stale data because of the replication lag.
It will also increase the load on the server you read the replication log from.
But the primary database could dump the transaction log to (S3/kafka) and have ReadySet instance read it from there instead of directly from Primary database.
So for a read mostly website (hackernews/reddit) this is indeed a free lunch.
I am curious how they handle queries that would overflow local main memory, like if I just had a PK lookup on a 10TB table you obviously can't store all that in RAM, and would still need to do some form of cache invalidation.
Today, machines are super huge (in terms of compute cores, memory and iops for storage and network) and a single Mysql or PostgreSQL database can do a lot of work. This makes it much much easier to build apps that don't have as many users at Internet scale – that is pretty much all enterprise apps – without resorting to distributed databases.
In Internet scale consumer app domains like e-commerce/delivery or fintech where relational databases are used heavily, most queries would have strict correctness requirements and won't tolerate staleness. Also, most query-results would be highly specific to each user and won't have much cache hits. Also, apps in general are increasingly personalised and have fast changing content.
In terms of technology evolution, I see people moving from single large machine databases to distributed sql databases as their use-cases scale.
And as distributed sql databases mature, I expect they will get built-in capability to generate user-defined materialised views with flexibility to manage their placement w.r.t class and number of machines to compute and serve them etc.
I think a lot of applications could probably benefit from this if they were built with a data model in mind that leverages it properly. But if you 1:1 migrate your code that relies on ACID transactions over to something with Strong Eventual Consistency... yeah, that's gonna be a bad time.
We frequently (~ once per second) run queries over relations that are increasing at a rate of ~100 rows per second (append only, no updates).
Could this cause any performance concerns for ReadySet? How much control do we have over the frequency of reconstruction of cached data based on the flow graph?
At least with scaling replicas or having a dumb cache layer it’s easy to understand the system.
I love the way you explain what's ReadySet. Congrats on the launch.
Huh? At a 3 year multiple that is less than $3,000 monthly revenue. I'm sure a SaaS can provide way more value for $3k/mo before you start running into scaling issues
Had a similar idea a few years back and it's nice to see it turn real. Congrats on the launch.