Dragonflydb – A modern replacement for Redis and Memcached
github.com
github.com
Also as a Redis replacement, it's not clear what durability is offered, and for most Redis use cases this is close to the first question
The tradeoff the way I see it - one needs to implement 200 Redis commands from scratch. Besides, I think DF has a marginally higher 50th percentile latency. Say, if Redis has 0.3ms for 50th percentile, DF can have 0.4ms because it uses message passing for inter-thread communication. 99th percentiles are better in DF for the same throughput because DF uses more cpu power which reduces variance under load.
Re-durability - what durability is offerred by Redis? AOF ? We will provide similar durability guarantees with better performance than AOF. We already provide snapshotting that can be 30-50 faster than of Redis.
If I recall ScyllaDB has some excellent examples of demonstrating this particular tradeoff visually. A simple option would be a scatter plot where X = latency, Y = load or similar, with points coloured according to the system under test. Probably there is a better option, but this would likely be enough to sell me at least
op r6g c6gn c7g
set 0.8ms 1ms 1ms
get 0.9ms 0.9ms 0.8ms
setex 0.9ms 1.1ms 1.3msOf course, it is also possible there are situations where it doesn't perform as well.
I'm not hip to how much new stuff is backport-able, so this may preclude Ubuntu 20.04, for instance. You lose the "LTS" part if you compile your own kernel, if you manage to make it functional at all.
Note: I never use kernel modules due to the issues rhel/debian and I have had with such things in the distant past.
I'm not sure that concern is justified. It seems io_uring was pushed as part of the 5.1 linux kernel release, and Ubuntu 20.04 LTS seems to have been shipped with 5.4.
https://packages.ubuntu.com/search?keywords=linux-image-gene...
Also, a quick Google search pointed to io_uring patches for Ubuntu 18.04 LTS.
It does sound like extendible hashing might have downsides in some scenarios also.
Curious: Why BSL? Why not open core [0] (or xGPLv3) like what most other commercial OSS projects seem to be doing?
https://www.gnu.org/licenses/license-list.en.html
https://opensource.org/licenses/alphabetical
That's a big red flag.
Still, I really appreciate that you didn't choose a copy-left license.
On the license front, what is the "change license" clause listed? It says something about changing in 5 years. Does this mean it will become Apache licensed in 2027? Why would you put that in there?
In 5 years the initial version becomes Apache 2.0 then the next version and so on and so forth. CockroachDB uses similar license. MariaDB uses that, Redpanda Data and others. You are right that acronym is confusing - it's not Boost license, it's Business License. Every major technological startup turned away from BSD/Apache 2.0 licenses due to inability to compete with cloud providers without technological edge.
However, I'm sad that instead of going with an Open Source license that protects against that, you're using a proprietary license. That alone is a nonstarter for many users, not because they want to compete with you but because they want to protect themselves and make sure they have a firm foundation to build on.
Much of the software you're citing as examples moved from Open Source to proprietary, harming their users in the process, and causing many users to seek alternatives.
No, there are plenty that still use permissive licenses.
GitLab uses MIT and a custom license for EE: https://docs.gitlab.com/ee/development/licensing.html
Deno uses an MIT license and has some secret sauce that is currently just in hosted services AFAIK: https://github.com/denoland/deno/blob/main/LICENSE.md
PlanetScale has hosted services and an open source tool called Vitess which is Apache licensed: https://planetscale.com/ https://github.com/vitessio/vitess
Finally Redis has a BSD licensed core, a source available license for additional modules, and a closed source license for enterprise. https://redis.com/legal/licenses/
Was excited to see the project but now seeing it is not Open Source it means 1/10th of value
Accomplishes the goal of preventing a cloud provider from stealing customers, but also ensures customers don't get caught in an "always tomorrow" trap when the deadline comes and the company realizes it only hurts them to fully share it.
Seems to align all interests pretty nicely.
(I'm as big of an OSS supporter as anyone, but we can't pretend we still live in a time where Google / Amazon / modern-Microsoft don't exist)
I ask this because I'm unsure if AWS Redis has any modifications on top of the Redis software itself, which would affect the speed, or even make it a bit slower. For example I know MS Azure's version of Redis restricted certain commands, and from a quick search AWS does something similar: https://docs.aws.amazon.com/AmazonElastiCache/latest/red-ug/...
(edit: added "affect the speed" for clarity)
Where's the benchmark compared to memcached?
Several years ago there was memcachedb, which could flush stuff to disk. While this operation was expensive, it was also useful, because you could restart instances without being overwhelmed by missing keys (data).
For the latter: your application quickly grinds to a halt if you need to build your cache from ground up after some kind of crash. This is a deal-breaker for many.
static constexpr unsigned NUM_SLOTS = Policy::kSlotNum;
static constexpr unsigned BUCKET_CNT = Policy::kBucketNum;
static constexpr unsigned STASH_BUCKET_NUM = Policy::kStashBucketNum;
NUM_, _CNT, _NUM, three different prefix/suffix for what seems to me like the same concept. That just tickled my inner nit-picker.Why not reuse seastar framework?
Can you describe your distributed log thing? Is it like facebook-logdevice or apache-bookeeper?
With multi-threading you need to think about all things holistically. How you handle backpressure, how you do snapshotting. How you implement multi-key operations or blocking transactions. So you need special algorithms to provide atomicity, you need fibers/coroutines to be able to block your calling context yet unblock the cpu for other tasks etc. All this was designed bottom up from scratch. Seastar could work theoretically but I am not a fan of coding style with futures and continuations - they are pretty confusing, especially in C++. My choice was using fibers - which provide more natural way of writing code.
I have not designed the distributed long thingy. Will do it in the next 2 months.
Out of curiosity, are you discovering any new bottlenecks to performance outside of the software, given Dragonfly is able to process far more qps than most systems? I imagine the network and disk I/O could become stressed, but also I wonder if it breaks any assumptions of cross-core performance, hypervisors, etc. I know that cloud offerings typically mean that you can attach ginormous disk IOPS and NICs, but surely there are limits.
I was mostly running on AWS. In terms of hardware, for small-packets loadtests, most systems are constrained on throughput, i.e. number of packets per second. Some instances saturate on interrupts reaching 100% CPU on all cores and some can not even saturate the CPU and you will see that CPU is at 60% but you can not go beyond in throughput. The best systems network-wise are c6gn family types. They are also better than instances that other cloud provide. btw, you mentioned hypervisors... About 8 months ago I opened a bug on AWS Graviton team https://github.com/amzn/amzn-drivers/issues/195 - about performance issue they had on their instances at high throughput. Recently they issued the fix. I suspect it was in their hypervisor.
In terms of my software I found many performance bugs at those speeds. For example, using a default allocator is a big no. I use mimalloc for uncontended allocations. In general, you can not use mutexes and spinlocks at those speeds. Those will just cripple the system. Sometimes it can be very annoying since you can not rely on a 3rd party library without carefully analyzing its design. For example, I could not use openmetrics c++ library because it was not performant enough. Even to implement a simple counter, say to gather statistics for INFO command becomes an interesting engineering problem: With share nothing architecture, I use a lot of thread-local counters that I aggregate only when stats are pulled.
As a general note, I expect that Dragonfly will stay very performant with the tailwinds from recent hardware advancements. For example, c7g (Graviton 3) is much better than c6g and DF shows it.
See https://s3-docs.fd.io/vpp/22.06/developer/extras/vcl_ldprelo...
dragonfly is a linking of a library and a main file dfly_main.cc so without this file you will have the lib.
Redis is "Remote Dictionary Server". You gonna loose the remote part :)
Looks awesome so far, though!
I will continue working on DF. Primary/Secondary replication is my next milestone.
We plan to implement everything but your votes can affect the priority of the tasks.
We do provide atomicity guarantees for all operations like Redis! We use an algorithm from a 2014 paper - see our readme, we provide the link to the paper.
Long story short, I do think Dragonfly is the fastest in-memory database in terms of throughput and latency today. We will see if we manage to stay this way when we extend our capabilities with SSD tiering.
They don't fsync the wal on every write, it's probably done as group commit every x seconds or every x MB.
I think came to conclusion that it could be interesting as an independent (novel) store but not something that can implement Redis with its complicated multi-model API, transactions and blocking commands. I do not remember all the details though...
Basically, I worked in a cloud company in a team that provided a managed service for Redis and Memcached. I witnessed lots of problems that our customers experienced due to scale problems of Redis. I knew that these problems are solveable but only if the whole system would be redesigned from scratch. At some point I decided to challenge the status quo, so I left the company and..and here we are.
One thing in common - we both thought that cache-based heuristics can be largely improved compared to memcached/redis implementations. We did it differently though. I think our cache design has academic novelty - I will write a separate post about it.
Can we see such disparity in benchmark even if we run Ncore instances of redis in parallel?
For SSD based storage, it’s getting 50k reads/sec PER core and scales linearly with # of cores you have in your cluster. (They achieved 8MM reads/sec with 384 cores)
For example, INCR would require one read followed by one write of the new value, and of course this will result in very inefficient mutation range conflicts (which must be retried for another couple of round trips) if you have frequent updates of the same keys in multiple concurrent transactions.
https://apple.github.io/foundationdb/api-python.html#api-pyt...
That said, I still don't think that it is necessarily the perfect match for implementing some of the Redis data structures.
Redis is basically a very performant, single-threaded (mostly) single-node in-memory datastructures system with an efficient and readable server protocol strapped to it.
FoundationDB is a completely different beast that has like 6+ distinct roles, and is optimized almost exclusively for interactive serializable transactions, range reads, and correctness.
They’re just completely different things, I recommend reading the FoundationDB paper to get a sense for its architecture. The amount of “steps” involved in processing an FDB write is much higher than in Redis.
So if I have 1 machine and increase from 2 to 256 core the throughout will scale linearly without the SSD ever being a bottleneck?
And we have more plans for using io_uring in DF in the future.
Could you get me a one liner on the helio library is it used as a fiber wrapper around the io_uring facility in the kernel? Can it be used as a standalone library for implementing fibers in application code?
Also it seems that spinlock has become a defacto standard in the DB world today, thanks for not falling into the trap (because 90% of the users of any DB do not need spinlocks).
Another curious question would be - why not implement with seastar (since you're not speaking to disk often enough)?
Re helio: You will find examples folder inside the projects with sample backends: echo_server and pingpong_server. Both are similar but the latter speaks RESP. I also implemented a toy midi-redis project https://github.com/romange/midi-redis which is also based on helio.
In fact dragonfly evolved from it. Another interesting moment about Seastarr - I decided to adopt io_uring as my only polling API and Seastar did not use io_uring at that time.
1. I speak fluently C++ and learning Rust would take me years. 2. Foodchain of libraries that I am intimately fimiliar with in C++ and I am not familiar with in Rust. Take Rust Tokyo, for example. This is the de facto standard for how to build I/O backends. However if you benchmark Tokyo's min-redis with memtier_benchmark you will see it has much lower throughput than helio and much higher latency. (At least this is what I observed a year ago). Tokyo is a combination of myriad design decisions that authors of the framework had to do to serve tha mainstream of use-cases. helio is opinionated. DF is opinionated. Shared-nothing architecture is not for everyone. But if you master it - it's invincible.
EDIT: another quick note: copy-on-write implementations on the user space, algorithmically, are cool in certain situations, but it must be checked what happens in the worst case. Because the good thing of kernel copy-on-write is that, it is what it is, but is easy to predict. Imagine an instance composed of just very large sorted sets: snapshotting starts, but there are a lot of writes, and all the sorted sets end being duplicated in the process. When instead the sorted sets are able to remember their version because the data structure itself is versioned, you get two things: 1. more memory usage, 2. a lot more complexity in the implementation. I don't know what dragonflydb is using as algorithmic copy-on-write, but I would make sure to understand what the failure modes are with those algorithm, because it's a bit a matter of physics: if you want to capture a snapshot at a given Time T0 of a database, somehow changes must be accumulated. Either at page level or at some other level.
EDIT 2: fun fact, I didn't comment something about Redis for two years!
Please allow the possibility that Redis can be improved and should be improved. Otherwise other systems will eventually take its market apart.
I appreciate your comments very much. I've wrote about you in my blog. I am an engineer and I disagree with some of the design decisions that were made in Redis and I decided to do something about it :) to your points:
1. DF provides full compatibility with single node Redis while running on all cores, compared to Redis cluster that can not provide multi-key operations across slots.
2. Much stronger point - we provide much simpler system since you do not need to manage k processes, you do not need to *provision* k capacities that managed independently within each process and you do not need to monitor those processes, load/save k snapshots etc. Our snapshotting is point in time on all cores.
3. Due to pooling of resources DF is more cost efficient. It's more versatile. We have a design partner that could reduce its costs by factor of 3 just because he could use x2gd machine with extra high memory configuration.
Regarding your note about memcached - while we provide similar performance like memcached our product proposition is anything unlike memcached and it's more similar to Redis. Having said that - I will add comparison to memcached. I do believe that memcached as performant as DF because essentially it's just an epoll loop over multiple threads.
Re you comment about snapshotting. We also push the data into serialization sink upon write, hence we do not need to aggregate changes until the snapshot completes. The complex part is to ensure that no key is written twice and that we ensure with versioning. I do agree that there can be extreme cases where we need to duplicate memory usage for some entries but it's only for the entries at flight - those that are being processed for serialization.
Update: re versioning and memory efficiency. We use DashTable that is more memory efficient that Redis-Dict. In addition, DashTable has a concept of bucket that is comprised of multiple slots (14 in our implementation). We maintain a single 64bit version per bucket and we serialize all the entries in the bucket at once. Naturally, it reduces the overhead of keeping versions. Overall, for small value workloads we are 30-40% more efficient in memory than Redis.
The complexity here can be seen in two ways: complexity of deploying more Redis instances, or complexity of the single instance. It's a trade off. But I think that Redis may go fully threaded soon or later, and perhaps your project may accelerate the process (I'm no longer involved, just speculating).
1. Your point about Cluster, I addressed it many times: the point is, soon or later even with multi-threading you are going to shard among N machines. So I believe that to have this problem ASAP is better and more "linear".
2. Already addressed in "1" and my premise.
3. Yep there are advantages in certain use cases related to cloud costs and so forth, that's why maybe Redis will end fully threaded as well.
About memory efficiency, what I meant is that to have versioned data structures, that is an approach to do user-space copy on write even in the case of multiple changes to large single keys (big sorted set example), you need more memory likely, to augment the data structure. Otherwise the trick is to copy the whole value, that has other issues. It's a tradeoff.
In a world where cloud providers offer instances with terabytes of memory and 128 vCPUS (e.g. aws x2iedn.32xlarge family maxes out at 4TB, gcp m2 family maxes out at 12TB) is that really inevitable? Applications serving 10s of millions of users likely won't come anywhere close to that limitation.
But the reality is that most companies and most use-cases do not need terrabytes of data. I would say that today the comfort zone for Dragonfly is upto 512GB per instance (1). So dragonfly solves the issue for... I would say 99% percent of the use-cases. Only the last percentile would need horizontal scale, and probably their business is already big enough, so that they can affort a high-quality eng team to work with horizontal clusters.
(1) We need to improve some things (mainly around serialization format of rdb) to reach another magnitude of 4TB. Nobody wants to wait for days to load a 4TB snapshot.
Glad to see you guys made a lot of progress, although a little disappointing you chose to go down the path of building yet another source available DB and not contributing to open source.
If chrome was not born you would still use microsoft explorer with aspx sites.
The conversation happened here, https://github.com/redis/redis/issues/8340, and it's not like the most pressing issue for the project. It's also not as complex as what was implemented for dragonfly, which basically has native support from the ground up for concurrent programming during command execution. It would be hard to do in C as well.
I am not going to get into what's better and what's not, especially because I haven't released the v0.1 yet and therefore it's not usable, but I am working on cachegrand which is aims to be (also) a redis compatible platform.
I have done A LOT of research and development before picking up the current architecture (you can see it from the amounts of commits) and I am trying to test as much as possible (Almost 1000 unit tests so far, but there is still plenty to do).
if you look at the repository please bare in mind that: - there is no v0.1, the code available in the repo only supports the basic GET, SET and DELETE (apart from a few additional commands like HELLO, QUIT, PING)
- the code in main currently supports only storing the data on the disk, which is also why the tests are failing, I am doing some general refactoring and need to bring back the in-memory storage (issue n. 88)
- there are some general performance metrics available on the repo
- don't enable verbose logging, it's currently synchronous :) - cachegrand is able to fully run in single thread mode so I can actually compare it to redis (well when it will make sense)
- only linux, requires a kernel 5.8 at least (e.g. it's provided by ubuntu 20.04.2 lts, but I didn't really care too much as it will take quite a bit more before I get the first stable version and by that point the kernel requirement will not be an issue anymore)
What I can say is that the project really focus ONLY on performances, therefore is not as memory saavy as redis or similar platforms, and it actually aims more to compete with Redis Enterprise long term than just Redis, on the other end it implements a number of things from the ground to boost massively the performances:
- cachegrand architecture follows almost the share nothing principle with the only exception of the hashtable because it has been built around that need
- I implemented from ground an hashtable capable to deliver lock-free and wait-free GET operations and which uses localized spinlocks for the SET and GET operations, basically the contention is spread across the hashtable instead of being bound to X queues
- the hashtable also support SIMD operations (AVX, AVX2 and AVX512F), it's heavily optmized to reduce the memory accesses is able to embed short strings in the bucket to further reduce memory accesses
- cachegrand will support both memory and ad-hoc backend for the storage that is going to be basically a time-series database (cachegrand is not bound to redis functionalities, the redis command set is just a way to expose these for now)
- I implemented from scratch a fiber library able to do a context switch in just 7ns
- the network and storage backend are modular, currently it really only support io_uring but the goal is to also add XDP+the FreeBSD network stack support for the network (e.g. similar to what has been done with F-Stack and DPDK) and then io_uring with the NVME passthrough for the storage (not sure if I will also add support for SPDK)
- I have also implement an ad hoc memory allocator which waste some memory but it's able to do memory allocations and free in O(1) (here a nice chart https://www.linkedin.com/posts/danielesalvatorealbano_dublin...)
- most of the code is built aiming to be zero-copy (there are a few places where it happens right now as I need to fix a couple of things)
Just to underline it, currently it's not possible to play with it, until I merge the branch I am working on, because performances would be terrible (only on-disk storage and currently without caching), the tests are broken for the same reason.
- If it's a normal configuration partitioning a single large node with multiple instances using Redis cluster
- A cost equivalent cluster of machines with a similar memory size running on Redis cluster
There was an optimization built for Redis 7 where we actually start return copied memory pages back to the kernel, https://github.com/redis/redis/pull/8974, I wonder if the testing provided on Dragonfly includes this optimization.
Will definitely follow this to see how it develops. Good luck.
A lot of projects say "faster" without giving some hint of the things they did to achieve this. "A novel fork-less snapshotting algorithm", "each thread would manage its own slice of dictionary data", and "core hashtable structure" are all important information that other projects often leave out.
I’ve seen the VLL paper before and I’ve wondered how well it would work in practice (and for what use cases). Does anyone know how they handle blocked transactions across threads? Is the locking done per-thread? If so, how do you detect/resolve deadlocks?
It also be good to see a benchmark comparing single-thread performance between DragonflyDB and Redis. How much of the performance increase is due to being able of using all threads? And how does it handle contention? In Redis it’s easy to reason about because everything is done sequentially. How does DragonflyDB handle cases where (1) 95% of the traffic is GET/SET a single key or (2) 90% of the traffic involves all shards (multi-key transaction)?
Having said that, DF also has a novel caching algorithm that should provide better hit rate with less memory consumption.
Get/set operations look like they don't need it.
This is our initial release and we just did not have resources to showcase everything under different scenarios. Having said that, if you open an issue with a suggestion of a benchmark that you would like to see I will try to run soon...
Anyway, I wrote lots of unit tests to cover those atomicity issues. You can also write a custom python/nodejs/golang/... scripts that simultenusly write and read from the same multiple keys in such way that some invariant is preserved. For example, "mset x {i}, y {i}" for random `i` and in parallel do "mget x y" and to check that the response returns same values. You can also test this for other families using transactions like "MULTI; lpush x ${foo}; lpush y ${foo}; EXEC" .. and then similarly test that x and y have exactly the same lists.
I wonder if there could be a tool that instruments in the same way that afl does to try to detect races or inconsistent states.
People not read docs neither know the consequences of words like "eventual" or "in memory" and star using this kind of software as primary data stores, instead of caches/ephemeral ones...
Conversely, plenty of DBs with programmable transactions (e.g. SQL) are considered work-a-day "ACID" enough, despite some massive gaps in their transactional model (no DDL in transactions, no nested transactions, atomic only when below a certain size, etc.)
For _any_ database there will be important information only available in the documentation.
I think that covers almost all the whole dev population, for what I see in relation with RDBMS. Lucky us most RDBMs shield the mistakes in their usage, a lot.
That is why I see is "dangerous" to call ephemeral/eventual stores as "db". Marketing/positioning have impacts...
Having said that we carefully choose to write everywhere in the docs thay we are in-memory store (and not the database).
Btw, I reserve full rights to provide full durability guarantees for DF and to claim the database title in the future.
Might try this out.
I am spoiled.
Elastic made change after it was very popular, MariaDB is same story, and even more so only uses BSL for "Enterprise" components which have very little community adoption.
We see however other folks, such as Neon picking permissive license for their technology https://github.com/neondatabase/neon
I think for Open Source Project just starting up concern of "Clouds will steal my lunch" is just stupid. If you're worth for clouds to Adopt you're in 0.1% of all Open Source Projects and already "winning" You can WHEN revisit your license, think how to get to that point, rather than create adoption barriers early on
If you run let say 32 instances of Redis ( not using HT ) with CPU pining will be much faster than DF assuming the data is sharding/clustered.
On r6g it's 1.4M qps and then it's saturated on interrupts due to ping pong nature of the protocol. This is why pipelining mode can reach several times higher throughput - your messages are big. c6gn is network-enhanced instance with 32 network queues! it's the most capable instance in AWS network-wise. This is why DF can reach there > 3.8M qps.
Youth of product makes it bit scary to use fully in mission critical systems - given how many problems with Redis started to show up under proper load. But definitely on my watch list.
Redis is fast enough. Read/write speed isn't usually the bottleneck, it's limiting your data set to RAM. I've long ago switched to a disk-backed Redis clone (called SSDB) that solved all my scaling problems.
The docs make some of the differences clear. Worth reading the GitHub repo readme.
I just deployed it on Northflank with your public docker image and wrote a guide here: https://northflank.com/guides/deploy-dragonfly-on-northflank... - works great!
Homepage: https://dragonflydb.io/
Benchmark: https://raw.githubusercontent.com/dragonflydb/dragonfly/main...
you can dm me at roman at dragonflydb.io
Could this architecture be extended to scale across multiple machines? What would be the benefits and costs of this?
For intra-process framework it's not an issue (as long as we do not have deadlocks).
I see only throughput benchmarks. Redis is single threaded, beating it at latency would have been far more impressive.
Do you have latency benchmarks at peak throughput?
Now, please take into account that DF maintains 99th percentile of 1ms at 3M! qps and not at 200K.
https://github.com/dragonflydb/dragonfly/blob/main/LICENSE.m...
I understand wanting to protect your work from someone else turning into a service, but I will need to get our org's legal team to review it first.
In Israel, nowdays, the price of a watermelon in a local supermarket is 1.8$.
But if you go to farmers in the north, you can probably buy it for 30-50 cents. But then you would spend 3 hours in traffic and 20$ on gas.
So, Dragonfly is a local-supermarket that sells watermelons for 50 cents. mic drop.
Why do you need a redis/memcache? Because you want to look up a shit-ton of random data quickly.
Why does it have to be random? Do you really need to look up any and all data? Is there not another more standard (and not dependent on a single db cluster) data storage and retrieval method you could use?
If you have a bunch of nodes with high memory just to store and retrieve data, and you have a bunch of applications with a tiny amount of memory.... Why not just split the difference? Deploy your apps to nodes with high amounts of memory, add parallel processing so they scale efficiently, store the data in memory closer to the applications, process in queues to prevent swamping the system and more reliable scaling. Or use an SSD array and skip storing it in memory, let the kernel VM take care of it.
If you're trying to "share" this memory between a bunch of different applications, consider if a microservice architecture would be better, or a traditional RDBMS with more efficient database design. (And fwiw, organic networks (as in biological) do not share one big pot of global state, they keep state local and pass messages through a distributed network)
https://laracasts.com/series/russian-doll-caching-in-laravel (subscription only)
Unfortunately, it looks like they retired their database caching lesson, which was a mistake on their part since it was so good:
https://laracasts.com/discuss/channels/laravel/cache-tutoria... (links to defunct https://laracasts.com/lessons/caching-essentials)
Laracasts are the best tutorials I've ever seen, regardless of language, outside of how php.net/<keyword> search and commenting was structured 20 years ago. They would be the best if they got outside funding to provide all lessons for free.
Anyway, one HTTP request can easily generate 100+ SQL queries under an ORM. Which sounds bad, but is trivially fixable globally without side effects via global "select" query listeners and memoization. I've applied it and seen 7 second responses drop to 100 ms or less simply by associating query strings and responses with Redis via the Cache::remember() function.
I also feel that there's a deeper problem in how web development began as hacking but devolved into application-heavy bike shedding. We have a generation of programmers taught to apply decorators by hand repeatedly, rather than take an aspect-oriented (glorified monkey patching) approach that fixes recurring problems at the framework or language level. I feel that this is mainly due to the loss of macros without runtimes providing alternative ways of overriding methods or even doing basic reflection.
So code today often isn't future-proof (requires indefinite custodianship of conventions) and has side-effect-inducing or conceptually-incorrect changes made in the name of secondary concerns like performance. The antidote to this is to never save state in classes, but instead pipe data through side-effect-free class methods, then profile the code and apply momoization to the 20-% of methods that cost 80+% of execution time.
Also microservices are definitely NOT the way to go for rapid application development or determinism. Or I should say, it's unwise to adopt microservices until there are industry-standard approaches for reverting back to monoliths.
Unfortunately, reality.
1. I speak fluently C++ and learning Rust would take me years. 2. Foodchain of libraries that I am intimately fimiliar with in C++ and I am not familiar with in Rust. Take Rust Tokyo, for example. This is the de facto the standard for how to build I/O backends. However if you benchmark Tokyo's min-redis with memtier_benchmark you will see it has much lower throughput than helio and much higher latency. (At least this is what I observed a year ago). Tokyo is a combination of myriad design decisions that authors of the framework had to do to serve the mainstream of use-cases. helio is opinionated. DF is opinionated. Shared-nothing architecture is not for everyone. But if you master it - it's invincible. (and yeah - there is zero chance I could write something like helio in Rust)...