Replacing Redis with "something faster" is a bit like removing the doors on a car because "lighter means faster!". It might look good on a racetrack, but it's about as pragmatic as climbing through a window every morning before setting off for work.
Replacing Redis with "something faster" is a bit like removing the doors on a car because "lighter means faster!". It might look good on a racetrack, but it's about as pragmatic as climbing through a window every morning before setting off for work.
I also find it interesting that the BSD license enables this 3rd party company to fork Redis and build closed source commercial software on top of it. One of the trade offs to consider when licensing a project.
Thanks!
As user of commercial software I am fine with it, not so sure if FOSS advocates at large will be so happy when only non-copyleft licenses survive and we are back in the shareware/pd libraries days.
Why do we assume closed-source software vendors contribute nothing back?
Speaking as an employee at a company that produces a closed-source software product that uses open-source libraries, I've contributed plenty back to various libraries, including publishing some of my own.
Libraries with permissive licenses get more users, and more users mean more opportunities for receiving contributions. Something like GPLv2 is really only truly effective at soliciting contributions that it wouldn't have received otherwise if there's no viable alternative.
Or to give another example, Rust is dual-licensed under the Apache License, Version 2.0 and MIT. This permissive license made it really easy for lots of people (including myself) to contribute to it. If it were released instead using the GPL, it would likely be a shadow of the language it is today, if even still alive at all.
when you're contributing to GPL software, you don't need your employer's permission for it to be upstreamed. If they release the code, and don't violate the GPL, then there's nothing the employer can do to stop it from going upstream.
Yes you do. Except the permission in this case is "permission to use the library" in the first place, which is a much higher bar than permission to upstream changes to a permissively-licensed library because using a GPL library has much farther-reaching implications than merely contributing changes back.
Yes, you do; if you are contributing to it, you need to have exactly that permission. If you are working on a derivative, you need that permission before you can contribute it to anyone else, including upstream.
> If they release the code
Plenty of people work on internal code for their employers, so this is not a given.
Like companies staying away from Linux or GCC? It's debatable if these projects would have been as successful using MIT/BSD.
Companies having problems with GPL is a problem of companies and not a problem of the license. If the library you want is GPL, then why blame the project and not your company's legal department?
For the curious: https://engineering.fb.com/data-infrastructure/scribe/
Edit: HN thread: https://news.ycombinator.com/item?id=21181982
Basically, everything that needs logging and post-processing by both real-time systems (e.g. Puma) and batch processing (e.g. all of the data that's ingested and sent to the data warehouse) goes through Scribe.
(disclaimer: I work in Scribe)
Scuba: https://research.fb.com/publications/scuba-diving-into-data-...
Puma: https://research.fb.com/publications/realtime-data-processin...
But it's not like that's through one pipe. We don't talk about how many zillions of tons of steel per minute are moved on freeways.
We picked the job queue framework long before we started moving tens of thousands of jobs a second through it with a later feature, and probably exceeded its design constraints - in that instance we effectively were using Redis + jobs as a pauseable and throttleable write buffer for MySQL.
With major tooling changes we likely could have come up with something a lot more elegant, I'm not longer with that company, but our plan was always to get rid of the need for that system to hit MySQL at all, which would eliminate a lot of our need to control throughput with our Redis buffer. Of course, doing that would have taken a lot more time and effort, and in our case doing "the simplest possible thing that would work reliably and serve our customers" meant that we probably could have used a faster Redis instead of multiplexing requests and workers across multiple Redis instances on the same hardware.
Sometimes an "improved" version of a tool you're already using can be really useful, if you find yourself in a bind and the alternative is major architectural changes.
Also, you didn't mention that, but I've seen this happening often, you should never use redis as a primary data store if your data is important.
I guess I don't know how redis deals with the network, but I'd assume that it handles concurrent requests. If not, then I guess that's the case where you'd see less than 1/n cpu utilization and it could still be CPU bound.
Because redis handles all requests serially you need to be very careful if the redis cluster is shared between different apps. You want to share the cluster among apps with similar usage patterns.
Basically redis can't necessarily guarantee data safety - which is perfectly fine for it's normal use case as a caching layer and data you don't necessarily care if you lose (for example rate limiting relies on tracking the number of requests/s within say an hour - if you lose the data you don't really care). User sessions too, worst case people have to log in again, meh whatever.
It would be great to remove limitations like RAM-only capacity in exchange for a slight performance hit, and while also gaining better core utilization. We used ScyllaDB (a very fast cassandra clone) in the past for the cpu/disk scalability but always felt Redis offered better APIs. Now it's a real option.
What are the advantages of "multi-threading" other than for performance?
Multithreading IO also reduces the latency hit from disk-persistence and provides more concurrency and throughput, which is a great tradeoff when you don't really need sub-millisecond performance but want the same API across a larger dataset limited by disk instead of RAM space.
That means nothing at all, if performance is not your concern.
> Taking advantage of all your cores.
Uh, is there any reason you'd want to "take advantage of all your cores" other than performance?
The question/proposal/dialog we started from, don't forget, was:
1. Redis is already as performant as I want, it's not even close to being a bottleneck, so I have no need for multi-threading, and this is a _very very common_ case, as redis is performant enough for a lot.
2. There are other advantages of multi-threading than performance. (Ie, other reasons you'd want it despite 1).
You seem to just be going around in circles. If 1, why would you care about "taking advantage of all your cores"?
However, “should” you in a world where you only have cloudwatch because you don’t want to roll your own monitoring or bring in a third party vendor? Probably not. It’s not a good thing, but it’s a reality. 1/n approximation is your go to here.
However yeah, 99% of all deployments could get by with SQLite/MySQL/Postgres performance amounts.