Garnet – A new remote cache-store from Microsoft Research
github.com
github.com
You can trade high-level expressiveness for low-level control where needed and you don't have to deal with any FFI to do it.
Nowadays, I think it's probably the most natural language and platform for teams that need to move on from TypeScript rather than Go or Rust given the similar constructs and idioms.
> the library ecosystem
I don't really see many gaps. There tends to be fewer libraries, but the libraries available generally feel more complete and well thought out because users tend to cluster around the known libraries.Many of the first party libraries are really, really good. EF Core is a prime example of possibly one of the best ORMs on the market right now in terms of productivity, ergonomics, and performance.
isn't that the "works on my machine" of this discussion? What's the Apache Tika for .net then? I don't mean, pdf parsing, I don't mean .docx parsing, I mean a framework for interacting with all their supported types <https://tika.apache.org/2.9.1/formats.html> with one surface area?
You're right, how silly of me, I'll install some rando's 6 year old build of it right away. But in seriousness, that readme did remind me of the thing I was thinking of: http://www.ikvm.net/userguide/ikvmc.html
Personally, I'd say it is probably the case that there are better alternatives. Azure Document Intelligence being one of those. Sure, it can't handle the variety of formats, but again, we have to come back to the underlying assumptions: .docx, .pdf, and a few handful of formats probably covers 99% of business use cases. Everything else would be considered niche.
There are exceptions in the form of well-written community-maintained libraries, but these exist for minority of Apache projects and vary a lot between languages.
Luckily, there are often (but not always) plenty of alternatives to whatever it is that you seek from Apache.
There's no need for VS for most (any?) .NET workloads these days. That's why I wrote it was an early mis-step. Nowadays, it's easy to do .NET dev on any platform. In fact, we ship our production runtime to AWS t4g (Arm64) instances.
It can be done, but it’s so inferior to VS that you might as well use VS.
1. Memory management with a GC often has higher throughput. With the downside that you can have high latency when a garbage collection occurs.
2. The JIT compilation can potentially do a better job of optimizing, because it has information about how the code has been run so far.
It is possible for c or c++ (or rust) to get get the first using alternative memory management strategies, and the second by hand optimizing, and/or using profile guided optimization.
Which is why WASM isn't that great novelty for grey beards.
Also limits on number of people in given time who can write optimal C code vs C# will also make managed code better.
These are then subsequently wrapped by `Span<T>` and `ReadOnlySpan<T>` respectively. This way a span can be a slice of memory that can have any origin:
var fromStack1 = (stackalloc byte[32]);
var fromStack2 = (Span<byte>)[0x20, 0x20, 0x20, 0x20];
var fromHeap = new byte[32].AsSpan();
unsafe
{
var ptr = NativeMemory.Alloc(32);
var fromMalloc = new Span<byte>(ptr, 32);
// Don't forget to free :)
}Ignoring the type systems from Eiffel, Objective-C, Oberon and Modula-3 linage, even though they were inspirations for Java, has been shown to have been a bad decision in hindsight.
Now, I could be talking out of ass, haven't done enough C/C++ coding in over a decade.
And the intel C++ compiler has provided significant performance for decades via optimization without needing to optimize by hand.
The inlining of virtual calls is the critical optimization that enables this. Because C/C++ is optimized statically and never at runtime, it is unable to optimize results of function pointer lookups (in C, and thus also virtual calls in C++). However, the JITs can inline through function pointer lookups.
In sufficiently complex programs, where polymorphism is used (i.e. old code that calls new code without knowing about it), this yields an unsurpassed speed advantage. Polymorphism is critical to managing complexity as an application evolves (even the Linux kernel, written in C, uses polymorphism, e.g. see struct file_operations).
People only have to learn how to use them, instead of placing all garbage collected languages into the same basket.
In this specific case, MSIL and .NET were designed to support C++ as well, and languages like C# and F# do have ways to access those features, and even if some feature isn't exposed at the language grammar level, you can emit the same MSIL that C++/CLI would generate.
https://ayende.com/blog/197412-B/high-performance-net-buildi...
I've seen Go code best Rust one, really when comparing languages you should look at the same implementation, if the design is completely different its not really comparable.
Modern garbage collectors are very good and handle many common use cases rather efficiently, with minimal or zero pauses, such as freeing short-lived objects. Many GC operations are done in parallel threads without stopping the application (you just pay the CPU overhead cost).
Also JITs in both the CLR and the JVM perform optimizations such as escape analysis, which stack-allocate objects that never escape a function’s scope. These objects thus do not have to be GC’d.
So really with a GC’d language, you mostly have to worry about pauses and GC CPU overhead. Most GCs can be tuned for a predictable workload. (A bigger challenge for a GC is a variable workload.)
Appreciate the detail about the stack allocated bits in .NET.
Unfortunately, object stack allocation was not one of them even though DOTNET_JitObjectStackAllocation configuration knob exists today, enabling it makes zero impact as it almost never kicks in. By the end of the experiment[0], it was concluded that before investing effort in this kind of feature becomes profitable given how a lot of C# code is written, there are many other lower hanging fruits.
To contrast this, in continuation to green threads experiment, a runtime handled tasks experiment[1] which moves async state machine handling from IL emitted by Roslyn to special-cased methods and then handling purely in runtime code has been a massive success and is now being worked on to be integrated in one of the future version of .NET (hopefully 10?)
[0] https://github.com/dotnet/runtime/issues/11192
[1] https://github.com/dotnet/runtimelab/blob/feature/async2-exp...
Redis is single-threaded, it’s simple and effective. I’m not sure it needs optimization, and we have 3 alternatives here.
Garnet however is the first alternative to actually outperform Redis at both low and high levels of concurrency, which is remarkable. I can’t wait to try it out.
1. Clients query a list of servers (IPs) and handle failover when you don't quickly get a reply.
2. Most of those servers at the root and TLD level are actually anycasted from multiple locations globally, so you connect to the closest instance.
3. Those instances are often clusters of physical servers. The big ones have fully redundant networking, so any router or switch failing doesn't take it down. Some run different DNS server software on each physical server, so even software bugs won't take down the whole system.
And then the problem is confounded by the fact that, ironically, DNS works so well that we don’t think of it as a primary point of failure. So inevitably, when there is a DNS problem, it’s the last thing we check. This just reinforces the idea that the problem is always DNS… because in those long, hard to troubleshoot instances… the problem was DNS.
All variants of ‘DNS is working perfectly, just your expectations of how it will work in your situation are not completely correct’.
// wrote what might have been first commercial-use dynamic DNS server for a regional ISP in early 90s, invented an unreasonably effective geo+latency balanced anycast-like DNS for global video delivery network in 00s
Except when Windows is the architecture.
This may be a place where Garnet is a good alternative.
DNS can fail when someone who is managing it have no idea what the hell they are doing and abuse it to do things it wasn't meant to do.
https://www.microsoft.com/en-us/research/blog/introducing-ga...
Seems they have tons of internal need and are willing to share.
For reference, something commonly used with IIS for ASP.NET apps was to have an out-of-process "session state" store, so that if the web app process restarted, users wouldn't lose their sessions and have to log in from scratch. Sure, you can put this somewhere central like SQL Server, but then every web page request sits there waiting for the session state to load before it does any processing at all. Session state is also typically locked in some way, which has all sorts of performance issues.
The typical current solution is to use Redis for both caching and session state, and this works... okay-ish. Throughput is high, sure, but Redis is a separate resource in Azure and is stupidly expensive. I really don't want to pay Oracle DB prices for something this simple. It's also a bit of a hassle to wire up.
In this article they talk about 300 microsecond response times, but that's irrelevant in any zone-redundant design because all Azure load balancers use random zone selection. So you'll have a web server picked in a random zone, then it'll contact a cache server in a random zone in turn. That server in turn may not have your key and have to contact yet another random zone to get your cache data! Your traffic ping-pongs between data centres. This introduces about 1-3ms of delays, up to 10x higher than the advertised numbers for Garnet.
The ideal scenario would be something like what Microsoft Service Fabric does: it has a "reliable collections"[1] service that runs locally on each host node and replicates to two other nodes. A web app can always read its cached values from the same physical host. The latency can be single-digit microseconds in some cases, which is thousands of times faster than any naively load balanced external service, no matter how well optimised.
I don't want 30% faster than Redis. I want 3,000x faster.
[1] https://learn.microsoft.com/en-us/azure/service-fabric/servi...
I am having trouble understanding this. Why wouldn't they wrap that in a transaction internally for you, and make the command atomic? What other atomicity "gotchas" are there.
It's also quite intriguing for me to see it's written in C#, as that's my native tongue. I'd be keen on dedicating some time to delve into the code.
"After thousands of unit tests and a couple of years working with first-party teams at Microsoft deploying Garnet in production (more on this in future blog posts!), we felt it was time to release it publicly" https://microsoft.github.io/garnet/blog
Thanks for your continuing work on memcached! I'd be very curious how garnet's benchmarks compare with memcached.
Would be interesting to know why it was forked, why the changes can't be incorporated and wether FASTER continues to be developed
Is it possible to use Aspire locally or is it just a cloud only 'framework'?
I trust Redis to not do something weird with their licensing or pricing in the future.
Plus Redis has billions of production hours under its belt.
It's easier to install and understand.
Also, this is Microsoft Research. This thing is code sharing not a product.
The real interesting part is what Azure will do with it.
> Garnet has been of sufficiently high quality that several first-party and platform teams at Microsoft have deployed versions of Garnet internally for many months now.
Aged like milk.
Of course, as a research project it doesn't have the same stability / support, etc. as redis, but I could easily imagine this rolling into a real product if it's popular.
...and as an MIT license, if nothing else, the code is a fun read. :)
[0] - https://github.com/microsoftarchive/redis?tab=readme-ov-file...
It's bizzare. What more rights could they possibly want on top of MIT to warrant CLA?