A curated list of .NET performance resources
github.com
github.com
In Java, every object is heap allocated - unless you are lucky enough for that to be optimized away. If you absolutely need the performance, you must resort to a weird kind of place-orientated programming, where you pass around mutable references to primitive types that get operated on by static functions. It kind of resembles C, but without the struct keyword!
https://openjdk.java.net/jeps/404
https://openjdk.java.net/jeps/402
Both .NET and Java did major mistakes at their v 1.0, by ignoring what other GC based languages were doing.
- Proper value types (Even .NET is only catching up since C# 7.0)
- AOT compilation alongside JIT (NGEN was just for startup and without optimizations, and on Java side only third parties)
- Seamless FFI with host OS (P/Invoke did better, JNI is finally getting a replacement)
Eiffel, CLU, Mesa/Cedar, Oberon linage, Modula-2+, Modula-3 were inspiration, but those capabilities were apparently considered not relevant enough, go figure.
Erlang's BEAM's nif[0] interface is the best I've seen in mainstream runtimes for FFI. You write an Erlang/Elixir/etc. function, and loading a native library may swap out the function body. This allows for workarounds in Erlang/Elixir/etc. in case of version skew in your native libraries.
In my day job, I work with a domain-specific language with a similar FFI (actually, 3 different FFIs, one similar to nifs), and it's very handy to have a fallback that does things with /proc/, calling external command line binaries, slower versions of FFI functions, etc. (This allows the native code release schedule to be decoupled from the domain-specific language code changes.)
The syntax isn't the same, but a Java-like syntax would be
// getpid is built in, but it could be done something like this
// look for XNativeGetPID() in mylib.so or mylib.dll depending on OS
@native_replace("XNativeGetPID", "mylib")
public static int getpid() {
// Throw an exception if you want to force
// native implemementation
return Integer.parse(Runtime.system( "/bin/grep pid /proc/self/status | /bin/sed 's/..." ));
}
[0] http://erlang.org/doc/man/erl_nif.htmlBy the way, it is now in preview.
Do you have any idea how easy it is to provide a Java fallback on platforms where a native implementation isn't available? (I find it's often possible to call out to a command line tool as a workaround for things I'm trying to do in native code.)
To read more about Panama, https://jdk.java.net/panama/
Check Foreign-Memory Access API and Foreign Linker API related entries.
Though I’m not sure I would want this behavior, or at most for really optional deps; ideally software should be “installed” in an ideal environment.
Native implementations with non-native fallbacks can be an important decoupling mechanism.
I used to be a developer for a Linux desktop application with a peak of something like 60 million unique users over a trailing 3 month window. We had Windows, OS X, RPM, and dpkg installers with native libraries, but people also ran our software on FreeBSD, Gentoo, etc.
It's handy to have gradual enhancement/graceful degradation in non-ideal installs where the native API's you'd like either don't exist or are buggy and shouldn't be used.
In my day job, we sometimes have a business need to iterate new features at a faster tempo than our native code release schedule, so it's nice to have a similar sort of fallback implementation.
Let's say you want to have a native-looking prompt window via wxWidgets, and fall back to Swing if wx isn't installed. Or, you want to use some particular optimized native BLAS library if it's installed, or fall back to doing the linear algebra in Java if it's not available.
In these cases, and many others, "detecting OS features" involves checking LD_LIBRARY_PATH (or PATH on Win32) for libraries, checking they're for the correct architecture, etc. The OS already provides dlopen() or its equivalent for this purpose. Properly detecting the OS and emulating the OS's algorithm for finding native libraries is just wasteful when the OS already exposes the functionality.
As an added bonus, an FFI that replaces a method body allows someone who doesn't have the C/C++/Rust/FORTRAN handy, or can't read it, to see roughly what the native code should be doing.
AOT compilation of a dynamic language like Java is hard to do without losing performance. You can AOT Java now with the same compiler used also for JITC (Graal) and you lose, I think, about 20% of the peak performance unless you do C++ style profile collection and double compilation. It's not clear that lack of AOT compilation held Java back any - although SubstrateVM/native-image are getting popular at the moment, that's taking place in the context of hype around "serverless functions". If you aren't using those (or writing small CLI apps, or doing other unusual things) the benefit is less clear.
With respect to proper value types, trying to do that in 1.0 would have just killed the projects. Biting off too much. Look at the complexity of Valhalla - only a small part of it comes from backwards compatibility. Most of it is trying to get the benefits of generics over value types via specialisation without the horrible consequences suffered by other languages. It's not about rectifying past errors but rather, doing something new.
Java is a static language.
In C# structs are value types, and they've been with .Net since the beginning IIRC. I assume maybe you're referring to System.ValueTuple? It was added in C# 7 and allowed tuples to be value types along with significantly better syntax for working with them.
There is so much more for value type programming, see sibling comment.
- proper memory slices
- ability to do stackalloc in safe code
- readonly structs
- generics for blittable types
- zero allocation pipelines
- using without IDispose for structs without implicit boxing
- return scope for zero copy struct parameters
- native function pointers
- static lambdas
I would agree that things like
> - using without IDispose for structs without implicit boxing
and
> - generics for blittable types
would make value types more "proper" in the C# language.
Just like C and C++ offer similar optimizations to avoid writing the tons of Assembly that was so common up to the mid-90's.
So yes, optimizations are definitly part of proper value types.
[0] https://docs.microsoft.com/en-us/dotnet/csharp/tutorials/exp...
See the bullet list on a sibling comment.
https://channel9.msdn.com/Shows/Going+Deep/Expert-to-Expert-...
Your 25 year old jars are supposed to keep working as much as possible.
And I don't even do Java development.
Not to mention android and gradle being broken and randomly failing to work between different JVM versions - depending on the gradle flavour of the day.
I don't really see JVM as a pillar of stability and backwards compatibility.
And modules was a single breaking change, hardly a regular thing.
The don’t upgrade thing on the server is a generally wanted stable environment, or do you upgrade libc in prod without testing it first? Also, it happens mostly in banks and other not-too technically up-to-date places that often times depend on bugs themselves. Trailing the latest OpenJDK is the best thing one can do.
Also, android is not/barely Java.
We literally had software in production that wouldn't work when you upgraded the JVM and their official documentation said it wouldn't work if you updated and you were unsupported (and IIRC downgrading was also a PITA) - this was in JVM 5 -> JVM 6 days which AFAIK was a big change (TBH I don't know back then I didn't do app development I did system integration).
Second, I didn’t say it is in every single case backward compatible, since that is impossible in a continuously evolving platform. But there are really few “user”-facing changes in the JVM. And so I still state that calling the JVM non-stable (for which you provided no evidence), or non-backward compatible is simply false. Please tell me a platform which is stable/backward compatible according to you to at least this degree.
(Also, we are at java 16!, java 5, 6 is so old that it is not particularly telling about anything)
What do I mean by that? Here's my idea. Make a Fluent parser. I have some few shallow structs that can be used interchangeably. E.g. you want to have following structure
Entry:
- Message
- Term
- Comment
How do I add polymorphism to my entries?Option 1) Go with interfaces. Yay. I've incurred the boxing penalty. I might as well as just write a class.
Option 2) Go with [FieldOffset]. Only do this if you are a masochist and you REEEEEEEEEEEEAAAAAAAALLY need that performance.
In Rust for instance, I could replace that stupid struct with an enumeration, which is impossible in C#. Hopefully in future there will be discriminated unions or some Rust enums equivalent.
If C# had the ability to define
List< struct A | B | C >
I would been singing a different tune.You either in for
List<IInterface>
or List<ABCFieldOffset>- arrays (which most collections use under the covers) will need to allocate `elements*size_of_struct`. Unless the structs are the same size, you'd have to do some level of thunking/offsetting.
- Even if they are the same size, when you're iterating the list, how would the runtime know whether it's `A`, `B`, or `C`? Structs don't have header information that can be used to infer the type, it's only whatever fields you've defined (with padding potentially). Part of 'boxing' is essentially thunking the struct by wrapping a methodtable for the struct into the object [0].
If you need something like polymorphism with structs, your best bet is to make the struct itself quasi-polymorphic (i.e. with enums) and/or start abusing generics in some fashion I haven't yet figured out.
[0] - https://mattwarren.org/2017/08/02/A-look-at-the-internals-of...
Sure, so let's make union or "struct enum" and do the same thing Rust does. Add a discriminator byte to the structure and use it to figure out type.
> If you need something like polymorphism with structs, your best bet is to make the struct itself quasi-polymorphic (i.e. with enums) and/or start abusing generics in some fashion I haven't yet figured out.
Well. I think there is the FieldOffset hack. Basically discriminated unions where programmer plays the role of compiler.
[StructLayout(LayoutKind.Explicit)]
struct DiscriminatedUnion
{
// Imagine we have enum like so
// enum
// {
// V1(byte a, byte b)
// V2(sbyte a, sbyte b)
// }
// The "FieldOffset" means that this Integer starts, an offset in bytes.
// sizeof(byte) = 1, sizeof(EnumType) = 1
[FieldOffset(0)] public EnumType TypeFlag;
[FieldOffset(1)] public byte V1a;
[FieldOffset(2)] public byte V1b;
[FieldOffset(1)] public sbyte V2a;
[FieldOffset(2)] public sbyte V2b;
}
https://sodocumentation.net/csharp/topic/5626/how-to-use-csh...Problem is, this is fragile. It might break on endianess.
I think that's what he is saying. That it's quite a leap from "structs are horrible in practice" to "structs are horrible if you need polymorphism in conjuction with structs" which is in my experience a pretty rare situation. If you approach structs from Rust or C++ you might expect to be able to form hierarchies of them, but approaching it from Java you are very happy to just be able to create a Point(X, Y) struct.
Even the rust type of enum structs with clever overlapping layouts I find pretty rare use for. The common use case is homogeneous data. If I have polymorphism, I'm likely going to incur some pretty expensive virtcalls etc anyway. At that point I just use classes.
In GCed languages the default seems to be that every object is referred to by pointer. Java took the "ugly" step of adding in primitive types so you could calculate 1 + 1 without causing dozens of cache misses. Struct types take this hack a step further. It generally makes the definition of the languages virtual machine a lot uglier, for example most jvm instructions operate on primitive types.
Structs are just that data stored in the same layout.
Classes are just a vtable pointer to some object on heap.
Not in C++.
An OO GC'd language "class" is a mechanism for run-time polymorphism. That's what it's for.
C++ has no "runtime" concerns of these kinds, so virtualization sits on top systems with compile-time semantics (ie., c++ classes).
If the GC language properly supported interior pointers (and first class pointers) for the GC there wouldn't be a need to force boxing in all cases. You would be able to make the decision at the point of instantiation of the object. It would make the GC more complex, but I think these days most GCs support interior pointers for other reason anyway, right?
Or am I missing something?
I feel like this is every wart in the C++ language!
As a C++-lander, I find it surprising that object allocation is tied to the object's type. IMO, the boxed/unboxed axis should be orthogonal to the type itself. It should be a property of the object instance, not of the object's type.
[<Struct>]
type Entry =
| Message
| Term
| Comment
And like in Rust, you have pattern matching on them: let computeFooForEntry entry =
match entry with
| Message -> 0
| Term -> 1
| Comment -> 2(Optimized) discriminated unions have a size that is the max size of the possible types, plus a flag to indicate the type. This is the best you can hope for without boxing.
Example of that approach: https://github.com/Const-me/Vrmac/blob/1.2/Vrmac/Draw/Main/I...
0: https://aspsecuritykit.net/docs/api-reference/aspsecuritykit...
Getting your application aligned with the NUMA model makes way more difference in performance than anything else.
Choose a better data format (binary format rather than json).
Compress data.
Remove repeated work (lift out of loops).
Get all the data from the DB up front.
Index.
Cache.
Parallelise.
cpusets (aligned with the NUMA model)
> Get all the data from the DB up front.
Between those two, I'd add:
Move computation to the data where appropriate. Let the data layer (particularly the database itself) do the heavy lifting.
So its a small amount of data, but large computation it might be tempted to do it on the application. But its a large amount of data, I/O maybe a problem in doing it on the app.
You're of course right, and it could be that DB being a compute bottleneck is the most common case - but I've seen or been involved in projects where the bottleneck was DB I/O and processing unnecessary queries, because the application was doing most of the calculations.
In this article, with some optimizations, the author was able to achieve on the order of ~70m operations per second using the .NET port. I think .NET might actually be faster than Java in some places here.
Isn't that a Java thing? So, this one is a port to .NET
Is .NET sandboxing effective yet? It's my belief that if I want to do something like this, that I'm way better off using V8 or some other sandbox to host your code. Or use Deno or something.
Or is it still not possible to run someone else's .NET code really safely?
Asking for a friend who wants to make an online game in the style of C-ROBOTS, but is afraid of running other people's code.
Personally, I would still consider it. For a game, the risks are not that large, it’s not a bank nor a nuclear silo.
You can use Mono.Cecil library to reflect assemblies without loading them as code or executing any parts of them. You gonna need to whitelist allowed imports (allow stuff like String/List/Dictionary but not much else), blacklist pointer types everywhere (function arguments, return values, and locals variables, CIL+metadata do have types of things), blacklist DllImportAttribute, and probably a few other things I forgot.
The performance should be awesome that way, however security-wise that’s not 100% failsafe. Also easy to DoS your server by consuming too much CPU or using all the memory. You might need workarounds to address these issues. Mono.Cecil can patch code making new assemblies, not just inspect/reflect.
Probably for these reasons MS discontinued app domains / code security, and recommends processes, containers or virtual machines instead. You can run another instance of .NET runtime inside these things, the startup/warmup time is not that bad, it's not a Java.
P.S. If you only allow to edit code in your editor i.e. don’t need support for visual studio + Microsoft’s compiler, you can make your own language where only safe things are expressible, and compile that one into CIL. Simplifies many things, you only need to worry about CPU time and RAM usage. You don’t need to emit DLLs nor mess with CIL directly, can compile scripts into delegates using System.Linq.Expressions or probably some third-party libraries.
also see https://dotnet.microsoft.com/apps/aspnet/web-apps/blazor
https://docs.microsoft.com/en-us/dotnet/framework/misc/code-...
As others in this thread have mentioned, CAS does not even exist in .NET Core and .NET 5.
dotnet rely a lot on the JIT to optimize code
So cases when you run your program once (middleware, short lived servers) the performance is abysmal
Never trust "ultra optimized & cheated" benchmarks where they run the same code 1_000_000_000 times
And the time/resources it takes to reach steady state is also something to keep in mind, warming up the JIT isn't free
Hence the reason why most people looked at Go as an alternative to both dotnet/Java/Scala
* You can AOT compile .NET, that has been true for a while but now it is a first class option
It's been a while i haven't used dotnet, but just take any available benchmark, and compare default results with
UnrollFactor = 1
IterationCount = 1
WarmupCount = 0
There are probably more options, but i forgot them
https://benchmarkdotnet.org/articles/configs/jobs.html#run
> You can AOT compile .NET, that has been true for a while but now it is a first class option
That's not true at all, if you refer to R2R, it's only partial, your code is still JIT'd at runtime, it only helps to speedup a little bit the cold startup time
AOT compilation? I'll believe it when they'll release it, until then, it's all speculation
Devil's in the details, but there -is- AOT compilation[0]. While it hasn't been released as an official product, it has been used for a few projects including a commercial game [1]. And yes, they're looking into the next steps to make it a 'released' thing.[2]
[0] - https://github.com/dotnet/corert/
[1] - https://github.com/dotnet/corert/issues/8233#issuecomment-65...
[2] - https://github.com/dotnet/runtimelab/tree/feature/NativeAOT
>This repo is for experimentation and exploring new ideas that may or may not make it into the main dotnet/runtime repo.
Talented people are working on CoreRT, sadly it doesn't seems to be the focus of the product manager, hence my concerns
.NET has had AoT compilation since v1. What, exactly, are you talking about?
Sure, rely on JIT for stuff like webapps and the like, but a CLI tool that is designed to run once would be a good candidate for AoT compilation, which, as I stated before, .NET supports.
Also you're dramatically overstating how mature AOT is for dotnet
https://github.com/dotnet/runtimelab/tree/feature/NativeAOT
>This repo is for experimentation and exploring new ideas that may or may not make it into the main dotnet/runtime repo.
[1]: https://docs.microsoft.com/en-us/dotnet/core/deploying/ready...
>R2R binaries improve startup performance by reducing the amount of work the just-in-time (JIT) compiler needs to do as your application loads.
and
>However, R2R binaries are larger because they contain both intermediate language (IL) code, which is still needed for some scenarios
and
>Ahead-of-time generated code is not as highly optimized as code produced by the JIT. To address this issue, tiered compilation will replace commonly used ReadyToRun methods with JIT-generated methods.
So if you're targeting .Net 1.0 this "PSA" is impactful, otherwise, it isn't really accurate. It could hypothetically apply if you created a new .Net process each execution (thus invalidating the GAC) but even then you could just add ngen to your build/deployment process.
Running the same code a lot is certainly not cheating, a lot of code where performance really matters is actually doing the same thing over and over.
.NET 5 added source generators, which are one part that is necessary to get to a point where ahead of time compilation would be feasible for typical .NET projects. So this is an aspect that Microsoft seems to be looking into now.
For people who have not been following .NET closely, there is a thing called Blazor that runs in the browser using web assembly. To get the download size down, a linker is used to remove classes and methods that are not being used. You can also use this when deploying other types of apps.
https://docs.microsoft.com/en-us/aspnet/core/blazor/host-and...
Also, who is that most people? Go is a tiny language compared to Java and C#.
Also, Java takes longer to reach steady state as it starts execution in an interpreter, then JIT compiles hot methods. .NET skips the interpreter and goes straight to JIT compilation.
https://docs.microsoft.com/en-us/dotnet/core/whats-new/dotne...
<RuntimeIdentifier>linux-x64</RuntimeIdentifier>
<PublishReadyToRun>true</PublishReadyToRun>
It will still Jit after 30 calls to refine further; but should start at a higher level of optimization.what C# lacks here?
is there huge gap between CLR and JVM?
Java's popularity depends how HF in HFT is and the threshold for pain developers are willing to go through to develop Java in a way to avoid GC and reduce memory use (and hence reduce cache misses)
Deploying to Linux boxen in colocation facilities is the norm and there is little/no culture of using dotNet on Linux. We all know it exists and it's open source now, but most Linux developers don't reach for dotNet.
On that note, how bad is .NET experience on Linux vs. on Windows? Besides the IDE (MonoDevelop is no Visual Studio).
I was about to dig into this topic soon - I'm particularly interested in how useful PowerShell is on Linux these days. On Windows, I appreciate the deep .NET interop it offers; I think it's closest experience to Lisp Machines that you can find in mainstream computing.
EDIT: Big thanks to the commenters who mentioned Rider, I didn't realize JetBrains had an IDE for .NET!
That said, to clarify my comment above, I'm more interested in the experience with the platform itself - how does .NET Core work on Linux in terms of performance, fragility, access to first-party and third-party libraries? Can I expect non-UI .NET code to port well to Linux? And, in PowerShell, do I get to do things like Add-Type "<insert bunch of C# code here>", or $foo = [Some.dotNET.Type]::new()?
JetBrains Rider is such a pleasure to work with.
If it wasn't written with tons of pinvoke to win32 or using Windows-only services (like WCF), then sure.
I've been building NET Core applications for Linux and exclusively on Linux for a few years now, mostly using Rider as the IDE. No complaints. I haven't seen any Windows-exclusive libraries worthy of any attention in all that time. Even libraries with tons of native code (like imagemagick) have Linux support.
The official MS tooling is behind what they offer on windows, but it's getting there. You can collect GC and memory dumps
https://docs.microsoft.com/en-us/dotnet/core/diagnostics/dot...
https://docs.microsoft.com/en-us/dotnet/core/diagnostics/dot...
traces:
https://docs.microsoft.com/en-us/dotnet/core/diagnostics/dot...
and here's the port of the veritable sos tool:
https://docs.microsoft.com/en-us/dotnet/core/diagnostics/dot...
it's pretty weird sentence "most Linux developers dont use dotNet", but on the other hands dotNet (Core) developers definitely do use Linux - I have dotNet things on Linux on prod for 2 years atm
- More support for UNIX and mainframe flavours
- Given the multiple implementations, there are plenty of GC, JIT and tuning options to chose from
- .NET Core is the Python 3 of .NET world, not everything from .NET Framework is available or portable to non-Windows platforms
- Several JIT optimizations quite common in JVM world are only now coming to CLR, like dynamic rewrite of native code
Source being compatible doesn't help when the libraries require code rewrites anyway.
Leave CLR in applications, where it belongs.
Also from my experience coaching and consulting there's a significant learning curve for .NET framework developers to work in the new core/.NET 5 systems.
https://docs.microsoft.com/en-us/lifecycle/products/microsof...
I bit the bullet and transitioned from WebForms to ASP.NET Core. I think it's easier than WebForms ever was. You can replicate 90% of what WebForms did with Razor Pages, but you also get much easier support for building RESTful APIs. And with DI built-in, I don't have to worry about junior devs forgetting to close/dispose database connections ever again.
If you've followed the .NET blogs at all, they have very good reasons they haven't prioritized support for these legacy systems. There isn't some cabal of platform designers in Redmond scheming on how to screw over developers at ossified orgs.
Not sure why I'm bothering. Everytime someone mentions .NET on here, someone comes out of the woodwork to complain about their pet features not being supported. Actually, come to think of it, it's often you in particular. We're going from .NET Framework being fundamentally tied to Windows to a new .NET that runs on everything. ASP.NET was built on HTTP.sys, a Windows driver for a web server. I don't know, is it terribly surprising that trying to migrate .NET off of Windows-only dependencies is going to break things downstream?
But if you still need this stuff that only runs on Windows, well, you got at least decade of support to look forward to.
I literally think you're asking for too much.
On my pet projects I use whatever I feel like.
It's certainly reasonable, but will take re-work for many applications beyond the project structure migration. I think in .NET 6/7/8 the .NET framework migration situation will improve to be mostly pure migration.
For example I migrated about 75% of a moderately large winforms system directly, but 25% needed rewrites due to missing core implementations for framework libraries. e.g. MS chart component not implemented in core yet, soap services changing to grpc, and various others. And this system was originally built to minimise third party dependencies and use the core framework or microsoft provided libraries wherever possible.
Is there anything in .NET 5 you're missing from .net 4.x?
MS is a big company, they aren't going to immediately migrate everything just like Google didn't migrate every single project to Go when they were heading it.
The biggest issue would be avoiding heap alloc, but that is actually not too difficult if you manage your data well enough. Most transactions in this business can be defined in terms of structs that are <1kb in size.
One other approach I have started to look at is sacrificing an entire high-priority thread to handling well-batched transactions (or extremely important timers) so that I never have to yield to the OS. In HFT, you always get an entire physical server to yourself, so eating 1 out of 32 threads is not a huge deal. In testing of this idea, I have found that I am able to reliably execute timers with accuracy of around 1/10th of a microsecond.
Java probably has more optimizations around GC suppression, but I feel like avoiding allocation in the first place is the most important bit. I believe there are already some ways to trick .NET into not running GC.
1 thread for ultra-low-latency timer execution
1 thread for processing the actual client event ring buffer, producing a consistent snapshot after each microbatch execution.
14+ threads for servicing HTTP requests (i.e. enqueuing client events), timers and redrawing client views using the near-real-time snapshots of business state.
My thinking is that if I can build something in this domain that is satisfactory, I could consider bringing it to an HFT firm as well.
[0]: https://www.youtube.com/watch?v=ug_UC4lxMr8 [1]: https://news.ycombinator.com/item?id=24896616
Java is not popular in HFT other than 2-3 shops using it.
C++ is best choice IMHO. It has become easy to use it and plus you can get compile time branch prediction with templates.
RAII also amazing. RAII looks like a GC but it's not but it's works like a GC ..
Branch prediction is a CPU level thing, and mis-predictions has a cost. You can do things like partitioning data beforehand so that a given if condition will take the same branch each time in the partition, but I don’t see how does it apply to templates at all.
You can’t avoid branches that depend on runtime infos, and those branches that depend on compile time known infos can be elided either automatically by the compiler if it can prove it is constant for example, or by things like constexpr, templates as you say. But it’s nothing too fancy, the dumb C-preprocessor macro can do similar things.
JIT-compiled languages can sometimes elide branches based on runtime data, so there is that.
Options. You are talking about the very tiny subset of Java users who actively worry about the relative differences in performance tweakability of different JVM implementations. Chances are that they routinely run multiple VMs and VM configuration sets in parallel on real world inputs to find the fastest configuration.
"Look at this nice VM, trust us, it's very fast, rewrite all your code and see for yourselves!" just won't be very convincing to those. With that crowd, .NET isn't competing against the JVM, it's competing against against an entire market of competing JVMs.