Introducing .NET IL Linker
github.com
github.com
It was originally used to reshape the entire Mono class library into a subset to expose the Silverlight class library API surface for Moonlight.
It was then used to link iOS and Android applications in MonoTouch and Mono for Android, and Xamarin continued to use and improve it.
The Mono Linker has an open architecture making it reasonably easy to customize how it processes code , and detect patterns specific to each platform to link them away. Xamarin added linker steps to do more than tree-shaking, and remove dead code inside methods. For instance:
if (TargetPlatform.Architecture == Architecture.X64) {
// ..
}
The entire if body can be removed if the linker knows that TargetPlatform.Architecture will not be X64.And now, it's the base for the .NET Core linker. Quite a journey!
"Self-contained deployment. Unlike FDD, a self-contained deployment (SCD) doesn't rely on the presence of shared components on the target system. All components, including both the .NET Core libraries and the .NET Core runtime, are included with the application and are isolated from other .NET Core applications. SCDs include an executable (such as app.exe on Windows platforms for an application named app), which is a renamed version of the platform-specific .NET Core host, and a .dll file (such as app.dll), which is the actual application."
So, this IL Linker is in support of that story, and aims to reduce the over all size of self-contained applications... Maybe it's time to start tinkering w/ .NET again.
Nowadays with dotnet core you can bundle the entire runtime with you binary. The linker is cool, in that it will create a strictly smaller binary, as in "hello world" might be 12mb instead of 50mb. That's cool, but it makes me wonder if .NET would be in a strategically better place today if the linker were present 10 years ago. How many companies didn't choose to write their app in c# simply because a customer might not have the runtime installed?
In any case I'm really impressed with everything Microsoft is doing with dotnet core and hope that I get to use more of it.
More than enough to make a difference against Java.
It took a Microsoft that understood open source and non-Windows platforms, and was serious about it, to ultimately get us to the point where .NET Core could be a thing.
Examples on our case, WPF, WCF, ODP.NET drivers for ADO.NET.
As far as I know, there are not plans for cross-platform GUI on .Net Core. And WPF is pretty much a dead technology.
I started my career working for a .NET shop a few years ago, and immediately fell into a love/hate relationship with the technology stack.
I loved C#, as it seemed to be a fantastically well-designed language that had learned a lot from Java's mistakes. Visual Studio was great, too.
However, as a longtime Linux and open source supporter, I loathed Windows and Microsoft's software culture. And honestly, the .NET open source culture - even now - is kind of lacking. C# shops typically wait for some canonical solution from MS rather than creating and sharing libraries themselves.
I really, really like the new direction that Microsoft is taking, and I hope that the open source community will embrace C# in a bigger way.
With NuGet I don't think it's as bad as it was, but I just caught myself hesitating when I found out this week that .NET has deprecated email sending functionality in 4.7 (or Core?) and is now recommending open source alternatives.
https://www.joelonsoftware.com/2004/01/28/please-sir-may-i-h... - Joel asked for this in 2004 - Better 13 years late than never!
Memory: The memory usage "looks" bad. I suppose maybe GC would kick in when it needs to. It has not been a problem for me, but when you see the code in C++ taking 1MB ram and a C# implementation taking 10MB, well, you have to wonder.
Performance: C# performs bounds checking, which I take it is a significant performance hit. I've not had this matter yet, but I've not converted everything over to .NET.
It's not to say that C# will always be 10x memory usage, however, it will always be somewhat higher. Even with a linker, when you run a .NET app, it's not just loading your C# code, but also .NET code that you call into (linker helps keep this slim) and of course the Common Language Runtime C++ code that executes your app and manages its memory.
This is currently being worked on. With the addition of Span<T>, many APIs will soon have a version that does not allocate and instead use a buffer you pass in (which can be managed or unmanaged).
https://docs.microsoft.com/en-us/dotnet/framework/net-native...
It does (The JIT that is), but it also often does not. For example a regular for loop does not bounds check anything. The newer x64 JIT ("RyuJIT") is a whole lot smarter than the legacy x86 jit was.
I imagine that once you start producing proper native binaries using whatever backend you want instead of a JIT that has a requirement of being FAST in addition to making code that is FAST, you can get even more optimization done.
Also, bounds check are becoming cheaper and cheaper as memory fetches become (relatively) more expensive with every hardware generation.
C++ code is usually faster than C# code but bounds checking is probably not the biggest culprit. If you use "idiomatic C#" with reference types, most operations have a lot of memory fetch overhead with pointer chasing etc.
It's possible to write pretty efficient C# code, and it can still look pretty nice. You end up using arrays more, structs more, obviously worrying about AoS vs SoA more. You should also look at the new ImmutableArray types that look like List<T> but have effectively no overhead compared to a plain array. After that comes Span<T>, which makes even more exotic things possible.
The optimal thing to do would be to memcpy several KB of data at a time from your heap-based memory, into your stack-allocated memory, do your processing in the optimized inner-loop, then commit the data back to the heap. I had to do this once to increase the throughout of a cluster of image processing servers.
I often scoff at the idea that C# isn't a good choice for performance-intensive tasks.
It's coming... https://github.com/dotnet/corert
And it most certainly didn't do anything about the dependencies on the runtime (VM, GC etc).
For compilation into a native binary, there is the corert project:
We've been compiling our large C# games to native code for some years now because iOS only supports signed code. The main problem is you can never JIT; this creates all sorts of limitations with generics, dynamics, expressions and reflection.
Xamarin has become better and better at handling generics but there are still a couple of things they just can't do anything about.
>Memory
Yes, the memory looks bad because it will not get completely collected until needed. The framework code is big, but that's a flat cost. It might take a little more memory for the runtime up keeping, but it is not 10x. However, memory profiling is a breeze compared to native, so it becomes easier to optimize on a large complicated project.
>Performance
If you don't want to bound check there are a lot of solutions, from simple to advanced. Most of the time we end up just iterating the whole array anyway, and in those case, there are no bound checks.
In the rare case where we really need raw power (graphics, path finding, physics, large file access), we use native code and thin wrap it with P/Invokes.
For the rest, we use System.Math. They are not used enough to be a significant bottleneck.
C# has value types, which follow the same memory model as in C (i.e. stack-allocated for locals, embedded directly into the outer object as fields). You can request explicit layout mode, whereby you can set an offset for every field manually - if you set them all to zero, you get a C union. Alternatively, you can request automatic layout mode, where the JIT is allowed to reorder the fields in arbitrary ways to optimize memory and/or access, which is something you don't even get in C.
It has raw pointers. Unlike object references, these have the usual C semantics - there's no GC involved there, and you're responsible for keeping the objects alive, so you can have dangling pointers etc. No boundary or null checks, either - it's zero overhead. You can do pointer arithmetic on them, including indexing with []. If you P/Invoke malloc or equivalent, this gives you C-style heap allocated arrays.
It has stackalloc, which is a language operator that's equivalent to the non-standard alloca() function in C - allocate a chunk of memory on the stack, and return a pointer. This immediately gives you stack-allocated arrays as flexible as C99 VLAs.
With generics, if a generic type parameter is a value type, the specialization for that type is separately JIT-compiled and optimized. This can be used in a way very similar to C++ templates, for zero-overhead inlined callbacks and similar shenanigans.
All in all, if you want to write high-perf code in C#, you certainly can. You'll hit the limits of the JIT optimizer soon enough - it can't be as good as a full-fledged AOT C++ compiler, say - but you can certainly shed most of the overhead associated with a managed language.
As for the rest at least we can hope Java 10 brings them.
You missed a few C# goodies from 7.0 up to 8.0, coming from the experience with Midori.
https://github.com/dotnet/roslyn/blob/master/docs/Language%2...
The binding is a bit funky (lots of `ref` and `gcnew` keywords sprayed here and there) but it works.
Source: discussion at https://github.com/dotnet/core/issues/915#issuecomment-32608...
https://github.com/gluck/il-repack
We do this at work with one of our utilities
It looks fairly active to me: https://github.com/dotnet/corert/graphs/contributors
Since reflection and dynamic invocation is a major feature of the runtime, they have avoided this optimization before.
Indeed. it's not that common, but I would expect it to happen. One example is that many IoC Containers allow you to scan assemblies and register all class + interface pairs that match given rules.
There has to be some safety-hatch to express "yes, I really want to keep that class and its methods. Don't strip them out".
This seems to be documented here: https://github.com/dotnet/core/blob/master/samples/linker-in...
https://blogs.windows.com/buildingapps/2017/05/19/introducin...
Putting aside the discussion on whether this is a good practice, I assume this would remove methods that are only called dynamically, or via reflection.
https://github.com/dotnet/dotnet-docker-samples/blob/master/...
That is still, absolutely speaking, twelve million bytes for not much more than "Hello World" levels of functionality, which means there remains plenty of room for improvement. I estimate the lower limit for this particular app is somewhere in the hundreds of bytes, most of it being string constants.
Hello World doesn't benefit from that 12MB. Your real app would.
12MB is not for Hello World. 12MB is for Hello World + Common Language Runtime (garbage collection, bounds checking, memory protection, execution environment etc.) + the .NET standard libraries (LINQ, task parallel library, etc.)
You're probably going to need those things in a real app.
If you're actually doing Hello World -- and that is all you're doing -- and if using 12MB of your 16GB of RAM is unacceptable for you, yes, by all means write some native code.
The number might shrink as the linker gets more aggressive but i wouldn't count on it much.
The application IL code is perhaps 50 bytes to for hello world.
A Hello World assembly in .Net Core 2.0 is 4.5 kB.
Out of that, the IL is 11 bytes, though that doesn't even include the "Hello world!" string.
You're probably going to need those things in a real app.
I'm not saying that. From the page itself: "The linker removes code in your application and dependent libraries that are[sic] not reached by any code paths. It is effectively an application-specific dead code analysis"
Going by that description, Hello World should be the most trivial of test cases for removing unneeded cruft, and yet it still leaves plenty behind.
Also, don't assume that everyone has 16GB of RAM (I have a tiny fraction of that), or that even if they do, your app should be using all of it.
- no removal of native code. - no removal of code from the base managed assembly.
We would like to fix both and both are a large win.