How Microsoft rewrote its C# compiler in C# and made it open source (2017)
medium.com
medium.com
Or imagine syntax reformatting but purely locally. Or semantically aware git diffs that can actually compare the underlying parse trees instead of just raw text. There's so much cool stuff you can do. I wish every language had a parser like this.
That's correct. Especially for parsers used in IDEs you most definitely care about exact locations of everything, including whitespace.
What I really like about Roslyn's AST (and I guess TypeScript's looks about the same) is that all those details are put into Trivia; ignoring those is just passing one flag to the visitor; you don't have to be aware of how many different types of AST nodes there are that might be ignorable. It also means that each and every whitespace and comment belongs to a syntax token instead of being other nodes interleaved with the normal nodes. This helps keeping comments where they belong even after re-arranging code in the AST – something that, e.g. ReSharper is atrociously bad at.
It's called abstract syntax tree to distinguish it from parse trees (a.k.a. concrete syntax trees).
Further, what Roslyn does is quite interesting even on top of preserving all of that. It does so in a way that lets it incrementally re-parse parts of the program without touching the rest- including position information for the rest of the file after the change, which has all shifted to match.
None of this was standard practice anywhere when Roslyn was designed, though it's starting to be picked up in other projects since.
You don't have to imagine it, this has been supported by IDEA for about a decade now at least for some languages (it works for Kotlin at least, I just re-checked before posting this).
> Or semantically aware git diffs that can actually compare the underlying parse trees instead of just raw text.
So like essentially all other corporations?
I say this as someone who happily uses a lot of the stuff they produce in the process. I really, really, intensely hope I'm wrong and my fears aren't realised.
Doesn't look like; they are doing EEE, just of a more 'benevolent' nature.
.NET between around 2010-2016 was a pretty boring/dire time for the community. Things were so bad that the Community started solving their own problems. This wound up giving us a bunch of great projects that the community stepped up and provided.
But now, when we fast forward to Today, .NET Core (Or, to put not too fine a point on it, ASP.NET Core) is doing a whole lot of what is at best steering developers away from good practices, at worst (un?)intentional EEE:
- EF Core: EF6 was terrible for a lot of reasons. EF Core tries to get people to access relational and nonrelational databases via the same API. Why? (Of course, with Cosmos the answer becomes obvious...)
- MS DI: Microsoft looked at the best examples of DI the community provided, and wound up deciding that breaking a number of established paradigms made it easier to build ASP.NET Core, even if folks who have written the best libraries/literature on DI in .NET explain why it's a bad idea.
- Serialization: 'System.Text.Json will be better than the other libraries out there' I heard that from multiple voices. It's still not really that good, but people try anyway because MS Stack.
And that's a bit of the problem. Many of these things _are_ needed for the 'safe shops'. I've been at places where deviating from MS Stack was specifically discouraged because they were worried about longer term support. Broadly speaking, if it had a big backer/sponsor (i.e. StackOverflow's Dapper) it was a non-issue, and it wasn't a -bad- way to make sure developers weren't just using $"{blogToolOfTheWeek}". Later on however, we shoved that mindset aside as much as we could. Instead we made it our goal to find the best libraries to solve or problems... and our productivity went way up.
"The world’s most valuable resource is no longer oil, but data"
https://devblogs.microsoft.com/dotnet/introducing-net-5/
> There will be just one .NET going forward, and you will be able to use it to target Windows, Linux, macOS, iOS, Android, tvOS, watchOS and WebAssembly and more.
But the fact that all this exists is a mess and thankfully MS finally has a good plan to clean it up.
https://www.hanselman.com/blog/MakingATinyNETCore30EntirelyS...
There's even fun experiments of AOT compiling to get interesting in "EXE golf" results from .NET Core applications such as getting them below 8 KB or running on Windows 3.1 (because why not):
https://www.hanselman.com/blog/NETEverywhereApparentlyAlsoMe...
.NET 5 will integrate further AOT capabilities as the Mono world is merging in, in addition to the raw marketing advantage that 5 > 4 for anyone still struggling to convince non-technical managers that .NET Core is a better investment in 2020 than .NET Framework.
I blame this directly on the immutable AST. While a nice concept in theory, it causes too many allocations, and is cumbersome to work with.
I predict another rewrite in 2 or 3 years.
I'm holding off on 2019 as much as I can. Between 'forcing' an upgrade for .NET Core 3.0 and the fact Resharper slows it down too much, I decided to give Rider a try.
I'm finding myself not missing VS a whole lot; on one hand Rider is taking way more RAM to start and load, but it stays pretty constant after the first debug session, winds up staying under VS for memory on longer loads (Especially if I've got multiple solutions open) and it's smoother than VS the whole time.
This isn't a cause of performance problems you're seeing. What's most likely happening is that the overall size of your code has increased a lot over time, causing performance issues due to issues that have existed for a long time, but weren't being felt yet.
The issue that's most directly related to this is here: https://github.com/dotnet/roslyn/issues/40300
However, immutable vs. mutable isn't related to the issue I linked. It's just about keeping more data around than is (perhaps) necessary. You'd see the same issue with a mutable AST. If you're curious about specific work that's being tracked you can use this label: https://github.com/dotnet/roslyn/issues?q=is%3Aopen+is%3Aiss...
And if you submit reports via the VS Report a Problem tool, with the option to collect a diagnostic trace, you'll generate exactly the data needed to fix issues that you're facing. The team is very keen on addressing performance problems, especially if there's diagnostic data that can pinpoint the source of a problem.
Not really. A simple empty project displays the same problems. You can try going back to 2017 right now with any project you're working on, you'll feel the difference instantly.
Intellisense simply takes longer to respond, and likewqise other editor functions.
Their feedback forums have hundreds of similar reports.
The issues are getting resolved in the public bug trucker, pretty quickly indeed, typically saying "can't reproduce, won't fix", sometimes "not a bug, won't fix".
I'm not sure the software problems are getting resolved.
You can also see that there are many more resolved issues in previews of that public release: https://github.com/dotnet/roslyn/milestones?state=closed
So there are a _lot_ of legitimate problems being fixed and enhancements being added over pretty short periods of time.
When something is resolved as "no reproduction" or "not a bug", that's because there was an earnest attempt to reproduce an issue with the latest bits set to go out to a release with no reproduction, or something is truly by design (e.g., user files an issue because they would prefer a feature to do something different than it does today).
https://developercommunity.visualstudio.com/content/problem/... (BTW, people have been reporting that bug for years now)
https://developercommunity.visualstudio.com/content/problem/...
https://developercommunity.visualstudio.com/content/problem/...
https://developercommunity.visualstudio.com/content/problem/...
I'm not familiar with this compiler but I'm currently writing a mostly immutable structure so I'm curious as to whether it's an issue.
Anytime a leaf node changes, all its ancestors have to be replaced, instead of just updating the leaf in place. (I'm aware of the red-black node separation, but I believe that in practice most of the tree is constantly regenerated all the time).
I realized it when trying to write a complex analyzer. I had to replace the tree all the way up to the project level. If you combine different chunks of the tree, each with a slight change, you're forced to recreate each of those chunks.
This is extremely wasteful, and no wonder the IDE behaves so poorly.
This is a fundamental problem that for various (outdated and bad) reasons the team hasn't fixed. They've been refactoring components to run in separate processes but it's slow progress and still won't solve the main thread running out of memory anyway.
So here is my base project config: https://github.com/capnmidnight/Juniper/blob/master/Juniper....
The most important part is the first PropertyGroup sets values for all build configs, in particular is setting LangVersion to 8.0. Framework 4.8 taps out at C# 7.2, but you can use most of the C# 8.0 features, including fully async streams if you manually set the language version. Features that aren't available are some minor things like the array ranges and indexing: https://docs.microsoft.com/en-us/dotnet/csharp/language-refe...
And here are my base project with the analyzers I use: https://github.com/capnmidnight/Juniper/blob/master/Juniper....
They're all ones provided directly from Microsoft, though there are a bunch more from other vendors: https://www.nuget.org/packages?q=analyzer
Then here is an example project using that targets file: https://github.com/capnmidnight/Juniper/blob/master/src/Juni...
You can see just how much the new SDK Style project file format simplifies things. There is no importing of any base Targets files hidden deep in Visual Studio's install directory anymore.
I manually import the .props and .targets file instead of using Directory.Build.props and Directory.Build.targets because I have other projects that use these configs, included via a git submodule.
Here is my .editorconfig file, where I set most of the rules related to Disposable types to errors: https://github.com/capnmidnight/Juniper/blob/master/.editorc...
And this Visual Studio extension makes .editorconfig files a lot nicer to work with: https://marketplace.visualstudio.com/items?itemName=MadsKris...
(BTW, I pretty much install all of Mads Kristensen's extensions)
And while I'm here, I'll give a shout-out to Viasfora for its syntax highlighting modifications that rainbow-highlight code blocks: https://marketplace.visualstudio.com/items?itemName=TomasRes...
And VSColorOutput for making the Output window in Visual Studio actually readable: https://marketplace.visualstudio.com/items?itemName=MikeWard...
(If somebody has--well, I wouldn't mind some lazyweb recommentations...)
Example: if my function was async and it had a .ToList() function call inside it would suggest to use .ToListAsync() and with a click auto fix it in your function
here is the repo for reference :
https://github.com/aviatrix/YARA
It took me couple of days to wrap my head around all of the new concepts, but it was quite fun! To get it integrated with the IDE, you need "CodeAnalyzer" class, and that gets executed automagically and provides annotations in the IDE :)
[1] https://github.com/louthy/language-ext/wiki/Code-generation
I take it you've never worked on a .Net solution with projects targeting both full framework and Core framework.
Core v1.x stuff was a nightmare - haven't had so many issues with versioning in 20 years. Core v2.0 was still pretty bad but each v2 point release made decent strides - and specific packages would get updated out of band at times to fix issues.
But to say .Net has never had DLL Hell is just wrong. Even pre-Core you could run into difficult situations with conflicting downstream dependencies of directly used packages.
https://en.wikipedia.org/wiki/DLL_Hell
The problem arises when the version of the DLL on the computer is different than the version that was used when the program was being created.
Other than the GAC, which was never recommended for use anyway, .NET has never had DLL HellAnd I've misspoken about GAC causing DLL Hell for .NET. It fixed DLL Hell, but introduced a new Strong Naming Hell.
It's entirely possible that I just don't know what I'm doing well enough to do this correctly, but I just couldn't get it to do the kinds of things that I wanted from it. My sense is that it is really great for injecting custom code that is used rather infrequently. I've had good results using it for building out reporting systems that are pluggable.
I feel defensive about calling it cumbersome though, and I can't imagine why something like your LUA vision isn't possible (though, I've never tried I just assumed someone would inevitably do this). For example, if World of Warcraft were to switch out their UI LUA extension system with C# I could totally imagine this being possible (though it'd be suicidal for their mod community). Likewise, if Unity were to begin using it for this kind of thing (if they don't already) I'd imagine it is possible.
would really appreciate bugs/comments on these docs pages of what else you would like to see.
See the section titled: Collecting a dynamically emitted assembly
It's a feature that's been around since .NET 1.1 and I used to use heavily 10+ years ago.
https://docs.microsoft.com/en-us/dotnet/api/system.appdomain...
The whole thing is now used in basically every library build we have at some point, even for the C# versions, as it ties in with our documentation writing process and places the correct API names and links for that product into the documentation, even though the docs start with mostly the same content for each.
I agree that lack of documentation makes working with Roslyn a bit daunting at times, although the API is very well designed and oftentimes it's very obvious where to look for something. I was also very impressed by their compatibility efforts. We started while Roslyn was in beta and upgrading through the releases worked without a hitch.
That said, large-scale rewrites don't happen just because engineers feel like doing them. In the case of Roslyn, there were multiple compilers (C# language compiler, VB language compiler, C# tooling compiler, VB tooling compiler, etc.), requiring an enormous cost to evolve C# and VB as languages while ensuring that end users of these languages had a good experience using tools like Visual Studio. Beyond that, there was a host of additional tooling to provide more advanced code analysis that was equally expensive to maintain and evolve as the languages evolved. This meant that the kinds of language and tooling innovations end users expected was more challenging to meet. A similar set of challenges - focused on problems centered around the end users - must exist to justify the enormous cost of rewriting a massive engineering system.
There's an old saying (from one of Unity's developers, I think?) that you can't run Hello World in C# or Java without an XML parser because so much configuration for things like locales ends up pulling in serialization libraries, reflection, etc. Things have improved in this area for both platforms but it's definitely still very difficult to trim managed code down as far as you can trim native code.
For something as critical as the JIT you also want to keep the generated code small (for cache efficiency) and if possible, PGO it - things that existing managed code generators aren't especially great at compared to clang or modern MSVC.
From my past experience writing/maintaining C# compiler and runtime code, I would never bother trying to port the JIT to C#. I don't think the returns are worth the massive investment. I'd sooner port it to a language like Rust for safety benefits or to some sort of hypothetical language that enables producing smaller/faster code to improve JIT performance. I'm not sure I'd invest in that either though because for most workloads JIT time is not the bottleneck (and you can optimize that out in many cases by pre-generating the JITcode). EDIT: Also, for workloads where JIT is the bottleneck, it's questionable whether it's possible to extract big gains out of optimizing the JIT because of the nature of JIT workloads - you may just be bottlenecked on memory bandwidth or instructions per clock. A JIT converting IR to machine code is not trivially vectorized.
People are figuring this out (again, from scratch) now with webassembly as all the modern browsers go through the churn and angst involved in answering the question 'can we actually JIT this whole 50mb executable from scratch at load time?' even though we already knew the answer back when WebAssembly wasn't a spec yet.
[DISCLOSURE: I get paid to work on Mono right now and was paid to help draft the initial WebAssembly spec, and before that my work on the JSIL MSIL->JS compiler was sponsored. So I have some massive biases here.]
We do know that the returns are worth it, because people are able to develop new optimisations in the Java version that people just won't attempt in the C++ version because the code is so much harder to work with (but possibly just due to the age of the C++ version) and its achieving 13% speedup over the C++ version in practice at places like Twitter, which is worth millions of dollars.
Also, it sounds like the old C++ JIT has to be maintained and shipped in the event that the Java-based JIT isn't available at startup, right? If so that makes sense as a stop-gap and it would be more of a tiered JIT, not a port. Tiered JITs are definitely proven, successful technology.
The claim is that the work to add new optimisations to the old C++ code is so difficult that people aren't prepared to do it. The Java code is easier to write and debug, enough so that people are managing to add new optimisations that they haven't added to the C++ code.
> the old C++ JIT has to be maintained and shipped in the event that the Java-based JIT isn't available at startup
Well only until the new Java JIT is mature. The AOT build of the Java JIT can be shipped in the binaries so it's always there.
> If so that makes sense as a stop-gap and it would be more of a tiered JIT, not a port. Tiered JITs are definitely proven, successful technology.
It's used as a new top-tier yes, replacing the C++ top tier. The C++ JIT is likely to go I think, in the medium-to-long term.
Note that you can still decompile the obfuscated code and look around (I’ve done so to attempt to debug a 3rd party library), but mangling all identifiers makes it quite hard to read.
Obfuscators are common if you're concerned about the symbols leaking.
If you don't include the .PDB debugging symbol files it's much the same as native binary code. The .NET virtual machine code is a little more expressive but is superficially similar to x86 native code.
In my opinion, if you're concerned about people reading your compiled machine code, the only solution is to run your app on a server and give users an API.
If the financial loss isn’t big enough to warrant a global license compliance scheme then I don’t see the source code would be that valuable (in a general commercial contex).
But I don’t think any form of software that is distributed to end users can be fully secured from unlicensed use. You always need a legal recourse if you actually want to stop unlicensed use.
The very-small niche where you can’t afford lawyers but want to force license compliance maybe isn’t a niche you can actually serve through a sound business.
So, rather than seek for automated technical compliance solutions (they don’t really exist without the physical lawyer component) maybe you should find the biggest market you can serve, use the most productive tool for the job and try to make sure unlicensed use can be noticed.
UWP makes use of AOT compiled .NET via .NET Native.
How do people secure they C++ code otherwise, it is quite trivial to use IDA or HexRays.