How a Zig IDE Could Work
matklad.github.io
matklad.github.io
Also see: comptime interfaces proposal https://github.com/ziglang/zig/issues/1268
It would also help creating a uniform standard library, for example creating a Writer interface instead of having specific functions for files, sockets, buffers and all kind of types
You could create a new type like Dynamic(InterfaceName) that handles dynamic dispatch somehow
From the article:
The whole process is lazy — only things transitively used from main are analyzed. Compiler won’t complain about something like
fn unused() void {
1 + "";
}
This looks perfectly fine at the Zir level, and the compiler will not move beyond Zir unless the function is actually called somewhere.And look, I can appreciate lazy evaluation being awesome in some circumstance, but `fn wat() void { 1 + ""; }` just spits in the face of static typing
It's also not nearly as bad as a dynamically typed language: you might have broken code you won't realize until you try to use it, but you'll always realize at compile-time.
In dynamic languages you refactor and don't notice you broke stuff until runtime!
I don’t know if I consider this a nice thing when thinking about it from a systems language lens. When working at the low level, I like the immediate clarity and compile-time feedback that preprocessor-like checks bring. They’re not perfect, but I think I prefer that instead of relying on laziness.
> It's also not nearly as bad as a dynamically typed language: you might have broken code you won't realize until you try to use it, but you'll always realize at compile-time.
This is correct, but to me, the elephant in the room here though is that this makes it much easier to introduce broken technical debt into a codebase. This likely means that when conducting large-scale refactors in a Zig codebase, as the amount of potentially broken code increases, the count of hair strands remaining in one’s scalp decreases.
Zig is an awesome language, but frankly, I think compiler laziness like this is pretty odd for a statically-typed systems language and not the behavior one would intuitively expect.
you might [reasonably] say you can't care about that kind of simplicity, which is fine; but radical simplicity is definitely one of the major charms of zig.
(and you don't use the zig feature by commenting/uncommenting, it's via inspection of the environment/flags in the comptime code.)
Moreover, I can grep for preprocessor defines, attributes, and such by name.
There's no way to grep a project for code commented out to prevent execution on some platforms.
You don't need a full preprocessor with macros and other evil stuff. Just #define and #ifdef would solve this without allowing wrong code that randomly disappears. And learning and understanding these two directives requires zero mental effort.
Yes. It's exactly that, because Zig very aggressively evaluates values that are known at compile time. It's as simple as:
if (builtin.os.tag == .linux) {
// Everything here only exists if compiling for Linux.
}
The standard library uses this quite a lot for OS/CPU specific things.If it's a comptime constant by "mistake" then it would have resulted in the same wrong behaviour at runtime.
The same feature is also used for specializing generic containers to specific types (an idea that became famous after the resounding success of std::vector<bool>) so that for example arraylists of u8 expose a Writer but arraylists of f64 do not. Most code that is conditional on T == u8 won't typecheck if T != u8.
I guess more generally these features are used for implementing fmt, json encoders and decoders, etc. None of these thing would work if code conditional on the type had to typecheck for all types.
At the end of the day, this is pretty equivalent. Not sure if this is a net gain or not.
Conditional compilation is needed, moving the feature from the preprocessor stage to a lazy evaluator is probably costlier for the compiler and it might also be costlier for the IDE to highlight active/inactive parts of the code.
From the user point of view, this is very close, mostly a cosmetic change.
How is this better than a line in the language itself?
All statically compiled languages need to skip compiling such code one way or another. Zig is just thinking this idea to the end and doesn't compile any dead code. I guess the recommended workflow is to have good test coverage. The test basically turns dead code back into live code.
Anyways it's not entirely crazy to think that maybe this sort of thing should be the province of static linters
There's support for invoking with qemu, wine, etc in the build system which allows you to run tests for other platforms.
A case where I've taken advantage if this is to have data structures and code adjust based on the expected page size and cache-line. There are also cases where things may be both comptime and runtime known depending on which platform that you target which is easy to handle by having the maybe comptime value be first in conditional branches.
As a comparison, what C++ does is that it differentiates between "dependent" and "non-dependent" expressions, where "dependent" means that it depends on some template parameter (analogous to comptime parameter here). "non-dependent" expressions are type-checked eagerly, while dependent expressions are only checked at instantiation time.
The design is certainly more complicated, even more so when you consider function overloading and name lookup, which can give different results if done late or early. But there are some upsides, and zig wouldn't need to deal with the additional complications of overloading.
Getting rid of the preprocessor and conditional compilation may not have been necessary. (as the feature is needed anyway and have to be implemented in a different manner)
The thing is that comptime, even if better than macros because everything is written in the same language, might end up not being much better than C macros.
Many of the macro pitfalls still applies, if you think about it...
In this line[0] a radix sort selects a function to read the next 8 or 16 bits of the input words. The C++ version of this code has to have readOneByte and readTwoBytes take an additional template parameter to control the return type:
static constexpr auto readBucket = (BYTES_PER_LEVEL == 2) ? readTwoBytes<idx, udigit> : readOneByte<idx, udigit>;
because both parts of the expression are evaluated by the compiler and must be the same type, including the function's return type. This means we end up instantiating nonsense functions like readTwoBytes returning u8. Elsewhere in the same file, some should-be-impossible template instantiations have had to have their static_assert's removed, because C++ will instantiate them and then not insert any calls to them into the generated code. So one cannot use static_assert to say "If this is reachable, that's a bug, please fail compilation and return an error" because template instantiation and the execution of the static_assert does not imply that the function won't be immediately discarded as unreachable.[0]: https://github.com/alichraghi/zort/blob/main/src/radix.zig#L...
I saw some people to propose that ternary should work potentially like if constexpr, if the condition is a constant expression. Honestly, IMO the ternary operator is already cursed, and it should not be overloaded with more responsibility.
I think this should work without nonsense return types:
static constexpr auto readBucket = []{
if constexpr (BYTES_PER_LEVEL == 2) {
return readTwoBytes<idx, udigit>;
} else {
return readOneByte<idx, udigit>;
}
}();I think it's possible to do comptime better, by putting comptime into separate files that can then "export" generated code.
At best it should give you a warning that the function is not used, that's it
Oh, and the compiler may not know it's unused until it processes all files if it's a public function.
Maybe we need to do the same here
I think Eclipse is mostly popular these days as a platform for companies offering custom tools. For example Team Center (which is awful btw). Honestly Eclipse was always pretty awful and I have yet to use any Eclipse based software upon which that awfulness didn't at least leave a mild stench.
Check out this list: https://en.wikipedia.org/wiki/List_of_Eclipse-based_software
Maybe universities still get students to use it?
However, I wanted to make the distinction that while @systems asked about Eclipse as a Java IDE, which it is objectively terrible at, Eclipse is like NetBeans in that it is multi-language[1] and is likely more accessible as an open-source C++ IDE for university use than trying to get CLion licenses (and, ahem, then explain CMake to students :-/ )
1: yes, I'm acutely aware IDEA is also multi-language but the open source distribution only covers Java and Python, unlike Eclipse and NetBeans
I was lucky to have one course in 2018 in which the professor actually included the most bare bones `git` usage.
I actually never had a class in which an IDE was recommended. In 2008 I did have a class on C (first half C, second half Java) in which we were restricted to using C89 because that was most widely adopted in industry. I really missed block-scoped variables in for loops (`for (int i = 0...)`).
Eclipse headless is also what powers VS Code Java support, with Red-Hat and Microsoft as main developers.
You know, the reason why JetBrains had to rush out Fleet.
To be fair, I use VS Code as well. It works well as a text editor.
I guess it depends on the locale/company/environment?
In most conferences, online videos, as well as among the people I know personally, JetBrains IDEs (IntelliJ IDEA for Java) seem to reign supreme: https://www.jetbrains.com/idea/ They have a community version, personally I pay for the Ultimate package of all the tools. They're slightly sluggish, want a lot of RAM, but the actual development experience and features make up for that. Hands down, the best set of IDEs that I've used from a pure development perspective - refactoring enterprise codebases mostly becomes something you can actually do with confidence. Running tests is easy. Integrating with app servers, containers, package managers or container runtimes is easy. Even remote debugging is easy to set up, as is doing debugging, or even testing web APIs. I'd say that all of the features that should exist do exist, which is more than I can say about many other IDEs.
I know that Eclipse is sometimes used more in an educational setting, however there are also both some specialized tools, as well as customized versions for something like working with Spring in the industry: https://spring.io/tools In my experience, the idea behind the IDE is nice (a platform that you can install whatever you want on, entire language support packages, or specialized tool packages), but the execution falls short - sometimes it's unstable, other times it works slow and so on. That said, it's passable.
I would say that personally I'd almost prefer NetBeans to Eclipse, even after it was given over to the Apache Foundation, which have released a few versions since: https://netbeans.apache.org/ It seems to do less than either Eclipse or IntelliJ IDEA do, but for general purpose Java editing and limited work with other stacks (PHP, webdev stuff, some C/C++) it is good and pleasant to use. However, if you have projects that get close to half a million lines of code, it does just kind of break and gets way slower than the alternatives. It still somehow feels more coherent than Eclipse to me, would pick it if IntelliJ IDEA didn't exist.
Some also try doing something like using Visual Studio Code with a Java plugin: https://code.visualstudio.com/docs/languages/java That said, I only used that briefly when I needed something lightweight for a netbook of mine, the experience was somewhat underwhelming. The autocomplete or refactoring wasn't as good as IntelliJ IDEA and just felt a little bit tacked on. Then again, that was a while ago, I don't doubt that progress is being made.
Eclipse tried to use the same approach for maven and wrote their own parser/processor to inject the Maven build file. This failed miserably.
I know that this is a standard assumption, but I personally believe that it is a mistake to support this use case. It is a holdover from the batch processing days when getting feedback on your program could take minutes or hours. In those conditions, it was absolutely essential to try and proceed with a partial compilation even in the presence of errors or else iteration would have almost been impossible.
Now we live in an era of fast personal computers. With a language that is designed from the ground up with IDE support/rapid iteration in mind, you can get feedback with every keystroke. Everything is easier to design if you abort immediately on the first error, whether it is syntactic or semantic. Designing a parser with error recovery from syntax errors is a particularly dark art. On some level, you have to resort to guessing. It may work _most_ of the time, but to me there is nothing worse than a tool that is maybe correct.
When you advance past an error all of the code below the error is suspect. To give a trivial example, let's say you rename a function with 100 call sites without some kind of refactoring (automatic or manual). What benefit is there in showing the 100 new failures when the root cause is exactly where your cursor already is? There are cases like this that are even subtler where you are bombarded with downstream error messages that are confusing and taking you away from the actual root of the problem (c++ template errors come to mind). You may as well just gray out everything below the first error and reprocess it only when the error is fixed.
Diffs often seem to fail to represent the actual change because they consider the delta from the last commit, which isn't the way we write code most of the time. If I go and insert a bunch of closing braces into the middle of a function, it's almost always because I'm dividing the function into two functions, or adding some missing error handling, or a corner case. So from the standpoint of a DAG representing the parse of the file, most of the time I expect the functions above and below to still be available in autocomplete even if I haven't balanced the tokens representing code blocks yet.
If you saw a bunch of functions and I break one, I expect the IDE to consider the function broken, not the file.
...
def add(x: Int, y: Int) = x + y
...
let x = ad|> I am using |> to denote the cursor location
...
I agree that where the |> is that the IDE should be able to autocomplete `add` for you. What I am saying is that it shouldn't process anything below the let statement because the code is broken there. Many IDEs and compilers will continue processing the file after the first error so that they can report all of the findings. That is what I am suggesting they not do.
That makes sense. What I am describing doesn't preclude this organizational structure. It depends on the language design. You could either support c style forward declaration that you put at the top of the file and would be available for completion even if the implementation is below the first error in the file. Or the IDE could provide folding of the implementation so that you can scan the api without drilling down into the details.
> That also means I almost exclusively call functions that are defined further down in the file.
Again, to clarify are you calling things before they are _declared_ or _defined_?
As a compiler author (I have written a self-hosting compiler for an unpublished new language), I dramatically prefer implementation simplicity to organizational flexibility. I respect your preference but believe that ultimately the more free-form and flexible a language, the more complex and slow the tooling, both of which lead to more bugs in the final artifact. But I certainly can't prove my point.
The obvious is that it requires forward declaration. However, depending on how back & forth interdependence is, it requires multiple "tiers" of forward declaration to effectively iteratively redeclare (or augment? Now that's some added complexity for all parties involved…) incomplete types until you can complete them. It's one thing to have a nice list of "here's what exists", but it's another to manually detangle dependency graphs. It's bad enough in C++ which already doesn't allow very much interdependence, but it'd be completely infeasible for any language more complex than glorified assembly.
Next, how would compile-time execution work? It's one thing to forward declare the existence of something, but how do you execute it without knowing its definition? You literally have to add another pass. Similarly, how do you make an inline function call?
One-pass compilation also relies heavily on searching and referring back to data from earlier in the pass, making it rather cache-unfriendly unless you build data structures as you compile, but now you're using massive amount of memory since you have to build these structures for everything in the entire input. This scales incredibly poorly. If you use one thing from some imported header, you now need to add that header and its entire dependency tree into your one pass. This isn't the 70's anymore; more passes ≠ more slower.
It's actually the compiler and IDE people pushing against this one-pass mindset, because it's "simpler" but just worse for everyone… except maybe the developer reading a header file instead of documentation. And do note that complexity comes in forms other than fancy data structures & algorithms in compilers/tooling. I'd argue that a manual flattening of a real world dependency graph is much more complex and harder to grok & maintain. Regardless, it's the compiler/tool developer's job to take the burden of complexity to better serve their users.
Why?