Incremental Compilation
blog.rust-lang.org
blog.rust-lang.org
Nowadays — and if you are doing interviews will understand — people expect you to write flawless code from scratch using a basic code editor (looking at you HackerRank) with no time for tests. Sixty minutes to write a solution for four challenges, without autocomplete, without tests. Even the genius co-worker that comes with an optimized solution for every problem would need more than fifteen minutes to solve those crazy/random challenges that at the end will measure your ability to resolve those specific questions but not problem solving skills in general.
So is trial-and-error bad or not?
[1] Much of a programmer’s time is spent in an edit-compile-debug workflow.
But the journey is part of the magic. Anyone expecting candidates to produce perfect code is doing it wrong. HackerRank is garbage.
I want to understand how you think. I want you to communicate and ask questions. Yes, it can challenging to ve without your tools, with only a whiteboard, but that's okay.
If I'm interviewing you and you make a typo, great! I don't care. Maybe I'll point it out to you, but I probably won't, and I certainly won't dock you for it. And yeah, you can debug. I often encourage candidates to walk through their code with sample input, which I usually want them to come up with too.
This demonstrates an understanding of how your software works. I don't actually care if it compiles. I don't care if you're a wizard in emacs or can type a million words a minute.
I also don't care if your solution is optimal. In the real world they rarely are. But, I want to talk to you about ways we could optimize it, and maybe I'll ask you to show me what you mean.
Ultimately I want you to explain stuff to me, talk technical, and display your train of thought.
Trial-and-error is fine. It's a tool, and there are no silver bullets.
In fact, later in your comment you allude to this yourself -- "with no time for tests". You're worried about interviews not having time for tests. But tests are a edit-compile-debug feedback loop. If a test fails, you need to debug it, fix the issue, and recompile. Often you have to do this many times. This is precisely what the blog post is talking about.
You're generalizing the term "trial and error" to mean something more than what your teachers probably meant.
(I do agree that those interview practices are silly. But the blog post is not talking about stuff like that.)
It depends.
Efficient trial-and-error should help you acquire a good understanding of the solution. When your program finally works, are you really sure why? If not, you might be doing it wrong.
You might want to google "test driven development". Basically, it's a programming discipline which heavily relies on shortening your feedback loops (i.e reduce the time between the moment you make an error and the moment you detect it). While it certainly is controversial, it can be seen as a serious formalization of "trial and error".
(Having to split up your code manually into separate .cpp [or what have you] files doesn't qualify as incremental compilation—that's just separate compilation. Incremental compilation is when the compiler automatically figures out what needs to be recompiled, even when all you hand it is a big blob of code that hasn't been manually split up by a human in any way.)
What makes this incremental compilation different from that of, say, Go, is that it works inside a package. Before incremental compilation, Rust and Go had the exact same model: package-at-a-time compilation. With incremental compilation, Rust is more fine-grained than that: Rust can compile specific parts of a package while leaving the other parts untouched.
If I recall correctly they demo it on this video
https://www.youtube.com/watch?v=pQQTScuApWk
Also IBM had a C++ compiler version for OS/2 based on their Visual Age for... stack that used an image store for code, offering a Smalltalk like editing experience for C++, but it was quite resource heavy for early 90's machines.
It had something to do with their CSet++ offering.
Delphi uses the source to figure out dependencies. That is, the compiler has make logic, and traces through the 'uses' declarations to recursively discover all the source and object files. It compares timestamps to discover which object files are out of date and need to be recompiled. Thus, if you touch a single file and do a compile, only a couple of files will be compiled - the main program source and the modified file. This dependency logic is in the compiler, not a separate make program.
If this isn't the compiler doing incremental compilation, I think we have a terminology problem. Because this is what is meant by incremental compilation as the term is applied to almost all production compilers, and if you persist in making a distinction, you may be confusing more than clarifying. ISTM that you're memoizing more parts of the compilation process than has historically been done in incremental compilers. It may be clearer to use a more qualified phrase to express this, rather than shift the generally accepted meaning of a term.
He explicitly says "Incremental compilation is when the compiler automatically figures out what needs to be recompiled".
(He does say that production compilers don't have incremental compilation, but that's not a matter of definition, and I read that to mean "usually don't")
He also says that merely using the separate files to deduce this (which the parent says is what Delphi does) is not incremental compilation.
There's no manual splitting of .cpp and .h files. No manual management of dependencies.
There is some (ongoing?) work on low-cost incremental computation by differentiating lambda calculi [1] [2] [3], and it seems like an incremental compiler could maybe be a good use case for it.
[1] http://www.informatik.uni-marburg.de/~pgiarrusso/ILC/
[2] http://bentnib.org/posts/2015-04-23-incremental-lambda-calcu...
If I was to implement incremental compilation, I'd start with per-module I guess, because it's the atomic unit for a compiler, so it would be a lot simpler. That's why I'm curious.
https://manishearth.github.io/rust-internals-docs/rustc/dep_...
Note that as per:
https://manishearth.github.io/rust-internals-docs/rustc/dep_...
The DepGraph that we actually care about is specialized to talk about DefIds. In rustc, DefIds are attached to the following things:
https://manishearth.github.io/rust-internals-docs/rustc/hir/...
This means dependencies are analyzed down to local variables and uses. But even beyond that it tracks things such as borrow checks, linting and more!
Right now, they aren't. This will change as things improve. Remember, this is just the groundwork to allow incremental compilation to work at all. There are lots of places where the compiler now needs to be updated to use that framework.
> build outputs are no longer deterministic functions of just the build inputs (which is useful for example for distributed build systems).
You don't have to use incremental compilation if you care about this problem. But you will want to, because it isn't a problem in practice :)
There is no reason why distributed build systems couldn't simply be made aware of incremental compilation.
Caveats and fine details missing all over the place. More focused on the general point.
Are you sure about this (in the current implementation)?
I would hope that the actual output is deterministic based on the inputs. That is, changing function C in a file that has A, B, and C should be the same as a fresh compile of the same file. Even if the actual work that is done is different.
On other architectures, it's the linker's job to convert symbolic addresses and it can choose all the addresses in one go as it's the final stage producting the executable.
There is no human benefit to having to copy and paste function signatures—and worse, the bodies of functions you want to be inlined, including all templates—into header files. There's also no benefit to having the compiler parse hundreds of KB of header files over and over and over again.
> You can build java files the same way without header files, worst case scenario is you have some fat compilation units.
Java (production implementations used in practice) performs whole program optimization via its JIT. It's nice for Java, but that isn't how Rust works.
> Java (production implementations used in practice) performs whole program optimization via its JIT. It's nice for Java, but that isn't how Rust works.
Whole program optimization isn't generally something you want on dev builds, which is where incremental compilation is useful. For a production build why wouldn't you do a full rebuild anyway?
By "whole program optimization" I also include things like generic/template instantiation. Whenever you use, say, a HashMap, you have to recompile the implementation of the HashMap specialized to the size, alignment, destructor, etc. of types you're using with it.
> For a production build why wouldn't you do a full rebuild anyway?
When I profile apps written in Rust, it's important for me to be able to get good turnaround time on optimized builds.
With project Valhalla: https://en.wikipedia.org/wiki/Project_Valhalla_(Java_languag... there's been talk about specialization for value types.
https://adoptopenjdk.gitbooks.io/adoptopenjdk-getting-starte...
In any case, Rust stores the interfaces to libraries in serialized binary form (including documentation, following a standardized format understood by rustdoc), and they can be readily extracted from crates using tools that ship with the language. So what you describe isn't a benefit of headers anyhow.
I definitely think there's benefit to being able to view interface as separate from implementation, although that might well be better supported by tooling (editor folding, documentation generation) than manual maintenance - which you seem to be doing well.
I don't know whether there is benefit to being able to edit interface separately from implementation. Probably not much of one, but I could certainly be convinced otherwise.
It's possible that you may want to make cosmetic changes (or maybe change how you're specializing something where the implementation is more polymorphic?).
Still, I certainly don't see much of an argument that being able to edit the two separately has much value. I'm just not entirely convinced of non-existence.
Yes you still obviously have to keep them in sync, but the point is you (as a commercial software vendor) can ship just the headers and keep the internal implementation details secret and change how things work completely. Which is useful for commercial vendors of software.
.h files are an extra amount of work you have to put in. In the case of plugin APIs, this is work you'd have to put in anyway (for a fraction of the .h files -- the ones which are part of the API). But that work doesn't have to be put in elsewhere.
This is an argument for interface description files for plugins, but that could always have been done by a language with idl files or something. Indeed, apps in most other languages define interfaces via .h files because that's the de-facto plugin API description format, but they don't use .h files everywhere in the source. They only use them where necessary for the API.
Most languages with native support for modules do have the necessary information stored in the binary file to enable this scenario.
The only problem is exposing those plugins across languages, but it isn't a big problem if the OS has a rich ABI instead of a C one. For example COM on Windows, or WinRT the improved version of it.
The C++ ABI is currently not portable anyway. So the concern about templates (or any C++ features) forcing you to put the implementation into the header file would not apply in the scenario dllthomas was referring to:
If you want to distribute your library as a binary object and a separate source-form interface/header today, you'll have to use a (wrapper) C API. Regardless of how "modern" the C++ code is.
Of course, it would be nice to have portable C++ objects some time in the future which would change things... :)
This is only true if you want to target different compilers. If you ship your binaries targeting a specific compiler (which is what just about every company does that I've ever worked with), you don't have any abi issues.
See this article by Herb Sutter on the topic: https://isocpp.org/files/papers/n4028.pdf
Not really.
Those of us on Windows make use of COM, or since Windows 8, UWP components (formally known as WinRT).
If you want users of any standards-compliant C++ implementation to be able to use your library today, you'll still have to go with c-abi symbols or ship the sources. All other worakrounds are vendor-specific and not part of the standard.
[Of course even objects containing only C symbols are not portable across platforms either, but at least the C ABI/calling convention is more or less strictly defined for any given target platform. Assuming no other platform-specific stuff like glibc is used]
This works at the abstract syntax tree level.
If you change a file that many other files depend on Makefiles will recompile all those files. This approach will only recompile the parts affected by the change in AST which will likely be significantly less.