C++ Modules Might Be Dead-On-Arrival
vector-of-bool.github.io
vector-of-bool.github.io
import std.stdio;
means look for: std/stdio.d2. Because Windows has case-insensitive filenames, and Linux, etc., have case sensitive filenames, we recommend that path/filenames be in lower case for portability. This annoys some people.
3. There are command line switches to map from module names to filenames for special purposes. They're very rarely needed, but invaluable when they are.
Overall, it's been very successful.
I'd say "valid X language identifier characters" should always be ASCII.
I never understood the BS fad for unicode identifiers.
Wanna allow some math symbols? Maybe. The full unicode gamut, so that you can have a variable named shit emoji? Yeah, no.
Or why not full unicode text support? Is there any real reason besides "some people might want emoji and I don't like that"?
But no, the totality of the argument always reduces to: "I'm not used to this and would find it inconvenient"
This means the code can be read by anyone anywhere in the world on any operating system and that string payloads can similarly be read by anyone anywhere in the world.
Uh, no? You are not supposed to be able to read this valid Unicode string literal in Korean: `"그뤼고 이 문좌열은 일부려 기ㅖ버역을 어럽게하러고 오타비문이 산개해 있구먼요."`
Also a significant portion (and possibly the majority) of codes would be ever read and written by a small group of people, often sharing a common language other than English, so non-English code is just fine for them. If you are saying that a public library should be written in English, I almost agree---there would be some exceptions though.
public class 그뤼고
{
private 이 이 {get;set;}
public 그뤼고()
{
var 문좌열은 = 이.문좌열은;
var 기ㅖ버역을 = Enumerable.Range(0, 10);
var 오타비문이 = 기ㅖ버역을.Select(요 => 요);
}
}
public class 이
{
public int 문좌열은 {get;set;}
}I have seen numerous instances of pseudo-English when it comes to naming. It is hard to name things in non-native tongues. When reasonable, reducing that overhead can be indeed beneficial.
Plus it is pain to alt+shift between languages all the time.
But in cases where the program will be dealing with some concept that doesn't exist in English, being able to refer to things by their actual name, in the native language (assuming that's also the native language of the customer and development team), is much better than inventing a confusing and unnatural English translation.
A programming language's identifiers is not the place to express one's national identity. They should be utilitarian, and easily understood by programmers across countries.
Since you're already supposed to understand the syntax of every major programming language (which is based on english) you can make do with english keywords too. Nothing worse than opening some code to find bizarro foreign language identifiers.
(And I'm no native english speaker, so I'm not speaking as someone who's ok with this because english is their language or ASCII fits their default keyboard layout: it just makes sense).
> Nothing worse than opening some code to find bizarro foreign language identifiers.
I'm not a native English speaker, and I have only to loose in this scenario.
Just like people know they have to write in English on HN to communicate, they also write their code in English when they plan to open-source it and share with the rest of the world. As for closed-source projects ... if your company doesn't conduct its business in English, why force the code to be in English? The only people whom that'd benefit are never going to see it.
We live in a multicultural world, one for which English as a default doesn't make sense in every context. Yes, it may be the case that Chinese, Spanish, Greek, German, Arabic, or other non-American ideogram using programmers write code primarily meant to be used and understood within their own culture. I see nothing wrong with that.
Also what about the fact that while Mandarin is the largest language/dialect it's far from the only one in China? Using English/ASCII means all the programmers in China can understand each others code...
I mean, yeah, this can be made a requirement of programming languages: After all, it was such a requirement for a long, long time. But it doesn't have to be anymore.
BTW, full disclosure, I'm French and I find code written in french completely fucking unreadable. And that's 100% ASCII. As I said I believe code should be written in English, but I also don't think we should have essentially-artificial barriers for people to enter something as important as programming; those barriers only end up eroding the culture in question.
Compare итератор and 迭代器, which are complete mysteries to me produced by Google Translate. If my intention were to reach as many people as possible, I'd use "iterator" (which, coincidentally, works for English and my native German).
If once worked in finance and there is a difference between GAAP accounting and German accounting rules. If my algorithms used English terminology to be consistent with technical terms this would be confusing inneach review. Using German terms (even combined with English "get" or "set", like "getBetriebsertrag") there was beneficial, even though it always confussed new members of the team.
(Indeed, we should arguably move away from the notion of a single character string as the only human-facing semantics that an identifier is associated with-- there should be a higher layer, perhaps with multiple choices of e.g. native language, formatting and the like. Human facing semantics are closer to "literate" documentation than to anything that compilers should have to deal with. Yes, the "native", underlying representation should still be something that we can somehow make sense of - I'm not saying that our identifiers should be GUIDs or anything like that! But it will only be resorted to in a pinch.)
I don't see a dichotomy (much less a false one). What are the two options I separate artificially?
I'm saying just don't impose regional alphabets (other than AZaz that's already par for the course with the syntax of all major programming languages anyway) and regional words into source code.
That's also a problem in most other languages; just a couple of days ago, someone else's C++ code didn't compile on my machine because they had accidentally included <something/whatever.h> when the file was actually named <something/Whatever.h>, because macOS is case insensitive. I had the same experience with JavaScript some months ago, that time because they were running Windows.
On the "filenames must be valid identifiers" thing; I really wish more languages would start allowing kebab-case in identifiers. That's also absolutely not a D thing, more of a common complaint about most languages.
Exactly which operators should require whitespace and which don't is up for debate, but in my personal opinion, requiring space around infix operators and letting prefix/postfix operators not require a space would be appropriate. Nobody wants to have to write `myArray [i]`, but I think most people would be willing to give up `i-1` and instead write `i - 1`.
t[x] = t[x-1] + t[x-2]
looks more readable to me than this : t[x] = t[x - 1] + t[x - 2]
Another example : y = a*x1 + b*x2 + cWould adding whitespace sensitivity really be a problem? You already need whitespace to separate identifiers, so it's not a totally foreign concept in mainstream languages.
It seems like we've been making a weird trdeoff, by disallowing kebab-case just so we can smash our operators together with our operands.
Some people don't want to bother putting a space between operators and operands, and proponents of kebab-case just don't want to push the shift key to get an underscore.
I agree that nobody would want to write `i ++` or `foo [10]` or `myvar . mymember`, but I think a lot of people could get behind `10 - 20` and `foo && bar` instead of `10-20` and `foo&&bar`.
Also, the reason to prefer kabab-case for me has nothing to do with avoiding a keypress. It's that I find kebab-case easier to read.
x = (-b + sqrt(b**2 + 4*a*c) / (2*a)
x = (- b + sqrt(b ** 2 + 4 * a * c) / (2 * a)
... Although, one could argue that allowing tightened multiplication and division are enough. x = (-b + sqrt(b**2 - 4*a*c) / (2*a)Aside from C-style type declarations ("unsigned int x;"), C-style syntaxes seem to always have ways other than whitespace to separate identifiers.
Like (using some JavaScript in a hypothetical example) I can't think of many concrete reasons why this is easier to parse:
let first_number=2, second_number=2, answer=first_number-second_number;
...than this: let first number=2, second number=2, answer=first number-second number;
Although, of course, some languages—most Lisps, Tcl, and Red/REBOL come to mind—actually do rely on whitespace and whitespace alone to separate identifiers in many situations, and something like this would likely be unworkable there. let let x = 5;
let x = 6;
// should this set the variable "let x"?
// or define a variable named "x"?
One could potentially design around situations like this, but allowing whitespace in identifiers likely does require being much more meticulous about treatment of reserved words, identifiers, and whitespace than more traditional syntaxes, and this is likely why not many people attempt this.I think the idea is worth experimenting with, though, and that a good implementation of it could be convenient enough for end-users to outweigh the implementation inconvenience.
CL-USER 115 > (let ((first| |number 10)
(second\ number 20))
(+ first\ number |SECOND NUMBER|))
30I think the best way to get identifiers with whitespace to work in a Lisp would be contrive a syntax for S-expressions that uses something other than whitespace to separate things. Perhaps letting (first rest-1 rest-2 ...) be written as as (first: rest 1, rest 2, ...) or (first, rest 1, rest 2, ...), so that example could be written as:
(let: ((first number: 10),
(second number: 20)),
(+: first number, second number))
I imagine it would be possible to write a macro in Common Lisp to transform this into runnable code, or a language in Racket to do so—although, I'm not sure how many people would actually want to make or use something like this.TXR Lisp:
1> (list 1"a"'(b(c)d(e)))
(1 "a" (b (c) d (e)))
Here we just have one space that prevents list 1 from being list1. foo-bar # kebab-case
foo−bar # subtraction
foo minus bar # subtraction? (infix identifier "minus")
foo - bar # subtraction (infix identifier "-")
foo − bar # subtraction (operator symbol)
using \u2212 as a explicit subtraction operator for people who really can't stand having 'extra' whitespace?Small Intro: https://perl6advent.wordpress.com/2015/12/05/day-5-identifie...
this could be solved by allowing strings in qualified imports
import "illegal identifier"."some more weird unicode" as someLib; const thing = @import("relative/path/to/thing.zig");
const package = @import("packagename");`extern crate foo`
`extern crate "foo-bar" as foo_bar`
This is all legal identifiers but
`use std::path::Path;`
`use std::path::Path as int; // For maximum confusion`
And not that anyone uses this part but:
`mod bazz;`
`#[path = "bazz-bar.rs"] mod bazz_bar;`
import A -- compiler reads A.hs
import A.B -- compiler reads A/B.hs
Thus, two semantically related modules are now in different directories.Python, IMO, handles this correctly by having __init__.py support inside directories. It's theoretically less elegant because of the special name, but in practice leads to better file organization.
Same for Rust, but even better, because one can define nested modules in the same file. So you can either define a new module in the same file, put it in a different file named by the module, or put it in the file `mod.rs` inside the directory named by the module.
> import A -- compiler reads A.hs
> import A.B -- compiler reads A/B.hs
That has to happen at some level, assuming subdirectories; otherwise, what would "import A.B.C" refer to?
require( "engine.shared.entities" )
or require "engine.shared.entities"
Which means look for: engine/shared/entities.lua
Additionally, with the `package` module, `require` can also be modified to look for: engine/shared/entities/init.lua`
which is a common Lua practice.Not that my memories of C++ are bad or that I'd avoid using it again! It's just that it would be like trying to reconnect with someone I haven't seen since college. I'm curious, but I don't know if it would be worth the awkwardness.
It's the best time to be a C++ programmer, because if there's something that annoyed you about C++03, there's probably a better way to do it in C++11/14/17.
In a word, Rust is becoming "C++ for 80% cases", but the remaining 20% is inherently difficult (much harder than most people's wildest imaginations - just try to implement a file system library, make it work on last two versions of Windows, Mac and major Linux distributions and you'll understand what it means).
- Go doesn't place requirements to the name of files defining a package [1]. However, it has a preprocessor neither, so the problems described in this article (specifying the module name within an #ifdef) are impossible.
- FreePascal has a preprocessor [2], but it defines a deterministic algorithm [3] to find the files containing a unit (the FPC equivalent of a module). Moreover, the compiler creates two files for every unit: a .o object file, and a "unit description file" [4], much like the C++ proposal.
It seems that FPC's case is the most similar. I think the author is right; the C++ committee should adopt a deterministic way to find the name of the files defining a module.
[1] https://golang.org/ref/spec#Packages
[2] https://www.freepascal.org/docs-html/current/prog/progse4.ht...
[3] https://www.freepascal.org/docs-html-3.0.0/user/usersu7.html
[4] https://www.freepascal.org/docs-html/current/prog/progse13.h...
In Go, the import path and package name are two distinct things: the import path locates and identifies the package, while the package name acts as the default name for scoping qualified name exported from that package.
Furthermore in Go the file names are not part of the import path: the name of the directory containing the files that together define a single package is part of the import path.
An imperfect analogy with C++ would be:
- Go import paths <-> path to the included file (header)
- Go package names <-> a namespace inside that included file.
- Go package filenames <-> sections within the included file
I'm generally disliking the need to maintain separate header and implementation files. Maintaining both is time consuming and putting everything in headers is no panacea, either. Now modules seem to add another type of interface definition to the mix that would need to be maintained after a project adopts modules.
It’s archaic and low level, but it’s also powerful and expressive. Replacing the CPP would probably just require a new language.
Yes it was tighly coupled with Make that took care of selecting the proper set of files to compile and link based on the platform.
The anti-module crowd seems to want to have it all, which to me is the same anti-exceptions and anti-RTTI crowd, and in that case modules are indeed dead-on-arrival.
I have also weitten and deployed what you’re denigrating as “preprocessor spaghetti code” across the above exact platforms, with a lot of success.
Perhaps we work(ed) at the same company.
That being said, if constexpr is great; iterating over either sparse or dense arrays with in Blaze without wrapping in two functions was mind-blowing when I discovered I could so easily.
One of the design goals of the module proposals is that there is a migration path from pure include to pure modules. The transition must not require a flag day.
In particular a program must be able to handle a mix of modularized libraries and old school libraries for the rest of the eternity (it is not likely that C is going to move to modules anytime soon).
If you have the other module as a build step it will compile. If not, you have a linker error.
I don't see the issue here?
Especially if you force module imports to be at the top of the file, you could only scan the start of the file to get an idea on whether you can continue and pause until it's possible or all modules are finished and you didn't get your BMI.
When a new import statement is encountered, Python will first find a source file corresponding to that module, and then look for a pre-compiled version in a deterministic fashion. If the pre-compiled version already exists and is up-to-date, it will be used. If no pre-compiled version exists, the source file will be compiled and the resulting bytecode will be written to disk. The bytecode is then loaded.
The article suggests using this idea in C++ and the parent comment objects, but then it sounds like you're saying it wouldn't be needed anyway (so you're disagreeing with the article too?).
Maybe I'm missing something?
It's not like C++ modules were designed by random nobodies, though; this has been worked on by build infra engineers at major companies with enormous C++ codebases like Facebook, and compiler maintainers e.g. the Clang maintainers. It's possible they completely forgot to think about parallel builds, but that seems at least a little unlikely.
The clang modules proposal had the concept of mapping files, mapping module names to file names.
Companies like Facebook will presumably use proper build systems that already encode the dependency information in the build files rather than try to autodetect it. In that kind of an environment this proposal probably isn't particularly painful.
That seems to leave us with just one conclusion: the article is right, and most of the ecosystem will never migrate to modules, leaving us with the worst of both worlds.
parsing C++ is mostly equivalent to becoming a C++ compiler.
It reallt isn't. Parsing a languagr just means validating its correctness wrt a grammar and in the process extract some information. Parsing something is just the first stage and a one of many stages required to map C++ source code to valid binaries.
template <int>
struct foo {
template <int>
static int bar(int);
};
template <>
struct foo<8> {
static const int bar = 99;
};
constexpr int some_function() { return (int) sizeof(void*); };
Now given the snippet foo<some_function()>::bar<0>(1);
then if some_function() returns something other than 8, we use the primary template and foo<N>::bar<0>(1) is a call to a static template member function.But if some_function() does return 8, we use the specialisation and the foo<8>::bar is an int with value 99; so we ask is 99 less than the expression 0>(1) (aka "false", promoted to the int 0).
That is, there are two entirely different but valid parses depending on whether we are compiling on a 32- or 64-bit system.
Parsing C++ is hard.
EDIT: Godbolt example: https://godbolt.org/z/yR3YHW
Here's an example of a program which compiles only if the constant N is prime, and otherwise emits a syntax error: https://stackoverflow.com/questions/14589346/is-c-context-fr....
In practice, compilers work around this by limiting template instantiation depth.
#if SOME_PREPROCESSOR_JUNK
import foo;
#else
import bar;
#endif
This has legitimate use cases, say #ifdef WIN32
import mymodule.windows
#else
import mymodule.posix
#endif
So in reality build systems will be required to invoke at least the preprocessor to extract dependency information.AFAIK the modules support in the Build2 build system does exactly this, and in fact caches the entire preprocessed file to pass to the compiler proper later.
Having the compiler produce header dependency information is possible, since the dependencies are just an optimization. If there's no dependency information available, you can just compile all of the files in an arbitrary order, and you get both the object file and a dep file. And then on further runs you use the old dep files to skip unnecessary recompilations.
With modules, you can't compile the files in an arbitrary order: if A uses a module defined in B, B must be compiled first. So you need to have the dependency information available up front even for the first build. And since it needs to be available up front, it can't be generated by the compiler. It must either be produced by the build tool which becomes vastly more complicated, or manually by humans.
What??? How would that happen? Are modules always compiled with zero flags because in non-module c++ how the dependent module gets compiled is defined in the build system so in order for the compiler to build a missing .bmi it would have to ask the build system how to build it .
That seems to answer the question. What happens if foo.bmi does not exist? Anwser: you get a compilation error (Missing foo.bmi or foo.bmi out of date). You then need to go fix the dependencies in your build system to make sure foo.cpp gets compiled before bar.cpp.
Right?
I get that might suck but it's not unprecedented. lots of builds have dependent steps. Maybe in order to implement C++ modules build systems will need an eaiser way to declare lots of dependencies where as now dependencies are an exception?
The build system just figures out how to invoke the compiler. The compiler does the actual building. When the compiler runs, it has all the flags.
Remember, headers in C / C++ are basically a file level construct. They happen before you even split the file into tokens. #include just means "do the equivalent of opening that file in a text editor and copy and paste it in place of this #include line."
The compiler is already compiling header files, as part of compiling cpp files.
modules work with the import statement (new) not the #include statement. They are not the same as include at all.
In fact this is spelled out in the article in the first goal
> The “importer” of a module cannot affect the content of the module being imported. The state of the compiler (preprocessor) in the importing source has no bearing on the processing of the imported code.
In other words, the flags passed in when compiling bar.cpp have no effect on foo.bmi. foo.bmi is the result of the flags passed in when foo.cpp was compiled and those flags can only be gotten from the build system if foo.bmi does not exist.
Sometimes C++ can use its age as an excuse to be super complicated but here, the modules implementation of C++ is younger than Rust's.
If I'm understanding the post correctly, the entire problem they are facing is that you have to scan all source files to build this module->filename mapping.
None of the "essential goals" listed at the top of the blog post requires that modules be imported by namespace instead of filename, as far as I can see. So why was this design chosen when it causes these problems?
Since modules operate at the language level, they need to operate on this notion, which precludes importing by file.
Also, going to this level of trouble to support systems that don't have files seems... odd. Targets that don't have files, that I can totally understand. But compiler toolchains that don't have the notion of a file? That sounds obscure beyond obscure. I'm surprised such a system would be a compilation host instead of a cross-compile target.
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2014/n421...
int main(int argc, char *argv<::>)
<%
if (argc not_eq 1 and argc not_eq 2) <% return 1; %>
return 0;
%>
https://en.wikipedia.org/wiki/Digraphs_and_trigraphs#CTo me this seems like a weird take on accessibility. In order to accommodate that one OS that has some serious disabilities, everyone else has to suffer the consequences. Why not build a ramp for that one OS, and build stairs for everyone else?
Still trigraphs were removed in the end; if there is enough support the committee is willing to break backward compatibility.
Admittedly that's not just a problem with old mainframes. Any system supporting file aliases (be it hardlinks, symlinks or the same FS mounted at several locations for instance) would be tricky to handle.
I always thought #pragma once was a bad idea for that reason, header guards with unique IDs don't require any compiler magic and are simple to reason about without having to read the standard or compiler's docs to figure out how it operates.
OCaml or Haskell do just fine with proper modules and parallel builds. I presume the same holds for Go, rust, et. al.
Clang’s module maps solve a lot of these problems. There is a known location to find the module map and all headers from the module must be reachable from the map.
Even at Apple, which originally designed them, they just focused on making it good enough for C and Objective-C system headers.
https://www.youtube.com/watch?v=K_fTl_hIEGY&feature=youtu.be...
I agree, but that's because C++ TUs are too small. Rust works like this "dead-on-arrival" way, but survives because TUs are larger.
For example,according to the clang documentation for modules [1], they are meant to obviate the necessity of #including headers of libraries you are linking to.
When you link with libraries, you expect them to be compiled already. If the library is a dependency of your current project (in Visual Studio solution parlance), then it will be recompiled if needed before your project is compiled, if your build system is set up correctly. You will also take care to point your build system to the correct version of the library you want to link to, including versions that change depending on CPP flags etc.
I don't see how modules are any different.
Please enlighten me if I'm not fully grasping the point of the article.
Header, source like before. Allow putting all implementations in the source (including templates). Allow private member functions to be defined in the source file even though they were not declared in the class. Every function that is not declared in the header is no-external-linkage (including these adhoc private members). Defines don't leak into header file from including files unless explicitly passed into it (using some new syntax). Defines can leak out of header files. If the define already exists it's an error (the point of this is to make sure that order doesn't matter), unless the origin of the define is the same (so if x includes y and z, and y includes z, then a define in z would go into x directly and also through y which would not be an error).
This seems like it should be a strict improvement.
Do you realize how hard this is to implement? This actually was part of the original C++ spec (export keyword), but no compiler successfully implemented it in a way that was compatible with other compilers, and it was deprecated.
This (and C++ modules) requires a standard intermediate representation of the language that all compilers share, otherwise compiler A can't use a module generated by compiler B.
This is why I've never really expected C++ modules to ever exist, or if they do, it will be in a form that is much more limited than most people want or expect. Either they'll only allow a subset of the language to exist in a module, or the feature won't be much different than the "pre-compiled header" feature offered by most C++ compilers), or modules won't be portable across compilers (and maybe even versions of the same compiler).
Currently every compiler except MSVC plans to make module files version locked. Different versions of the same compiler will use different module files.
Personally I find little value in portable module files, as you need to be able to rebuild modules anyway to handle pretty much any change to compile flags.
But I have to wonder, is adding one more big feature wise? Will modules be the feature where it becomes impossible to create a compiler that's both useful amd conforming?
It seems to me that the module-interface unit (MIU) would be pretty much the equivalent of present-day header files. So for two modules foo and bar, there is no dependency on the order foo.cpp and bar.cpp are compiled, because only the MIU needs to be compiled for a given module and the design ensures two MIU are isolated. If they are mutually dependent it's a bad design, but teh solution is the same as in Java: you need to build the MIU in the same compiler invocation. (In fact, that would probably happen automatically, using the equivalent of today's -I include directory directive to find MIUs.)
Yes, that means you need to split your module into a clean MIU andthe actual implementation file, just like now you split them in a header file and a cpp file.
Yes, you need MIU to be "available" to be translated to BMI gobally, just like you need header files to be available globally when compiling.
My immediate though was create a SQLLite database per some collection of modules (a crate of modules?).
Each row in that table will have unique 'business id' and unique row_id. business Id is a composite key on
module_name+exported_function_names_with_attribute_and_result signature
every BusinesID as defined above, can point to source location, compiled object location, last compilation time, compilation state.
Moving crates to other locations will not change the BusinessID. Moving compiled object location will not change the BusinessID.
The sqllite database can have a network interface for multi-machine compilation. can be backed up/restored into a new environment and be unchanged as long as compilation flags for all the modules remained the same, and compiler versions are same.
C++ does not. The build system will also take care of this problem, just as we currently e.g. define libraries in cmake.
I didn't get why - one could imagine parallelizing BMI generation, then parallelizing "normal" compilation.
The only issue I see is that you wouldn't know which BMI you'd need, so you would need to generate all of them (or regenerate those who are out of date), or specifically list those that needs to be generated in a build tool. Given how the rest of the compilation pipeline works, is that undesirable?
I'll keep my irrelevant eyes to myself then, thanks.
It is a binary module interface not binary module implementation.
Well before reaching 300 I'd have to ask WTF are you doing? I mean really, seriously, if your dependency chain is 1/10 that deep I'd look for something wrong.
I've often thought that including headers from inside headers in C or C++ is a mistake. And that thinking is probably wrong. It makes sense when using a library that may itself have a lot of components - I just include the top-level header for the library. But even that is different from having a really deep dependency chain.
Maybe - just maybe - people have shifted the spaghetti out of their code and into the file structure.
class Foo {
};
Then if I must write bar.hpp as: #include "foo.hpp"
class Bar {
Foo f;
}
I cannot forward declare Foo because in order to size Bar, I must know the size of Foo. I must therefore include `foo.hpp` in `bar.hpp`, thus I must include headers in headers, unless my headers are not allowed to contain class definitions.This is the type of complexity that a good software "architect" should be trying to reduce rather than manage.
e.g. a blanket rule that a header file isn't allowed to include another header file is trivial to enforce, one which says it can't be more than n deep, is subject to boundary pushing.
Are there other languages _like Rust_? Have you heard of large, production code bases maintained at large successful companies that are written in Rust?
Google, Amazon, Facebook, Microsoft, Dropbox...