Translating a C++ parser to Haskell
haskellforall.com
haskellforall.com
E: As some quick examples, peeking an istream in a loop is extremely inefficient[1], yet is done pervasively. The Derivation type has a bunch of separately allocated strings, and a lot of types like std::map<string, DerivationOutput> are pointer filled. To check whether a string starts with "r:", they allocate a _new string_ and compare it, rather than just doing a normal string comparison.
https://github.com/gcc-mirror/gcc/blob/1cb6c2eb3b8361d850be8...
which itself redirects to a whole bunch of stuff.
The virtual call alone is a lot more expensive than a simple pointer dereference, and though it's possible proper inlining might have alleviated some of this problem, you're not particularly likely to see it come close to something that's efficient by design.
[0] https://stackoverflow.com/questions/4340396/does-the-c-stand...
Nope; people use C++ for all the times ... when what they know best is C++. Which is all the time for a lot of C++ developers, I suspect.
Then, for a whole bunch of other times not covered by that case, they choose C++ because of tooling and surrounding code into which that code must integrate. If you need a parser, and it has to integrate into a C++ program, you probably choose C++ for that reason, and then write the parser in whatever way you are able while trying to keep it understandable and maintainable, which may not be the most efficient approach. You do that even if you can write the parser in six other languages.
The idea that C++ is always selected for speed and used accordingly rings fallacious to me. It's not even the way it should be; choosing the ideal language (and ideal way of using that language) for every little subtask in a large project is going to be harmfully counterproductive. Unless the language count can be kept to no more than two or three.
What I meant is more that the initial choice of C++ tends to come from a requirement somewhere in the program for the low-level control that C++ needs, rather than the whole thing, and making an argument that Haskell is performant enough to replace C++ only makes sense if it's either fast enough for that case too, or you're willing to have a heterogenous codebase and you don't really care about Haskell's performance anyway.
Even that doesn't always come from such a requirement. Maybe it's just, "the common denominator of expertise on this team is C++ so we will use that".
That original requirement for low-level control is now 5% of the code base, 95% of which is still C++ because of the effect that it's easier to just add more C++ than to balkanize the project with new languages. The parser adds 1%, and scans a config file only about every 100 days days when the C++ server is restarted; why would it have to be fast ...
I am old enough to have heard the same complaint, but with Assembly, C and Turbo Pascals as protagonists.
One should program in idiomatic productive code and only if the profiler shows there is a bottleneck, that is actually relevant for the use case at hand, one should bother to spend time optimizing.
To add my voice to the other comments, there are many reasons we get to use C++, not only for bleeding edge performance.
Portable code across mobile OSes, access to the JVM and CLR monitoring APIs, UWP APIs not exposed to .NET, medical devices DLL and COM libraries are just a few examples that come to my mind.
Consider a case where somebody is using C++ for mobile OS portability or to hook into another large project, rather than the performance it offers. In that case, the argument that the parser would be equally fast in Haskell is irrelevant because that was never the use-case.
Consider a case where somebody is using C++ for performance. In that case, the argument that the parser would be equally fast in Haskell is irrelevant because that's not the part that needs to be fast.
The only time "X in Haskell is equally fast to X in C++" is an argument that you care about is when you care about the performance of X, and in that case the code wouldn't look like this.
> Note that Haskell type synonyms reverse the order of the types compared to C++.
This is true in the context of the article when compared with the presented typedef definitions, however the modern way to do type synonyms in C++ is via using declarations, which are very similar to the Haskell ones.
using Path = string;
using PathSet = set<Path>;I imagine a lot of people are in that place with C++. Personally I've followed C++ for a long time as a hobby and it seems like it's finally getting to where it wants to be. But I'll talk to someone who works writing C++ and the idea of being able to use C++17 is a distant dream.
For Open Source developers and hobbyists though, you might as well take advantage of whatever features Clang+GCC+MSVC offer. I wish someone would standardize `#pragma once` so I don't have to feel guilty using it.
Modern c++ is awesome. I definitely appreciate being able use c++14 where I work (as well as 'pragma once'). I really look forward to being able to use c++17.
It is called C++ Modules Technical Specification. :)
How do the 2 parsers compare w.r.t error messaging?
Haskell has Parsec/Megaparsec which have better error messages, but are extremely slow.
Trifecta has the best error messages, but takes some more work to fully make the error messages great.
Took around 20 seconds to parse a few hundred thousand lines (almost no backtracking, too).
Attoparsec/parsec have very similar APIs. I wonder why it cannot parse optimistically with attoparsec, and in case of failure - re-parse with parsec to get a nice error message. This should yield the best of both worlds?
[1]: https://github.com/merijn/lambda-except [2]: https://github.com/quchen/stgi/blob/master/src/Stg/Parser/Pa...
For this specific post, the main thing I do is ignore the error messages and just look at where parsing fails by retrieving the leftovers when it fails (using the lower-level `parse` function). Usually that gave me enough of a clue to figure out what was going wrong.
However, I may be able to fix the zooming. This was triggered by the switch to syntax highlighting using Pandoc which wraps each code block in a scrollable frame if it exceeds the given size. I can probably fix that with appropriate CSS although it will take me some time to figure it out.
To propose another hack, I've found that 'Request Desktop Site' in your mobile browser's menu works well for articles like these (it just reverts to desktop layout).