Can C++ become your new scripting language?
nu42.com
nu42.com
C / TCC was great, but TCC doesn't attach much importance to final speed optimization. I had a working prototype but ended up scrubbing it for that reason.
C++ / Clang / LLVM has the best end speed and optimizer, and as the article points out, the language is looking better and better. However, as a library it's pretty massive, difficult to embed, and compilation would probably be too slow for a REPL/JIT type of situation, though I haven't tested this with a working prototype.
Lua / LuaJIT is what the project is currently using. Since this is HN, I guess I don't need to say anything about how fast it is. However, my application (and many others, I suspect) would get a large performance improvement (I estimate 3x) if the compiler was capable of working with float32 operations and of optimizing them into vectorized SIMD instructions. This is why I'm currently looking at the next possibility:
Javascript. With the recent introduction of different float widths and SIMD in major Javascript JIT compilers, this option is starting to be the fastest possible (in my case). I'm planning to make a prototype to verify this, and I'm not too keen on joining the JS bandwagon, but if LuaJIT development continues to be basically halted, I'll have to get with the times. What I need is at the bottom of an epic TODO list for LuaJIT [2] that hasn't really moved for years. I don't know who's in a position to make a move on those LuaJIT open sponsorships, I really wish they would do it! I for one don't really feel up to the task.
If you are using Luajit you can game the typesystem, by using the FFI interface, making math much more performant by using the FFI C interface. I think i would go through that route first
And ideally, you shouldn't need to always specify vector operations, the compiler should optimize your loops (for example) into SIMD instructions. But as far as I know, very few compilers are capable of doing this now. I think some C/C++ and Java compilers are among the only ones, though the example given here suggests that Spidermonkey should also do it to some extent:
This is what is being used now: https://root.cern.ch/drupal/content/cling.
But if I remember correctly, then CINT wasn't actually developed by the root-team, but rather already existed when work on root started.
[1] Read: complicated.
For example, this Python is both better behaved and much, much simpler:
import sys
for filename in sys.argv[1:]:
try:
with open(filename, "rb") as file:
word_count = sum(len(line.split()) for line in file)
except IOError as e:
print(e)
else:
print("{}: {}".format(filename, word_count))
I'd suggest you think up an example that justifies its implementation before making such a comparison. If the aim is to make both implementations do the same dance, not just to make them give the same results, even Haskell won't look much different to C++. #!/usr/bin/env perl6
sub MAIN(*@filename) {
for @filename -> $file {
say "$file: { +$file.IO.slurp.words } words";
CATCH { default { say "Unable to count words in $file"; } }
}
} #!/usr/bin/env perl6
sub MAIN(*@filename) {
for @filename -> $file {
my $bag = $file.IO.slurp.words.Bag;
say "$file: { [+] $bag.values } words";
CATCH { default { say "Unable to count words in $file"; } }
}
}ref: http://www.reddit.com/r/cpp/comments/369lcn/can_c_become_you...
First, this can be done with one line in Python.
Second, it's almost a page of C++.
Third: if done in a 10 lines in python it is around 30% faster than the C++ version and 2x faster than the perl version.
Here is my reference Python version that also counts the words in the same non-code golf kind of way:
import sys
from collections import defaultdict
def wc(path):
d = defaultdict(int)
with open(path) as f:
for line in f:
for w in line.split():
d[w] += 1
return sum(d.values())
if __name__=='__main__':
path = sys.argv[1]
print(path, ':', wc(path))
EDIT: They are counting new lines correctly. for my $file ( @ARGV ) {
my $words;
open my $fh, '<', $file or die;
local $/;
say "$file: ", scalar(split /\s+/, <$fh>), " words";
}
EDIT: Changed an explicit count of the `split` list into an implicit count via `scalar`.* You need to loop over sys.argv[1:].
* You could use binary mode to avoid decoding (which the C++ avoids).
* You could use collections.Counter() and its `update` method, which should be both cleaner and faster on Python 3:
d = Counter()
with open(path) as f:
for line in f:
d.update(line.split()) +$path.IO.lines.wordsI'm wrong, though; the optimizations Counter applies only make it faster when each `update` takes an iterable of a large number of elements. This is unlikely to happen in our case.
You can actually bypass the Counter wrapper and use its accelerator directly:
from _collections import _count_elements
def wc(path):
d = {}
with open(path, "rb") as f:
for line in f:
_count_elements(d, line.split())
return sum(d.values())
However, this uses implementation details and is as such bad code.Nevertheless, defaultdict == win for counting things like if we actually wanted unique words in this case.
int word_count(const char *const filename)
{
std::ifstream file{filename};
return std::distance(std::istream_iterator<std::string>{file}, {});
}C++ is improving, which is to be applauded. But all of these examples seem to be "look, C++ used to be a lot worse than Python, now it's only a little bit worse than Python". Show me the USP, the compelling use case where only C++ will do.
I think that was the case here.
Also, you can simulate large parts of the program in IPython and not only find what the problem/bug is, but solve it right there and copy the solution back to the original program.
int sum = 0;
for(auto&& wc : word_count) sum += wc.second
return sum;
It's shorter, more general, more familiar, possibly has better performance[[citation needed]], and gives better errors if you mess something up.Or at least it would in an ecosystem where functional programming is embraced, but I don't know how practical it is in practice. Its not composable if you can't expect any other given piece of code to be written in a way conducive to composing with accumulate.
In other languages, though, it's pretty neat. For example, in Rust one just does
data.iter().fold(0, |s, e| s + e)I used to write scripts in Ruby since I liked the syntax, it was very easy, nowdays I have tried javascript since I use it anyway in web projects and Node/NPM has a lot of modules ready. One thing I like with scripting languages is that I usually need to tweak the script a bit over time and that is very easy to do with scripting languages, even on different computers which may or may not have a compiler.
Ah yes, auto! I've been using that a lot since I saw Herb Sutter's talks.