Such a test will favor implementations that have extremely svelte boostrapping, allowing them to immediately begin executing the relevant code and return a result.
I feel a more useful test would be for the relevant string processing code to be run tens to thousands of times within the process itself, so as to diminish the relative importance of boostrap/warmup code and/or runtime optimization performed by VMs. Unless, of course, part of the intent was specifically to measure the impact of the bootstrapping.
IMHO, VM startup time should be included and the first few passes (before JIT kicks in) should also be included.
Java has supposedly “nearly C-like performance” until you read the fine print.
(This should apply to C# as well.)
Is that just your assumption or have you measured that penalty for this tiny tiny program?
Might "the JVM's slow start time" in this case be insignificant?
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
But why does Swift have a long startup time in the first place? Shouldn’t it start near instantly like the C and Rust programs?
echo 'print("hello world")' > hello.swift && swiftc hello.swift -O -o hello
time ./hello
[1] https://www.swift.org/server/But I don't think it would be fair: I use Python a lot, and if it's for scripting, you are happy with the fact it's very easy to write. However, it's slow to start, and you pay that each time you run the script.
To me, it makes sense in this exercice, which is heavily leaning toward scripting, we see the price of the VM start in the overall profiling.
Otherwise, let's use pypy, warm it up, a few 1000 times, and you may get closer to Go times.
But we don't use pypy for scripting.
1.33 s 56.7% specialized Collection<>.split(separator:maxSplits:omittingEmptySubsequences:)
263.00 ms 11.2% Substring.lowercased()
226.00 ms 9.6% specialized Dictionary.subscript.modify
189.00 ms 8.0% readLine(strippingNewline:)
As you can see, calling split on the string is really slow. The reason it is slow is that it's using the generic implementation from Collection, rather than String, which doesn't really know anything about how String works. So to do the split it's doing a linear march down the string calling formIndex(after:), then subscripting to see if there's a space character there using Unicode-aware string comparison on that one Character.Swift is really nice that it gives you "default" implementations of things for free if you conform to the right things in the protocol hierarchy. But, and this is kind of unfortunately a common bottleneck, if you "know" more you should really specialize the implementation to use a more optimized path that can use all the context that is available. For example, in this case a "split" implementation should really do its own substring matching and splitting rather than having Collection call it through its slow sequential indexing API. Probably something for the stdlib to look at, I guess.
The rest of the things are also not too unfamiliar but I'll go over them one by one; lowercased() is slow because it creates a new Substring and then a new String, plus Unicode stuff. Dictionary accesses are slow because Swift uses a secure but not very performant hash function. readLine is slow because the lines are short and the call to getline is not particularly optimized on macOS, reallocs and locks internally, then the result is used to create a new String with the newline sliced off.
It's sad to see that Swift is so slow in this particular case given that it has the potential to be such an optimized language (strong typing, compilation, backing from Apple).