:-)
:-)
Having said that, compilers do pretty badly on C-to-simd optimisation. The best you get is loop vectorization if it's really simple logic. You can usually get some pretty good wins there. The fact you lay out your memory for simd usually is a win all of its own due to cache prefetching even if you don't actually use any simd instructions. Compilers need heuristics to manage cache when you know what you're trying to do, (eg when should it use non-temporal writes, for example?) Fast C code is written while having a really clear mental model of the underlying architecture and the assembly that the C will produce with -O3 (or whatever flag is relevant to your compiler) and then checked with -S or objdump -D, profiled with callgrind/cachegrind, perf, rdtsc etc...
The compiler really can't "Do it for you" You /can/ use a compiler as one of your tools when /you/ do it. As Randy Hyde points out you can always beat the compiler because you can use its generated assembly language in every case you can't beat, so the absolute worst you get is a tie.
So yeah, you can totally smoke clang, Intel, microsoft and gnu C compiler and get paid something for doing it in certain industries too. :-)
Mike Acton being aggressively opinionated on the subject, but the lecture is really good (despite/because of) the bits you'll disagree with and the manner he'll rub you the wrong way. https://www.youtube.com/watch?v=rX0ItVEVjHc
Requests/sec: 100.00 Transfer/sec: 11.91KB
for a JPEG image (200KB) the results are similar:
Requests/sec: 99.86 Transfer/sec: 19.27MB
[trent@ubuntu/ttypts/4(~s/wrk)%] ./wrk -c 1 -t 1 --latency -d 5 http://localhost:8080/Makefile
Running 5s test @ http://localhost:8080/Makefile
1 threads and 1 connections
Thread Stats Avg Stdev Max +/- Stdev
Latency 35.00us 0.00us 35.00us 100.00%
Req/Sec 10.00 0.00 10.00 100.00%
Latency Distribution
50% 35.00us
75% 35.00us
90% 35.00us
99% 35.00us
1 requests in 5.10s, 1.57KB read
Requests/sec: 0.20
Transfer/sec: 314.55B
Note the 1 request. For 1000 clients, it's only doing 1000 requests: [trent@ubuntu/ttypts/4(~s/wrk)%] ./wrk -c 1000 -t 1 --latency -d 5 http://localhost:8080/Makefile
Running 5s test @ http://localhost:8080/Makefile
1 threads and 1000 connections
Thread Stats Avg Stdev Max +/- Stdev
Latency 307.93ms 552.88ms 1.63s 87.10%
Req/Sec 414.29 439.07 1.34k 85.71%
Latency Distribution
50% 4.11ms
75% 407.63ms
90% 1.63s
99% 1.63s
1000 requests in 5.01s, 1.53MB read
Socket errors: connect 0, read 41, write 0, timeout 0
Requests/sec: 199.79
Transfer/sec: 312.96KB