This isn't a parallelized implementation, no. To have that, imagine that instead of a size_t for ret, that it was something like size_t[4], and that each operation you did acted on all 4 items at once (this is how SIMD operations work, they do the same thing to each item in a vector of values).
You would end up returning a size_t[4], which would have the results of your 4 searches in it, which ran in parallel.
Regarding the ternary operator, you can always check your assembly (here's some that shows it https://godbolt.org/g/BM6tfc), but the article also shows how to do it with a multiply instead of the ternary operator.
The biggest thing sticking out of this as being slow would be the cache unfriendliness of binary searching, but the link at the end shows how to clear that up (which is not the point of the article, but the ideas are compatible and combinable).