HNHacker News
TopNewBestAskShowJobs

def-pri-pub

214 karma · joined January 3, 2017

C++/Qt Developer. Personal website: https://16bpp.net
submissionscomments
def-pri-pub··on Even Faster Asin() Was Staring Right at Me
>> It also gets in the way of elegance and truth.

Where did that come from in the article?

def-pri-pub··on Even faster asin() was staring right at me
Thank you for linking that!
def-pri-pub··on Faster asin() was hiding in plain sight
The reason for writing out all of the x multiplications like that is that I was hoping the compiler detect such a pattern perform an optimization for me. Mat Godbolt's "Advent of Compiler Optimizations" series mentions some of these cases where the compiler can do more auto-optimizations for the developer.
def-pri-pub··on Faster asin() was hiding in plain sight
I did see that, but isn't the vast majority of that page talking about acos() instead?
def-pri-pub··on Faster asin() was hiding in plain sight
I did scan some (major) open source games and graphics related project and found a few of them using `std::asin()`. I plan on submitting some patches.
def-pri-pub··on Faster asin() was hiding in plain sight
Thanks!
def-pri-pub··on Faster asin() was hiding in plain sight
When I was working on this project, I was trying to restrict myself to the architecture of the original Ray Tracing in One Weekend book series. I am aware that things are not as SIMD friendly and that becomes a major bottle neck. While I am confident that an architectural change could yield a massive performance boost, it's something I don't want to spend my time on.

I think it's also more fun sometimes to take existing systems and to try to optimize them given whatever constraints exist. I've had to do that a lot in my day job already.

def-pri-pub··on Faster asin() was hiding in plain sight
You'd be surprised how it actually is worth the effort, even just a 1% improvement. If you have the time, this is a great talk to listen to: https://www.youtube.com/watch?v=kPR8h4-qZdk

For a little toy ray tracer, it is pretty measly. But for a larger corporation (with a professional project) a 4% speed improvement can mean MASSIVE cost savings.

Some of these tiny improvements can also have a cascading effect. Imagining finding a +4%, a +2% somewhere else, +3% in neighboring code, and a bunch of +1%s here and there. Eventually you'll have built up something that is 15-20% faster. Down the road you'll come across those optimizations which can yield the big results too (e.g. the +25%).

def-pri-pub··on Faster asin() was hiding in plain sight
Funny enough that fdlimb implementation of asin() did come up in my research. I believe it might have been more performant in the past. But taking a quick scan of `e_asin.c`, I see it doing something similar to the Cg asin() implementation (but with more terms and more multiplications, which my guess is that it's slower). I think I see it also taking more branches (which could also lead to more of a slowdown).
def-pri-pub··on Faster asin() was hiding in plain sight
These are books that my uni courses never had me read. I'm a little shocked at times at how my degree program skimped on some of the more famous texts.
def-pri-pub··on Faster asin() was hiding in plain sight
Wait, what? Do you have a resource I could read up on about that? That is moderately concerning if your math isn't portable across chips.
def-pri-pub··on Faster asin() was hiding in plain sight
> floating point bithacks

The forbidden magic