When writing GPU code I often optimize/ tailor the algorithm to the details of the architecture. Trying to write generic GPU code that is performant across multiple platforms is ugly, painful and produces larger and less maintainable code. One actually chooses different algorithms based on the architecture.