It's not out of hand a terrible idea to avoid flushing data to disk and there's no free lunch here as any workload you optimize for will have a different workload that suffers. People try to come up with general heuristics that work in most situations on consumer machines, but there's no one size fits all for all HW + use-case combos. That's why hyperscalars tune the kernel beyond that / have kernel developers writing code to optimize for their use-case. It's telling that the performance analysis in the article is pretty hand-wavy without any clear demonstration of a concrete problem.
As for sync, I believe the author is mistaken. You can do fsync instead which is more efficient as it only creates a barrier for writeback of the file descriptor rather than a system-wide sync. And invoking fsync I believe is more common than sync. You should be able to have multiple parallel fsync happening concurrently for unrelated files that don't block on each other so much (ideally the kernel would prioritize those writebacks and interleave for fairness, but I doubt it does).