You're mostly right in that it is similar to the problem of buffer bloat in networking. However, it's not quite the same thing because for example you can do things like write a file to disk & unlink it or overwrite some portion of contents, meaning that by buffering for longer you can avoid the writeback in the first place. By buffering for longer, the kernel is trying to balance things landing on disk and avoiding touching the disk if the dirtied data will be dirtied again. Granted not a common use-case these days, but consider the case where you have object files being created during a build over & over again. It's easy to construct scenarios where you indeed wouldn't want to write to disk so quickly.
It's not out of hand a terrible idea to avoid flushing data to disk and there's no free lunch here as any workload you optimize for will have a different workload that suffers. People try to come up with general heuristics that work in most situations on consumer machines, but there's no one size fits all for all HW + use-case combos. That's why hyperscalars tune the kernel beyond that / have kernel developers writing code to optimize for their use-case. It's telling that the performance analysis in the article is pretty hand-wavy without any clear demonstration of a concrete problem.
As for sync, I believe the author is mistaken. You can do fsync instead which is more efficient as it only creates a barrier for writeback of the file descriptor rather than a system-wide sync. And invoking fsync I believe is more common than sync. You should be able to have multiple parallel fsync happening concurrently for unrelated files that don't block on each other so much (ideally the kernel would prioritize those writebacks and interleave for fairness, but I doubt it does).