1. On multicore machines you want to process timers in parallel on multiple cores. With userspace timers you either set the same timeouts on all threads and have unnecessary wakeups or distribute timers to cores ahead of time which leads to increased latency if a thread is stalled for any reason. I think this is unfixable without a dedicated timer API.
2. Good timer APIs let you set a time _interval_ for when the timer expires, which is essential so that the system can group timers and reduce wakeups (i.e. you process all timers where the lower bound has been reached before going to sleep, but don't wake up until the upper bound arrives). Most or all "wait with timeout" APIs only have a single timeout, although this could be fixed.