That is a common misconception. Or rather it is a tautalogy. No threads = thread-safe. But, it is not concurrently-modifying-data-structures safe -- which is the main painful point.
One can get just as easily tangled over a set of callbacks.
Here is a set of callbacks all started from some select/poll/epoll loop. Some call it a reactor (namely Glyph's own Twisted Python).
cb1 -> cb2 -> cb3|eb3 then cb3->cb4 and eb3->cb5
Processing starts with cb1 and ends with cb4 or cb5. Notice at some point cb2 function could result in generating an errback (eb3) which then ends up calling another callback cb5.
Understanding that the above, in a large system is just a messier, uglier concurrency structure than a thread/goroutine/task/actor is crucial.
It doesn't necessarily save you from simultaneous access to same shared data.
Imagine processing starts at cb1 and by the time it reaches cb3 (say cb2 calls some io or sleep operation), cb1 gets called again. cb1 through cb3 end up modifying some shared data (hey no need for lock, we are using callbacks remember!). Now there are two callback chains modifying shared data.
Yes you need locks and semaphores with the above just as you do with threads
For example this exists -- Twisted's own Sempahore:
http://twistedmatrix.com/documents/10.1.0/api/twisted.intern...
I had to use it, and not just for throttling concurrency, but also to protect critical data from being modified concurrently.
Asynchronous/callback/promise/future based concurrency looks really good in small examples. In large application they get messy quickly.
Threads/actors/goroutines etc are still nicer from a logical, application point of view. You can even build them on top of the same epoll/select/kqueue system calls if the language can support some kind of a coroutine structure (which for example python gevent/eventlet) is doing.