Celery boils down to a nice abstraction over forking another process to do your work. You still have the scalability problems that might make you want to write concurrent single-threaded code, just at a different layer of your stack.
Take the author's example of scraping 1000 webpages. Let's say their computer running celery has enough memory to run 50 celery processes. This means they can request 50 web pages concurrently. Most of the time those celery processes are going to be sitting idle, waiting on the remote web host to respond. This is terribly inefficient compared to coroutines with asyncio/aiohttp, or even using a thread pool since urllib will release the GIL and let another thread run when it's blocked on network i/o. You could have a single process performing the work just as fast as 50 celery processes.