I like the thread per core design because kernel context switches are expensive. But userspace context switching scheduling (such as an unbuffered golang channel) for a 64 bits of data at a time makes me uncomfortable too unless it represents a large amount of work backed by a pointer to work data.
I think I like to multiplex sockets over threads and multiplex IO over threads so that you can do CPU work while IO is going on and you can process multiple clients per thread.
You cannot scale memory mutation by adding threads. You want ideally one thread to own the data and be safe to mutate it uncontendedly.
When you send data to another thread, don't refer to that data again. Transfer ownership to that thread.
If you can divide your request into phases of expansion (map) and contraction (synchronization) you can do intrarequest parallisation.
I've been looking into runqueues of go and tokio where you have a local runqueue without a mutex and a global runqueue with a mutex.
I've been trying to think how IO threads that run liburing or epoll can wake up a Coroutine or an async task on a worker thread without the mutex.
The worker thread is looking for tasks to resume that are unblocked and needs to be notified when there is IO finished. I think you can have a Coroutine that is always runnable to read from a lock free ringbuffer. You can have that Coroutine that checks for finished IO's ringbuffer yield if its contended by writes by the IO thread that is trying to make work available to the worker threads.