For a plain-text benchmark, you are ultimately measuring the ability of sitting on top of a socket and reading/writing to that socket as fast as you can.
However, Phoenix/Cowboy, whenever there is an HTTP connection, spawns a VM light-weight process to own that connection, and then each individual request runs in its own light-weight process too. This comes with its own guarantees in terms of state isolation, reasoning about failures, and you get both I/O and CPU concurrency for free - if you were to do any meaningful work on those requests.
Even though the VM processes are cheap to spawn and lightweight, it is overhead compared to something that is just directly reading and writing to the socket. I actually wrote a proof of concept where we just sit on top of the socket without spawning processes and it performs quite well - although I don't think it has any practical purpose. There is also an interesting article from [StressGrid](https://stressgrid.com/blog/cowboy_performance_part_2/) that shows how removing the per-request process speeds it up by ~60%. But once you are going through long requests (i.e. 1ms because you need to talk to a database or API, encode JSON, etc), these differences tend to matter less and the concurrency model starts to give you an upper hand.