A port identifies a process, an IP address identifies a machine. The machine receives the packet, and then passes it to the process, since it's the machine that needs to receive the packet before routing it to the process, it's the kernel with ring 0 privileges that needs to process the packet and map the port to the process.
I don't think it's a good idea to put more of the stack than necessary into each individual program though, because then you're reliant on each program implementing everything correctly and you have to worry about bugs and security issues in every single program you're running rather than just in the kernel.
Considering that even something as simple as opening a socket (which involves calling one function to give you three numbers which you pass unchanged to a second function) is largely a clusterfuck in existing programs, I don't trust them to implement the whole TCP stack.
I was thinking more along the lines of network namespaces -- although actually, I should have been thinking of TUN interfaces. Any packets sent into one of those are delivered to the attached program, which can do what it likes with them. Those are usually used by e.g. OpenVPN for L3 things, but nothing stops a program from handling the L4 headers.