IBM was doing something like this in their software load balancer ~1996. The technique was used in the large sporting events websites of the day (olympics, wimbledon, ...)
went something like this:
- a routing node (a kernel extension) would receive packets and forward them (verbatim, same src & dest IP) to a backend worker
- backend worker was configured for same IP, which meant:
a) arp had to be disabled to avoid conflict, or
b) the loopback interface had to be aliased to same IP (using loopback avoids arp conflict and worker still happy to accept the packet and respond)
- response then goes back directly from worker to original source, which was a good approach for scaling web traffic back then, small http requests going through routing node, large responses going back direct from worker to source.
more details:
https://patents.google.com/patent/EP0838931A2/en