If you run an actual binary, the OS will not only cache the code pages, it will share them across processes. In the perl case, the actual interpreter is loaded, cached, and shared across the CGI invocations, while the actual script code will be cached from the disk, but each invocation will (likely) get their own copy of the code (i.e. they will duplicate it from cache when they read it in).
Finally, of course, the runtime needs to start up again, compile/interpret the code, and, finally, execute it.
A binary will have the fastest startup. I'm sure there may be things that could be done. I don't know if perl mmap'd the .pl file as a readonly buffer, if each invocation would share that file (probably), and thus prevent that initial duplication. Don't know if any of the interpreters do this. Odds are this gain is lost in the whole startup of the interpreter anyways.
Also, using a binary, by a similar mechanic, lowers overall memory consumption. Have 100 connections all running the binary, they all share the code pages, and thus only "pay" for their individual data use. Share 100 connections with an interpreter, and they each pay their own cost for the entire program file plus any runtime data. Again, not so much a problem today with our ample memory resources, and mmap the source may remedy this, but this actually was an issue back in the day.