they are already using the same compiler infrastructure: pretty much every relevant shader or otherwise GPU oriented compiler out there is based on LLVM. As for the parts of the toolchains that are different: that's simply because GPUs and CPUs are inherently very different.
> so that CPU and GPU code can be easily mixed in the same source file
what advantage would it have to mix together into the same source file the code for totally different targets with totally different parallelism models? I find it cleaner and simpler (apart from maybe trivial one-liner shaders that you can embed in a string) to keep them separated. Also makes debugging much easier.