I'm just guessing it makes more sense to separate them. In my naive no knowledge version. if it was
RGRGRGRG
GBGBGBGB
RGRGRGRG
GBGBGBGB
Then, if I compress losslessly just grayscale I know at the other end I just interpret the first pixel as Red, next as green, red, green, etc., Then on the next row I interpret the first pixels as green, next as blue, next as green etc...
BUT, if I compress lossily then the unrelated colors next to each other are going to bleed into each other while compressing. VS if I separate them out so
RGRGRGRG
GBGBGBGB
RGRGRGRG
GBGBGBGB
becomes
RRRR BBBB
RRRR BBBB
GGGG
GGGG
GGGG
GGGG
The compression would make more sense and less bleeding of the signals from each color into other colors.
It would also seem to compress better. Imagine you're filming something pure green. Exaggerating but you'd get this
01010101
10101010
01010101
10101010
Which after separation is
0000 0000
0000 0000
1111
1111
1111
1111
That's just a guess though.