Sampling via just enabling it for some hosts/partitions is one solution (if you're producing 100M entries a day ... probably could just grab 1/100 of those for parsing).
Another solution is pre-processing (serial dupes are not forwarded).
Another solution is heavily reduced logging (ERR or higher only on prod hosts).
These can be used together and be very helpful.