Is data being inserted into CH as it's received, or is there an intermediary buffer? A general overview of the flight of telemetry data through the system would be very welcome.
Tail-based sampling will require buffering spans in memory for some time, but tail-based sampling is not implemented yet.
Cloud version also uses Kafka to survive surges in traffic, but I guess "personal" / company version does not need that as much. So no need to introduce additional dependency.
There is some discussion at https://github.com/open-telemetry/opentelemetry-collector-co...