We spit out some 30k+ spans per second, FWIW. :-)
Edit: Disclaimer, we're not using Tempo.
I think I must be missing something, but it seems like the big difference between Tempo and a traditional tracing system is the storage indexing & database (ES/C* vs object store and index all fields vs key/value lookup by ID). I vaguely remember reading something that latency even from EC2 -> S3 can be around 200-300ms. Wouldn't this cause the overhead to rise?
Feel free to point me to any documentation that clear this up!
Disclaimer: I am using Tempo :) (and from Grafana)
I was testing Google Pub/Sub's Go client for publishing internal API event data for later ingest to BigQuery, and it turns out Pub/Sub publishing is not that much faster than writing directly to BigQuery. The buffer sizes we'd need to avoid adding latency to our APIs would have to be ridiculously high; the Pub/Sub client buffers and submits batches in the background (its default buffer size is 100MB!). I don't like the idea of having huge buffers that increase with the request rate.
Conversely, pushing the data to NATS in recent time without any buffering or batching turned out to be fast enough to not add any latency. You have to be able to receive messages very fast on the consumer side (as NATS will start dropping messages if consumers can't keep up), but you can simply run a few big horizontally autoscaled ingest processes that can sit there ingesting as fast as they can, which never impacts API latency at all.