The 44ms you’re referring to is the time it takes our API to respond to an incoming call and ingest an event. It’s certainly not the wall clock time, but it is a good measure of the overall health of the system. Your feedback about its prominence is definitely good–it’s our goal to be transparent, not misleading. We’ll change the area where it’s displayed shortly.
If you’re looking for the ‘end-to-end’ wall clock time, you can find that a little further down the page. For every event coming through our pipeline, we average the time from ingestion to successful delivery and display those metrics on a per-destination basis.
You can see right now that the Google Analytics end-to-end delivery latency is ~400ms within the past hour, and is pretty consistently near that number. We’ve also developed internal systems which break down exactly how many events from a given source were delivered to a given destination, and what the latency distribution for those events was.
The ordering problem you mention is indeed tricky. Like TCP, if we wanted to keep a loose ordering, we’d have to keep a window and ensure the partner API would then re-order messages appropriately. If you want a total ordering, this window of delivery _has_ to be 1 for a given user. It does have some pretty serious implications on the throughput of the system, so we’ve been working with both Mixpanel and Intercom (we’re users ourselves) to try and solve the issue just with timestamps. Ideally, partners would be able to re-order events received based upon time, which is what we do inside of customers’ data warehouses.
As far as the deliverability issues you mentioned, I’m terribly sorry to hear that we failed you here. We’ve hit some scaling bottlenecks that we’ve been working hard to fix–and we do our best to keep the status page updated whenever we have production incidents.
All that said, reliability is our top focus as a company. Teams present their SLA metrics on a weekly basis at all hands, and it’s a key part of our monthly board reporting. We’ll be surfacing these metrics and event traces inside the webapp so as a customer can see exactly where your data is, and what has been delivered. And we’re in the process of building an entirely revamped pipeline that will provide better deliverability guarantees. We plan on sharing the architecture on the blog once it’s all rolled out. Giving you transparency into where your data is stored and how it is processed is exactly what we want to achieve as a company.
If there’s anything I missed–please reach out! I’m calvin at segment.