If you're running usage-based billing for AI, infra, or API-heavy platforms— How do you deal with high-throughput event ingestion (say, 10k+ events/sec) without dropping events or messing up customer metering?
We’ve seen setups struggle hard with:
Event ordering guarantees
Idempotency at scale
Handling retries without double-counting
Would love to hear what infra patterns, queues, or storage choices worked (or failed) for you—especially?