AFAICT this system feels like a decent choice. Alternatives?
AFAICT this system feels like a decent choice. Alternatives?
Even google cloud and others let you wait for longer search queries. If not business ciritical, you can definitly wait a bit.
And the write system might not need to write it in the endformat. Especially as it also has to handle transformation and filtering.
Nonetheless, as mentioned in my other comment, the interesting details of this is missing.
So subsecond I would say is a requirement.
And no, it doesn't have to be the same system that ingests/indexes the logs.
You can easily entertain users to show them that the system is doing something in the background without loosing them and if they are collegues who actually need to search, you don't even need to keep them as they have to use your setup.
I think achieving sub-second read latency of adhoc text searching over ~150B rows of unstructured data is going to be quite challenging without a high cost. Clickhouse’s inverted indices are still experimental.
If the data can be organized in a way that is conducive to the searching itself, or structured it into columns, that’s definitely possible. Otherwise I suppose a large number of CPUs (150-300) to split the job and just brute force each search?
Re: advanced querying: the recommended way to do this is to build an index out of band (like Redis (or a fork) or SQLite or something) that references the stored messages by sequence number. By doing that, your index is just this ephemeral thing that can be dynamically built to exactly optimize for the queries you're using it for.
Re: sharding: no, it doesn't support simple sharding. You can achieve sharding by standing up multiple NATS instances, and making a new stream (KV and object store are also just streams) on each instance, and capture some subset of the stream on each instance. The client (or perhaps a service querying on behalf of the client) would have to me smart enough to be able to mux the sources together.