65 karma · joined August 7, 2023
It is impossible to detect anomalies in log patterns using deterministic methods alone. That is why companies like Datadog, etc have stopped at anomaly detection at the metrics layer. Because they are numbers you can predict. And you can't feed your entire firehose to LLMs because they blow up in compute.
I developed Rocketgraph that generates "snapshots" from billions of logs so that your agents can query and root cause without hallucinating and burning your engineering budget. First, we fingerprint the logs by masking away all the PII stuff, then use fuzzy matching to group together similar logs using TF-IDF. Then we apply IsolationForest to rank the logs with an anomaly score. By now, we have condensed them to 100-1000 log patterns we call a "snapshot". Here is the interesting part: we just make an LLM call with the service graph dependency map to root cause over the log patterns, and the result is something like what is shown. They are highly accurate.
It can be used to detect anomalous retry loops, weird call patterns, unseen formats, etc.
Basically, I'm building Datadog - but the user is an AI agent. A monitoring tool whose output is queryable and consumable by an AI agent.
I would love to hear your thoughts on this.
I built a tool that flags out anomalies. The rarest of the rarest logs by clustering them. This is how it works:
1. connects to existing Loki/New Relic/Datadog, etc - pulls logs from there every few minutes
2. Applies Drain3(https://github.com/logpai/Drain3) - A template miner to retract PIIs. Also, "user 1234 crashed" and "user 5678 crashed" are the same log pattern but different logs.
3. Applies IsolationForest(https://scikit-learn.org/stable/modules/generated/sklearn.en...) - to detect anomalies. It extracts features like when it happened, how many of the logs are errors/warn. What is the log volume and error rate. Then it splits them into trees(forests). The earlier the split, the farther the anomaly. And scores these anomalies.
4. Generate a snapshot of the log clusters formed. Red dots describe the most anomalous log patterns. Clicking on it gives a few samples from that cluster.
Use cases: You can answer questions like "Have we seen this log before?". We stream a compact snapshot of the clusters formed to an endpoint of your choice. Your developer can write a cheap LLM pass to check if it needs to wake a developer at 3 a.m for this? Or just store them in Slack.
It's in the following 2 links on the OpenAI website:
- https://cookbook.openai.com/examples/question_answering_usin... - https://cookbook.openai.com/examples/embedding_wikipedia_art...
I trained a chatbot on the Godfather script for example and asked Open AI to give the exact response to a dialogue in the film. I'll open-source that too pretty soon if you guys are interested.
I Previously wrote about Postgres and pgAudit in detail here: https://news.ycombinator.com/item?id=37247275 https://news.ycombinator.com/item?id=37082827
Since a lot of you are interested, I put together an article on how to install pgAudit on your own Postgres instance here: https://blog.rocketgraph.io/posts/install-pgaudit
Using pgAudit will help you debug your postgres instances better. Further you can add cloudwatch to query these logs in a timeframe.
I wrote about it here: https://blog.rocketgraph.io/posts/install-pgaudit
That's why I automated the process. Now every project comes with Postgres and pgAudit configured, Authentication, GraphQL and front-end SDKs right out of the box so it will be easier to develop web applications.