I think it comes down to two metrics: flexibility and structure.
64 karma · joined June 24, 2022
I think it comes down to two metrics: flexibility and structure.
Particulary -> https://github.com/cloudflare/skills/blob/main/skills/cloudf...
Think: fetch me the Rust skills for kube.rs for building the Kubernetes operator -> then you can relate what you need to your own skills, or see examples you can build from your specific cases.
That search feature is on the roadmap.
A bit of a personal reflection on this, from a non-technical standpoint, Agent Skills are the procedural memory you keep in your head on how to solve a particular problem, which can be very valuable, whether you are aware of that or not. You are constantly adapting and changing that procedural memory since the task is usually not fully deterministic, hence it cannot be a script. Parallel to that, skills have caused some controversy for being a security vulnerability and hallucinated LLM brain fog, but more on that in the future.
Staying on the positive side of things and ignoring the negatives for now, agent skills could hold all the operational knowledge, allowing agents to operate semi-autonomously or autonomously to solve the particular operational problem. An example of that would be compiling a Memgraph Rust query module, which is not an easy task since you need the environment, the Memgraph query module API dependency, and knowledge of how to actually do it. Most advanced LLMs, like Codex or Opus, succeed at this after many tries and failures. This is why we build skills for compiling and deploying C++, Rust, and Python query modules that let LLMs practically single-shot the whole process.
Back to the topic of the graph of skills, what is the actuall problem here? So if you have hundreds or thousands of skills in your organisation, the question is: how are you going to maintain them, how will they learn and evolve, and how will agents access them? If the tool's API changes, so should the skills, which causes a cascade of events across the files. Then the question becomes: how are those connected and correlated? This is what graphs as a structure are built for, and this is what we in Memgraph are trying to solve from different angles.
The graph of skills will serve as our test bench for running the evolution, traceability, and access to the skills, while improving Memgraph as the graph database that serves as a real-time context engine for AI.
If you deal more with large datasets, the spectrum would ideally move to the right for you, as you have described, since you probably need a trillion on that scale.
I assume you really like the RedisGraph, now if you think about how much time of development did the RedisGraph had, 5 years? For example, Postgres is 37 years old :D
Horizontal scaling is hard in GraphDBs due to the nature of how the graph is structured and how you interact with it. Jumping across different servers is expensive. This is not a trivial problem that requires a lot of work and investment that Graph-based engines didn't have time to receive.
When we come to the Cyhper, it is easy to shoot yourself in the foot; that is true. It takes time to get used to tricks with Cypher and to know how to use it properly.
Performance, ingress, memory usage, etc is very specific per each system architecture so it is hard to comment.
This brings me to the actual workload, GraphDB/Graph analytical engines, and Knowledge Graph solutions will have their place under the sun, the more complex the use-case the more value the products will bring.
Just look what Deep Mind is doing with GNN and graph tech today: https://deepmind.google/discover/blog/millions-of-new-materi...
But what Peter said it should not be used in simple use cases where the cost of ownership does not bring value because the underlying problem is fairly simple, and Postgres is Free after all.
Disclaimer: I work for Memgraph https://memgraph.com
Based on your feedback, we have updated the process with new queries, datasets, vulcanic mode, more workers, etc.
The biggest improvement is that you can run benchmarks on your data[4].
If you are more of a visual person, we have a video tutorial on this topic[5].
Feel free to drop any feedback. If you are interested in the latest results, take a look at the latest blog post[6].
[1] https://memgraph.com/benchgraph/base
[2] https://news.ycombinator.com/item?id=33813781
[3] https://news.ycombinator.com/item?id=34342371
[4] https://memgraph.com/blog/benchmark-memgraph-or-neo4j-with-b...
[5] https://www.youtube.com/watch?v=vzc1iVtHgeE
[6] https://memgraph.com/blog/benchgraph-backstory-the-untapped-...
What are you doing with Dgraph, and what are the requirements for use-case? Of course, if you can share more info?
If you want to continue using the Neo4j driver. We actually also have GQL Alchemy, which is ORM for Python. Take a look here: https://memgraph.com/gqlalchemy
Both on Memory usage tracking and precise data on load/input.
Regarding scale, we are aware of the issue, listed in limitations: https://github.com/memgraph/memgraph/tree/master/tests/mgben.... Next versions will probably have a billion nodes/relationships.
Actually, Neo4j is particularly slow on writes, import/load times were 50x faster on Memgraph, but we didn't show it. Will do it in the next version for all vendors.
What is specific in graph space is that things are still quite early days compared to relational database space. This means performance differences are big, and playing a more important role.
So far, on this dataset and scale, we didn't encounter but we have plans for a bigger dataset and more complex queries, you can take a look at the limitations part of this benchmark.
What are the downsides, Memgraph and Neo4j are currently a bit different vendors, Memgraph is an in-memory database, while Neo4j is on disk. So in Memgraph's case, you are exclusively using RAM as a storage but gain speed, we have snapshots for disk for recovery, etc. While Neo4j is on disk, not-in memory but they are loading a bunch of stuff in RAM and consuming waste amounts of RAM, so it is hard to give pure distinction.
Inspiration came from ClickBench, a Benchmark For Analytical DBMS.
We previously developed mgBench as in-house testing infrastructure to benchmark Memgraph, and now we are adapting it to support other graph database vendors. In order to test graph database performance, mgBench executes Cypher queries on a given dataset. Queries are general and represent a typical workload that would be used to analyse any graph dataset. Running this benchmark is automated, and the code used to run benchmarks is publicly available. You can run mgBench yourself to validate the results on the BenchGraph platform. The methodology is explained in detail on GitHub repo [2]
As you can see, at the moment, we have two vendors on the platform. We would like to add more vendors to our platform. If you want, feel free to contribute.
Let me know if you have any questions or suggestions.
[1] https://memgraph.com/benchgraph [2] https://github.com/memgraph/memgraph/tree/master/tests/mgben...