HNHacker News
TopNewBestAskShowJobs

mapleeman

64 karma · joined June 24, 2022

Software engineer at Memgraph
submissionscomments
mapleeman··on What is the best data model for AI?
The Bitter Lesson says general methods win. So which data model is least "human-sophisticated" for AI?

I think it comes down to two metrics: flexibility and structure.

mapleeman··on Show HN: We build a Graph of public Skills
I would highlight the whole repo approach -> https://github.com/cloudflare/skills

Particulary -> https://github.com/cloudflare/skills/blob/main/skills/cloudf...

mapleeman··on Show HN: We build a Graph of public Skills
So this is not necessarily the tool that should help you with building your own skills, but rather the tool that will help you search so you can dynamically fetch skills that are relevant to you.

Think: fetch me the Rust skills for kube.rs for building the Kubernetes operator -> then you can relate what you need to your own skills, or see examples you can build from your specific cases.

That search feature is on the roadmap.

mapleeman··on Show HN: We build a Graph of public Skills
Hi all, author of the skillinsight.io[1]

A bit of a personal reflection on this, from a non-technical standpoint, Agent Skills are the procedural memory you keep in your head on how to solve a particular problem, which can be very valuable, whether you are aware of that or not. You are constantly adapting and changing that procedural memory since the task is usually not fully deterministic, hence it cannot be a script. Parallel to that, skills have caused some controversy for being a security vulnerability and hallucinated LLM brain fog, but more on that in the future.

Staying on the positive side of things and ignoring the negatives for now, agent skills could hold all the operational knowledge, allowing agents to operate semi-autonomously or autonomously to solve the particular operational problem. An example of that would be compiling a Memgraph Rust query module, which is not an easy task since you need the environment, the Memgraph query module API dependency, and knowledge of how to actually do it. Most advanced LLMs, like Codex or Opus, succeed at this after many tries and failures. This is why we build skills for compiling and deploying C++, Rust, and Python query modules that let LLMs practically single-shot the whole process.

Back to the topic of the graph of skills, what is the actuall problem here? So if you have hundreds or thousands of skills in your organisation, the question is: how are you going to maintain them, how will they learn and evolve, and how will agents access them? If the tool's API changes, so should the skills, which causes a cascade of events across the files. Then the question becomes: how are those connected and correlated? This is what graphs as a structure are built for, and this is what we in Memgraph are trying to solve from different angles.

The graph of skills will serve as our test bench for running the evolution, traceability, and access to the skills, while improving Memgraph as the graph database that serves as a real-time context engine for AI.

[1] https://skillinsight.io

mapleeman··on Handling Large Graph Datasets
Yeah, size classification is always tricky since the reference point always moves.

If you deal more with large datasets, the spectrum would ideally move to the right for you, as you have described, since you probably need a trillion on that scale.

mapleeman··on Ask HN: Were Graph Databases a Mirage?
Relational databases have a much longer history of development, and much more engineering time has went into designing RDBMS. It is not a surprise that they are mature on more levels. By looking at the age of a product, you can get a sense of how mature RDBMS systems are compared to most GraphDB projects.

I assume you really like the RedisGraph, now if you think about how much time of development did the RedisGraph had, 5 years? For example, Postgres is 37 years old :D

Horizontal scaling is hard in GraphDBs due to the nature of how the graph is structured and how you interact with it. Jumping across different servers is expensive. This is not a trivial problem that requires a lot of work and investment that Graph-based engines didn't have time to receive.

When we come to the Cyhper, it is easy to shoot yourself in the foot; that is true. It takes time to get used to tricks with Cypher and to know how to use it properly.

Performance, ingress, memory usage, etc is very specific per each system architecture so it is hard to comment.

This brings me to the actual workload, GraphDB/Graph analytical engines, and Knowledge Graph solutions will have their place under the sun, the more complex the use-case the more value the products will bring.

Just look what Deep Mind is doing with GNN and graph tech today: https://deepmind.google/discover/blog/millions-of-new-materi...

But what Peter said it should not be used in simple use cases where the cost of ownership does not bring value because the underlying problem is fairly simple, and Postgres is Free after all.

Disclaimer: I work for Memgraph https://memgraph.com

mapleeman··on Show HN: Benchgraph – Run graph database benchmark on your own dataset
A few months ago, we shared an initial version of Benchgraph[1] here on HN[2]. There was a lot of debate about benchmarking, both positive and negative[3].

Based on your feedback, we have updated the process with new queries, datasets, vulcanic mode, more workers, etc.

The biggest improvement is that you can run benchmarks on your data[4].

If you are more of a visual person, we have a video tutorial on this topic[5].

Feel free to drop any feedback. If you are interested in the latest results, take a look at the latest blog post[6].

[1] https://memgraph.com/benchgraph/base

[2] https://news.ycombinator.com/item?id=33813781

[3] https://news.ycombinator.com/item?id=34342371

[4] https://memgraph.com/blog/benchmark-memgraph-or-neo4j-with-b...

[5] https://www.youtube.com/watch?v=vzc1iVtHgeE

[6] https://memgraph.com/blog/benchgraph-backstory-the-untapped-...

mapleeman··on Bullshit graph database performance benchmarks
This is true regarding the transactions and cypherl. All data is cypherl transactions because memgraph can handle a large volume of transactions. mgbench was designed to run in-house CI/CD, and mgBench is still tightly coupled with Memgraph. That is the reason we are still running everything in transactions. We did open an issue where we plan to improve things, adding CSV support for faster imports being one of them. https://github.com/memgraph/memgraph/issues/689 Feel free to suggest things, some things Max suggested we will add. Agree on the more complex queries, and different vendors.
mapleeman··on Bullshit graph database performance benchmarks
This is indeed a good industry leading benchmark.
mapleeman··on [dead]
Hi! I'm the author of the blog post, and I'd be more than glad to answer any of your questions. I've already covered a lot of things in related Show HN last week (https://news.ycombinator.com/item?id=33813781), and we got a lot of feedback. We will try to implement as many of your suggestions as possible
mapleeman··on Show HN: We have built a benchmark platform for graph databases
The difference comes from several things, JVM being the first and obvious one. It takes as much memory as whole Memgraph + small dataset in this case. Second is the overallocation that JVM/Neo4j is doing, taking a bit more memory to have future space to grow. The third one comes to the implementations of storage and Neo4j cache implementation, not sure what Neo4j does on that side and how impactful it is. We use a skip list.
mapleeman··on Show HN: We have built a benchmark platform for graph databases
Exactly, you are right. We were able to decrease Memory usage to 1GB in Neo4j case, but then experienced some crashes. We then just removed the limit, and let it take us much as it needs.
mapleeman··on Show HN: We have built a benchmark platform for graph databases
Yep, you are right. We will expand the quantity, complexity, and variety of queries in future versions.
mapleeman··on Show HN: We have built a benchmark platform for graph databases
Glad you found them. Thanks for the tip!
mapleeman··on Show HN: We have built a benchmark platform for graph databases
No, currently, Memgraph is loaded 100% in RAM. Agree with the point about latency regarding I/O. The story with Neo4j is a bit more complex one. They are loading graph in memory, and using disk storage, doing both. But take a look at memory usage from them, it is the magnitude higher then Memgraph.
mapleeman··on Show HN: We have built a benchmark platform for graph databases
You are right about this 100% percent, performance is not the only factor, especially in an established DB space such as a relational DB world. There is a lot of things to consider before moving/or picking the database.
mapleeman··on Show HN: We have built a benchmark platform for graph databases
Added Dgraph to the reported ideas/request for the next benchmark version: https://github.com/memgraph/memgraph/issues/689 :D.

What are you doing with Dgraph, and what are the requirements for use-case? Of course, if you can share more info?

mapleeman··on Show HN: We have built a benchmark platform for graph databases
Yep, that is possible. Actually, one of the community members few days ago did this: https://gigi.nullneuron.net/gigilabs/using-the-neo4j-bolt-dr... You just need to tell driver, it is not Neo4j.

If you want to continue using the Neo4j driver. We actually also have GQL Alchemy, which is ORM for Python. Take a look here: https://memgraph.com/gqlalchemy

mapleeman··on Show HN: We have built a benchmark platform for graph databases
Thanks for all the comments and inputs. I have added the suggestions that we plan to implement: https://github.com/memgraph/memgraph/issues/689.

Both on Memory usage tracking and precise data on load/input.

Regarding scale, we are aware of the issue, listed in limitations: https://github.com/memgraph/memgraph/tree/master/tests/mgben.... Next versions will probably have a billion nodes/relationships.

Actually, Neo4j is particularly slow on writes, import/load times were 50x faster on Memgraph, but we didn't show it. Will do it in the next version for all vendors.

mapleeman··on Show HN: We have built a benchmark platform for graph databases
So these results were not cherry-picked since the same queries were run previously in our CI infrastructure. Yep, as you have said, performance is just one of the things that matter. A lot of things matter when picking a vendor, some are mentioned in the comments.

What is specific in graph space is that things are still quite early days compared to relational database space. This means performance differences are big, and playing a more important role.

mapleeman··on Show HN: We have built a benchmark platform for graph databases
Thanks for the input. It is a fair point. It is hard to create benchmark that is universally fair. But both things serve the same/similar purpose, on top of that, Neo4j also loads a bunch of stuff into RAM, consuming more memory than Memgraph, if your dataset fits in RAM, it because quite a fair comparison.
mapleeman··on Show HN: We have built a benchmark platform for graph databases
Yes, good point, but both Neo4j and Memgraph(https://memgraph.com/product) are ACID-compliant and have on-disk persistent storage. Memgraph is currently RAM constricted, while Neo4j is not but it is hungry for RAM.
mapleeman··on Show HN: We have built a benchmark platform for graph databases
This is a great idea, with one relational database as a reference point for every query on the graph database. I added it to the backlog: https://github.com/memgraph/memgraph/issues/689 Thanks for the idea! :D
mapleeman··on Show HN: We have built a benchmark platform for graph databases
We plan to add more graph database vendors to this benchmark, this will not be just Memgraph vs Neo4j comparison, hence the name "bench graph". You are 100% right about comparing architecturally different database systems, it is hard to compare them, but they serve the same/similar purposes, we are mentioning that part in limitations methodology: https://github.com/memgraph/memgraph/tree/master/tests/mgben... We first added Neo4j since we are both compatible.
mapleeman··on Show HN: We have built a benchmark platform for graph databases
So is it because of C++ vs Java? Well, not everything can be written just to these differences, a lot of stuff can influence results, architecture, database isolation level etc. One of the many reasons is definitely a C++ and Java argument. Take a look at memory usage, here, you can see how JVM is memory hungry. Also, it takes time to warm JVM and Neo4j, so it is another penalty for the same reason.

So far, on this dataset and scale, we didn't encounter but we have plans for a bigger dataset and more complex queries, you can take a look at the limitations part of this benchmark.

What are the downsides, Memgraph and Neo4j are currently a bit different vendors, Memgraph is an in-memory database, while Neo4j is on disk. So in Memgraph's case, you are exclusively using RAM as a storage but gain speed, we have snapshots for disk for recovery, etc. While Neo4j is on disk, not-in memory but they are loading a bunch of stuff in RAM and consuming waste amounts of RAM, so it is hard to give pure distinction.

mapleeman··on Show HN: We have built a benchmark platform for graph databases
Yes, testing and benchmarking is a full-time job! The extra issue here is making it work under a single client to minimise latency penalties across measurements, then, there is also a protocol issue. Nice project with hash-db, I guess it is quite the learning experience(distributed-multimodal)?
mapleeman··on Show HN: We have built a benchmark platform for graph databases
Haha, adding to a backlog of things to do, add support for TigerGraph https://github.com/memgraph/memgraph/issues/689
mapleeman··on Show HN: We have built a benchmark platform for graph databases
Hey, thanks for the comment, you are 100% right, this is just the initial version since we are compatible with Neo4j, so it was a least effort to do it. It is just the initial setup, making it language agnostic will take a bit time. If you peak at the methodology and future part: https://github.com/memgraph/memgraph/tree/master/tests/mgben... You will see that we have the plan to add more database vendors + make it language-agnostic. We are also keeping track of all comments regarding this, I have opened an issue: https://github.com/memgraph/memgraph/issues/689, there is a language agnostic note in there. If you have any other input, it would be highly appreciated.
mapleeman··on Show HN: We have built a benchmark platform for graph databases
Hi everyone! I’m one of the co-authors of BenchGraph[1], a platform for Graph Database Performance Benchmarks. Our platform shows the results of running benchmark tests (via mgBench) on supported vendors. It shows the overall performance of each system relative to others.

Inspiration came from ClickBench, a Benchmark For Analytical DBMS.

We previously developed mgBench as in-house testing infrastructure to benchmark Memgraph, and now we are adapting it to support other graph database vendors. In order to test graph database performance, mgBench executes Cypher queries on a given dataset. Queries are general and represent a typical workload that would be used to analyse any graph dataset. Running this benchmark is automated, and the code used to run benchmarks is publicly available. You can run mgBench yourself to validate the results on the BenchGraph platform. The methodology is explained in detail on GitHub repo [2]

As you can see, at the moment, we have two vendors on the platform. We would like to add more vendors to our platform. If you want, feel free to contribute.

Let me know if you have any questions or suggestions.

[1] https://memgraph.com/benchgraph [2] https://github.com/memgraph/memgraph/tree/master/tests/mgben...