HNHacker News
TopNewBestAskShowJobs

adsharma

543 karma · joined March 19, 2016

ladybugdb.com
submissionscomments
adsharma··on Why is Hacker News like that?
Jensen Huang is neither a Nazi, nor far right.

HN has other biases that mirror the biases of the avg tech worker in the valley and those of the people called out in the article.

But they don't neatly fit into "far left" and "far right" buckets. People who use those terms usually are professional agitators.

https://youtube.com/shorts/h3dqFk3FcRk?is=uxh_cCfkuhj8yWBP

adsharma··on A warning about 'model welfare'
Yes, which is why the focus should be on how they're connected together. Not what happens after we connect them to create Jason Bourne.

SQLite was mentioned as a placeholder to make it a queryable database. Genome as an analogy.

We need to develop new type of databases and connect them to ML.

adsharma··on A warning about 'model welfare'
A 10TB SSD is not conscious. SQLite is not conscious. A wafer is not conscious.

But connect them all together...

adsharma··on We must pace the frontier
The problem is that you're not offered a choice. No one is making the "err towards explainability" choice.

Training data is treated as IP. Distillation is seen as an attack.

Open data, open training based systems such an Marin are just getting started. Explainability is not a priority there.

The ones who do discuss these ideas are confrontational about LLMs and not effective spokespeople.

adsharma··on We must pace the frontier
What you're using is a transformer based database index. An obfuscated one.

Training data is the database. Model is a lossy compressed index.

adsharma··on We must pace the frontier
I'm sympathetic to this argument.

All I'm saying is: if you have a choice between two systems with equal power of discovery and one is more understandable than the other, we choose the more understandable one.

Limits of Human comprehension and quest for power are two different motivations that could lead to black box systems that are marketed as semi-explainable.

We need to verify that human comprehension is actually limiting progress before allowing such things and even when we do, do it responsibly on an explainable foundation.

adsharma··on We must pace the frontier
Oh you want tech that helps discover new science instead of parroting existing wisdom?

There is little evidence that the RSI we are discussing is capable of inventing the theory of relativity (or the more advanced equivalent). All we have seen is pattern matching in a much larger space than humans can, with some human provided verification tech.

I would argue that human-AI collaboration with explainable tech has a better chance. Continuous learning can be done in a way that doesn't violate IP or privacy.

adsharma··on We must pace the frontier
If there is guillotine in the weights of the model, it needs to be properly labeled so you can look it up by name using a database index (or a graph-vector index).

It helps both the bad guys and good guys. Like responsible disclosure in cyber security, we need to have a conversation around it.

adsharma··on We must pace the frontier
You can have a transformer based index built on top of a graph database. Much of the explainability tech (people who prefer "alignment" also prefer "mechanistic interpretability") is reverse engineering the real world graph that's hiding in the weights.

https://research.google/pubs/the-case-for-learned-index-stru...

adsharma··on We must pace the frontier
Statistical learned indexes can exist in explainable tech such as a database.

In 2017 Google was writing papers about it. Then something changed.

I don't think it was the tech. It was a realization around the power and societal impact.

adsharma··on We must pace the frontier
The gun comes with a manual on how to use it safely. I'm sure it has some complexities, but at the high level:

> A gun is a metal tube that uses a tiny, controlled explosion to shoot a small piece of metal (called a bullet) forward at very high speed.

If the gun doesn't work as intended, you can take it to a shop and someone can fix it so it works as designed.

All I'm saying is AI should be designed the same way. Treat AI as normal tech like any other and use similar language.

adsharma··on We must pace the frontier
The real threat is that we uncritically adopt language such as alignment.

Implicit in this is the idea that AI is a inscrutable matrix and going to remain that way and we'll need expert interpreters to make sense of it.

We need to insist on building tech that's explainable by design.

adsharma··on We must pace the frontier
So many words. But missing the one that matters the most: Explainability.

AI slowdown is worth it only if it can be made more explainable.

Changing the language we use to discuss it is a good first step.

We need to stop using inside baseball terms like alignment and mechanistic interpretability. Replace them with explainable tech. Graph Databases, Causality, Shared semantic spaces.

Previous writings on the topic (also on LinkedIn, but can't find urls):

https://x.com/arundsharma/status/2005338775468282339 https://x.com/latentpedia

adsharma··on The Emergent Symbolic Structure of Artificial Neural Networks
I'm encouraged by this result. It's the primary hypothesis behind latentpedia.org.

Instead of distilling the geometry of a model into a huge knowledge graph, we start from the largest known open source graphs and build it up towards something that resembles this geometry.

Come and join us. Discuss on github.com/latentpedia. We have the basic tech covered. Need more compute, storage and enough business to cover the cost of serving.

adsharma··on Markdown Database Pattern
Simplicity is in the eye of a beholder. It's not an argument backed by facts.

Many single file graph databases out there. They should be "simpler" since it's one file vs many.

Text files vs binary is what you're arguing. Why don't you store your phone address book as a text file?

Everyone know the answer to that one. SQLite is a solid, well accepted answer. You already have 20 of them on your phone.

We need to build a similar option for graphs instead of the markdown detour.

adsharma··on Markdown Database Pattern
Many competent embedded graph DBs out there. More suitable than graphs on SQLite projects.
adsharma··on Markdown Database Pattern
For what benefit though?

Much better off storing this in an embedded graph database and use cypher.

MVCC, Indexes, Strong typing, WAL, changefeeds.

The simplicity of markdown and JSON are deceptive. SQLite is a solid choice, but then you deal with graph-relational impedance mismatch.

adsharma··on Show HN: LatticeDB – Like SQLite but for graph databases
https://www.youtube.com/@ladybugdb - what content would you like to see?
adsharma··on AWS Acquires DuckLabs
Several "Graph on DuckDB" efforts started on DuckDB and ended up reinventing a basic columnar codebase to innovate on because of these reasons. Even though people didn't document why, lack of outside contributor friendly flows is likely one of the reasons.

KuzuDB folks worked on something called GRainDB in 2022: https://vldb.org/cidrdb/2022/graindb-a-relational-core-graph...

But circa 2023 decided to write their own. Work continues as LadybugDB. One of my long term goals is to find an integration point with DuckDB's table implementation as the "node table". Conversely at some point DuckDB could implement all the join algorithms and LadybugDB's REL table in their code base.

For now, any talk of Graph on DuckDB is limited to DuckPGQ and the more recent entrant DuckGQL (both of which don't touch the storage layer which is the main reason why LadybugDB exists).

adsharma··on AWS Acquires DuckLabs
Surprised that there isn't more discussion of DuckDB source distributions.

Why do we need a source distribution for a well regarded MIT licensed project? Because it's not easy to contribute code to DuckDB if you don't work at DuckLabs. The CI used to take 5 hours for a simple bug fix last I looked (may have improved since I flagged it on social media).

There are two that I'm aware of:

  * Haybarn: https://rusty.today/blog/duckdb-extension-distribution-gap/
  * Pygmy-Goose: https://github.com/Pygmy-Goose/pygmy-goose
Pygmy-Goose is focused on making agentic workflows faster by splitting the repo, making git worktrees cheap and 5 minute cached CI on 3 major platforms.
adsharma··on Show HN: LatticeDB – Like SQLite but for graph databases
I measured this on a M4 mac mini (base model):

  zig build sqlite-benchmark

  Medium (100K nodes)

  +----------------------+-----------+---------+---------+
  | Workload             | LatticeDB | SQLite  | Speedup |
  +----------------------+-----------+---------+---------+
  | 1-hop traversal      | 5.7μs     | 16.1μs  | 2.8x    |
  | 2-hop traversal      | 30.1μs    | 59.4μs  | 2.0x    |
  | 3-hop traversal      | 171.1μs   | 228.8μs | 1.3x    |
  | Variable path (1..5) | 82.2μs    | 5.8ms   | 70.4x   |
  +----------------------+-----------+---------+---------+
Very different from the comparison on github and the website.

Given that on-disk data structures are similar to SQLite, I expect the competition from other "graph on sqlite" projects when they co-opt the techniques in LatticeDB.

adsharma··on Show HN: LatticeDB – Like SQLite but for graph databases
Wikidata doesn't imply RDF/SPARQL. Cypher works fine too. A columnar storage engine means you get indexes and the relational goodness for free.

https://huggingface.co/datasets/ladybugdb/wikidata-20260401

adsharma··on Show HN: LatticeDB – Like SQLite but for graph databases
LadybugDB maintainer here.

Couple of corrections:

* LadybugDB has revamped the Kuzu WAL design. It shouldn't be hard to build WAL based replication

* 19ms vs 39us - like the author says these are vastly different systems and the benchmark methodology may not be comparable.

We've mostly focused on query plan optimizations, not so much the micro query operator optimizations.

The 0.20.x end of the month release should have some interesting optimizations.

  * Prepared statements will cache query plans and result vectors. So you don't pay malloc costs
  * SIMD optimizations for filter. More to come in the next release.
adsharma··on Knowledge Compressor
https://huggingface.co/datasets/ladybugdb/github-knowledge-c...

These methods could not beat the 480 token result because of the small dataset. But for a sufficiently large corpus with graph reordering and columnar parquet compression, they could be quite competitive.

The parquet files are self explanatory. For knowledge.lbdb.zst, first decompress with zstd and then run a cypher such as:

  lbug> MATCH (a:E)-[t:TRIPLE]->(b:E) RETURN a.name AS subj, t.predicate AS pred, b.name AS obj, t.props AS props, t.scope AS scope;
adsharma··on Knowledge Compressor
Did you consider extracting knowledge from the doc into a graph database (while passing the Q&A) and compressing the graph instead?
adsharma··on The Mojo language (by Modular, now Qualcomm) is now open-source
[ mojo explaining why they didn't create an embedded DSL]

> This is particularly problematic if you're trying to introduce fundamental new concepts because you can't change the grammar of Python or C++.

Yet, this is exactly the approach py2many is taking.

Keep it 99% python compatible. Deviate from the python grammar only for new features and where python made a decision incompatible with compiled language backends.

adsharma··on The Mojo language (by Modular, now Qualcomm) is now open-source
Mojo as MLIR++ is how I've thought about it as well.

Many users will find transpiling static python to mojo an interesting path.

Updating py2many --mojo to 1.0 to see what breaks.

Also looking forward to translating existing python adt module to mojo Variant in phase 2. I hope the mojo pattern matching proposal looks more like rust (expression) and not like python (statement).

adsharma··on Mojo 1.0
Given the shared heritage with MLIR, one way to think about mojo: higher level IR, but still an IR. Rpython for GPUs.

But mojo is a superset, not a subset. So why not use a subset, infer what you need and generate mojo?

I've never seen this question actually presented to the company and discussed in more detail.

adsharma··on Mojo 1.0
https://github.com/py2many/static-python-skill

You can stick to python and generate mojo or rust or lean.

There is more than one solution to the two language problem.

adsharma··on What I love about Django
If you must use Django, use it through an abstraction layer like this:

https://adsharma.github.io/django-fquery/

Your models can be plain old python data classes, declaratively mapped to Django primitives.

Page 1 of 14Next →