Do you also use rolling checksums (like bup) to prevent re-storing data when only a few bytes have changed?
--> What are the major use cases you imagine for noms?
I've read the section on Github, but can you give some specific examples that you see as good use cases?
https://github.com/attic-labs/noms/blob/master/doc/intro.md
We were also heavily influenced by camlistore (which I hacked on for awhile), irmin, ipfs, and others who have done a lot of interesting work in this space.
We do use rolling checksums, but I think we have done some novel work here: https://github.com/attic-labs/noms/blob/master/doc/intro.md#...
PS: I love the name "Prolly Tree".
[1] <https://github.com/attic-labs/noms/blob/master/datas/databas...
Do you have any mechanisms for securing integrity, specifically repairing the store in case of inconsistencies?
Is there any plans to support any data retention policy/functionality?
You can take the JSON output of an API and drop it into Noms, then do the same thing tomorrow, and Noms will automatically deduplicate the data as well as give you a nice structured API to read and interact with it.
We have an example of this here: https://github.com/attic-labs/noms/tree/master/samples/js/fl... but it's not working atm due to a bug introduced right before launch. You can look at the code though.
However, this falls down pretty rapidly. In order to get reasonable diffs, the data has to be sorted, and line-oriented. Also Git just doesn't scale well to larger repos or individual objects.
Otherwise, we see the competitors as the way that people distribute data today - custom APIs, zip files full of CSV, etc.
From my quick inspection, it looks like Noms shows some focus towards working in multiple branches, whereas Datomic, at least in its marketing materials, just talks about preserving a single timeline.
I guess the above means that Noms does scale to larger repos... Do you guys have any numbers, comparisons, benchmarks against git?
If so, it would be useful to include in the readme.md as it would be kind of a big thing, and quite attractive to many people.
Schema inference is very difficult to do correctly and safely, especially with small initial samples of instances (source: work on https://github.com/snowplow/schema-guru).
In Noms every value has a type. It's an immutable system, so this type just is. The type of `42` is `Number`. The type of `"foobar"` is `String`. The type of `[42,44]` is `List<Number>`. And if you add "foo" to that list, the type becomes `List<Number|String>`.
We don't try to infer a general database schema from a few instances of data. We just apply this aggregation up the tree and report the result.
That all said, we do want to eventually add schema _validation_, by which I mean the ability to associate a type with a dataset and have the database enforce that any value committed to the dataset is compatible with that type (following subtyping rules).
Are you planning on writing complete reference documentation at some point, like https://www.sqlite.org/limits.html, https://www.sqlite.org/howtocorrupt.html, https://www.sqlite.org/lang.html, https://docs.python.org/2/reference/index.html, and https://golang.org/ref/spec? Or is using Noms going to be more of a UTSL kind of thing? The documentation I've found so far seems to be purely tutorial and introductory in nature.
(I'm really glad you're writing Noms, by the way. There's an enormous need for it.)
Given that schema validation ("does this instance match this type?") is simpler to implement than schema inference ("what is the type of this instance?"), it's surprising to me to deliver inference first...
Schema validation for us is just looking at the type requirements of the dataset and the type of the value and seeing if they are compatible. How can we do that without first knowing the type of the value?
IPFS is essentially providing a globally decentralized filesystem. Noms is providing (or hopes to provide) a database.
By database, I mean:
- small individual records
- efficient queries, updates, and range scans
- ability to support complex queries
- ability to enforce structural data validity
These are all things that IPFS could eventually grow to support, but in order to do it, I think it would have to grow into or layer something like noms on top.