Noms – A versioned, forkable, syncable database
github.com
github.com
EDIT: At least one team is investigating layering Noms on top of IPFS [1]. I guess the idea would be to construct something similar to GitTorrent [2]; layering various version-controlled datastores on various p2p protocols could result in several viable architectures.
[1] https://github.com/attic-labs/noms/issues/2123#issuecomment-...
Why not http://localhost:8000/dbname ?
https://tools.ietf.org/html/rfc3986 (see "3.3. Path").
data Database = InMemory | LevelDB Path | ViaHTTP URL
newtype DataSet = DS Text
newtype Hash = Hash Text
data Accessor = AccessDS DataSet | AccessValue (Either DataSet Hash) Path
type DBAccessor = (Database, Accessor)
They elected to basically encode a DBAccessor above as a string which you can split on "::", with the URL above being stored on the left in the case of the ViaHTTP databases. Syncthing does recognize conflicts. When a file has been modified on two devices simultaneously, one of the files will be renamed to <filename>.sync- conflict-<date>-<time>.<ext>. The device which has the larger value of the first 63 bits for his device ID will have his file marked as the conflicting file. Note that we only create sync-conflict files when the actual content differs.
https://docs.syncthing.net/users/faq.htmlIn the meantime, as long as I've got your attention, here's a few new stuffs we've been working on since last time Noms was discussed here in August:
- A prototype query language, and a demo of how to create indexes in Noms: https://www.youtube.com/watch?v=fv6_T5yaWns
- Support for merging concurrent (and potentially conflicting) changes: https://www.youtube.com/watch?v=--7dgoJBdjU
I mean: instead of syncing whole the database, only syncing the parts that a user has access to, and being able to define those accesses. The standard use case for consumer apps.
So yes, in principle, we can definitely do this. In practice we are missing some conveniences that would make it a really easy drop-in feature.
If it isn't ACID it needs to make a very strong case for itself to even be played with by most DBA's, including myself.
Noms doesn't manage its own storage - it relies on an underlying key/value store that must provide strongly consistent reads for at least one key. In other words, we delegate most of the hard part to somebody else.
With that all said...
Currently our intent is that:
- Transactions that read and write from a single dataset have strong serializability
- Transactions that read from multiple datasets and write to a single dataset have snapshot isolation
- Transactions that write to multiple datasets aren't possible
In the future, we will probably allow additional configuration, such that, e.g., one could choose snapshot isolation within a dataset for additional concurrency, or strong serializability for transactions that span datasets.
For extra fun try doing this with geographic data and try merging geometry changes correctly.
But I agree, there are certain kinds of data that you can't so efficiently merge. Your best bet is to try to adapt it into some sort of known, proven CRDT. (Dunno how you do that with this database, haven't really read up on it.)
The world contains logical conflicts because physical constraints mean that processes can operate disconnected from each other. No database can wave that away.
Noms will automatically, efficiently, and correctly merge changes that don't logically conflict. Which is a pretty cool and unique property in a database.
If any conflicts are found, there is a callback to user software to perform a resolution.
More info in the documentation:
I created a bug for this just now: https://github.com/attic-labs/noms/issues/2718
Please feel free to get involved there.
But...
* It's a big jump from relational or noSQL DB's, so there aren't (m)any adapters that I can see for it for JPA, ActiveRecord, etc.
* I'd really like to see a benchmark for each noms implementation compared to postgres, mysql, oracle, and mssql server, if there is a way to do apples-to-apples.
* "noms" is unfortunately is really bad for SEO because noms is a common word in French. If it could be nomsdb or nomnomnoms or something less exactly French, that'd be better. It's going to be tough to find support online easily otherwise.
* SQL compatibility.
* Fault tolerance (how easily does it corrupt), HA, mirroring, full/partial replication, sharding, archival, partial history truncation, etc.
It seems a little like a dolphin jumping into a pool of hungry sharks. It might be more evolved and more capable in some ways, but it's going to get its ass handed to it on speed and lack of features.
Still- I can't wait to try it.
I'm inclined to agree for large, centralized databases, but I wonder if this would be a good fit for places where sqlite is used? This seems like it could be a good foundation for situations where you want to sync information without a central server, like between devices. An Access/Filemaker clone built on top of this would be cool, too.
I do want to note some tradeoffs with Content-Addressed and Append-Only systems, as my work on a similar project ( an Open Source Firebase, https://github.com/amark/gun ) made me move away from those ideas (even though they are great ideas).
- Content-Addressed stores are going to revolutionize data integrity and efficiency. But they do have a trade off, it makes it a lot harder to read the data if you do not already know the data you are trying to read! The bottom of the repo metions for instance that a query system has not yet been built. From my experience the reason why is because it is difficult to build query systems on Content-Addressed stores, which is a tradeoff from all the gains you can get from it.
- Append-Only gives you rich features like offline-first support and (if implemented) lovely things like rewind/fastforward data time travel. All very cool. However, do not forget that this then also makes it difficult for you to retrieve the latest whole snapshot of your data. So you are not going to get the read performance that you could.
But the only possible way for us as a community, and people playing around with databases, can figure out what the best system is - is for people to build and experiment. Which is part of the reason why Nom is so cool. It is an invitation to others to actually join, play, and experiment with database technology in an open and encouraging environment. That is incredibly valuable and needed!
2. It's true that content addressing can exacerbate data locality which can hurt read performance. However, there are thing you can do to get a lot of that back.
Python code, if interested: https://github.com/WorldMaker/tokdiff
The nice thing about git is that it doesn't really require much of the server. I run my own "git hosting" with linode, apt-get, and ssh.
Most of the solutions for immutable big data don't have that level of convenience, as far as I know. Anything which requires a lot of sysadmin and ops work will eventually become a commercial cloud service... git on the other hand is useful by itself, and good companies can obviously be built on top of it as well.
The difference between this and a version control system is a major part of why the UI is so awful.
* Convert the versioned data to tab-separated values
* COPY it into Postgres every time
* Hope Postgres can act immutable enough even though it wasn't designed to be
The closest I've come to improving this situation was Kyoto Cabinet (unusable license) and rolling my own damn hashtable (it worked okay but adding new kinds of indexes was just unmaintainable, there's a reason databases should be made by experts).
Pachyderm is git for data. We work hard to make sure we can store data of different types (binary, text, json) efficiently. We also work hard to give you good mechanisms to read the data in a distributed way. I'd be curious how this suits your purposes.
While I can see how a git-based filesystem can help with some use cases, does it do any kind of indexing at all? I see that the FAQ recommends exporting the data from Pachyderm into PostgreSQL, which leaves me where I am now.
One way to think of Noms is that is an index optimized for computing diffs.
Here's a screencast that shows off Noms diffing things fast: https://www.youtube.com/watch?v=Zeg9CY3BMes
This seems like an interesting project that tackles some of the data versioning stuff. However, I believe that, at least in data science, we need data versioning closely tied to the analyses themselves for complete reproducibility.
That is, we need the versioning tied to the inputs/outputs of data pipeline stages, such that we can reproduce pipeline runs at any time and incrementally improve and run pipelines based on diffs in data.
As mentioned elsewhere in the comments, Pachyderm (http://pachyderm.io/) does exactly this. Working both as git for data, but also enabling data pipelining and analyses with the data versioning.
I've hand-rolled something a lot like this already for the Shaxpir backend, but it would be really nice to have a well-engineered database that already supports this kind of model, out of the box.
Why would I use this if I can't use it everywhere?
Also, supporting Windows is often a pain in the ass compared to Linux/Mac. I don't fault them for not supporting it officially, especially in the project's life.
Just getting a machine to test on is expensive. If you look for Mac OS VMs in the cloud, you find they start at $1 an hour (https://www.macincloud.com/) or around $80 / month (http://xcloud.me/pricing-signup/). Compare that to around $5 / month for a Linux VM.
And then you have to go the whole dance of getting xcode and homebrew just to have a development environment. It's been a few years since I used Mac OS X to develop stuff, but it wasn't intuitive at all back then.
But Macincloud has ~$20/month plans (with 8GB of RAM which isn't a $5 option for most Linux or Windows VMs) and there are several others going for between $30-50, so you hardly have to go with $80/month.
http://www.geekwire.com/2016/mac-overtakes-linux-as-develope...
So my comment should be correct in a year or two, hopefully.
Look again at that graph:
http://stackoverflow.com/research/developer-survey-2016#tech...
OS X is treated as a single category, but Windows is split over multiple versions. When you add up all the Windows versions, OS X isn't close to 50% market share.
Also, I'd question why it should be 'hopefully' correct. Other than better support for Unix command line tools, what gives OS X the edge over Windows 10 as a dev environment?
What was claimed is that most developers don't use Windows, not that they use macOS. If over 50% of developer use either macOS or Linux, then the claim is true.
Also, that doesn't answer my other question. Other than better support for Unix command line tools, what makes OS X (or Linux) a better platform for devs than Windows 10?
Yes, but the post you replied to said "in a year or two".
Also, that doesn't answer my other question.
That's because I don't have an answer, it wasn't me who said "hopefully" :) As long as Linux is considered a first-class platform, I don't really care who's on top.
I said close to 50/50 so it's not false either, if you take in account the fact that there is probably some margin of error anyway.
> Other than better support for Unix command line tools, what makes OS X (or Linux) a better platform for devs than Windows 10
Maybe it depends what kind of developer you are/who you talk to, but most devs I interact with tend to live in the command line (and need proper package management as well) - things that Win10 does not do too well yet.
Enterprise devs make up the largest segment of pro developers and Windows rules the enterprise.
Win 10 also has the Ubuntu command-line now. I've been using it since beta and it's glorious. Macs can't compete with this - they don't even ship with new GNU utils and you'll have to fight with Apple if you want them because updates will break your setup.
Meanwhile it takes 5 minutes to get a modern Unix command-line in Windows.
I'm glad to say that I'm quite certain that your hopes of a non-Windows world will never be realized.
Some that come to mind: Granular packaging systems with everything developer under the sun in them (including binary and source packages). Better support and easier install of a vast array of developer tools and languages (just one example: git). Much more automatable (eg not every install on windows can be automated. Many require GUI interaction). Containerizable. More powerful filesystems like layered filesystems or content addresses filesystems like ZFS. Cloud orchestration tools work better (puppet, chef, ansible). Tiling window managers to streamline screen work. Much wider choice of code editing environments and code manipulation tools (Windows is much more centric around the offerings of Microsoft). Better interoperation with other tools and filesystems (Linux plays much nicer with windows than windows plays with linux). Less bugs in the APIs and development systems themselves (a result of open source enabling bugfixing independent of a vendor). Better system debug and development tools (eg. strace/dtrace/ktrace). Almost every dev tool included in the distributions (no need to go download some dodgy .exe of tucows or where ever). More example open source code to reference and work with makes coding similar ideas less error prone.
Yes, that's true.
> "Better support and easier install of a vast array of developer tools and languages (just one example: git)."
This falls under Unix command line tools for me, but okay.
> "Much more automatable (eg not every install on windows can be automated. Many require GUI interaction)."
This is really just the same point as the package management one you already mentioned.
> "Containerizable."
Windows Server now has native support for Docker.
> "More powerful filesystems like layered filesystems or content addresses filesystems like ZFS."
Linux's support for ZFS isn't exactly a strong point. Perhaps you had OS X in mind? In any case NTFS is a fairly decent file system, I don't really see it as a weak point for Windows.
> "Cloud orchestration tools work better (puppet, chef, ansible)."
Automating Windows configuration is easily done through PowerShell. I know that Chef and Ansible both use PowerShell to get their Windows support. I'd suggest taking a look at Desired State Configuration if you're unfamiliar with how these tools utilise the existing infrastructure on Windows.
https://msdn.microsoft.com/en-us/powershell/dsc/overview
> "Tiling window managers to streamline screen work."
In my experience, tiling window managers are nice if you've got a keyboard-heavy workflow, but not that much more efficient when you switching around GUI apps. Windows has some basic tiling windows shortcuts built in, plus it now has virtual desktops built in and shortcuts to switch between them, so I don't feel like I'm missing out on much.
> "Much wider choice of code editing environments and code manipulation tools (Windows is much more centric around the offerings of Microsoft)."
Which code editing environments are you thinking of that you like that aren't also available on Windows? As for the MS tools, if you can show me a better IDE than VS on any platform then I'll be impressed.
> "Better interoperation with other tools and filesystems (Linux plays much nicer with windows than windows plays with linux)."
The upcoming Linux subsystem for Windows 10 should go a long way in addressing that.
> "Less bugs in the APIs and development systems themselves (a result of open source enabling bugfixing independent of a vendor)."
Hmm, I don't think you can back that up. Let's put it like this, I've tried Linux multiple times, but I always come back to Windows, and generally speaking that's because of bugs I've found in Linux or Linux software. I have far fewer problems with Windows. Can you share some of the problems you've had with Windows?
> "Better system debug and development tools (eg. strace/dtrace/ktrace)."
Sure, I'll admit these tools are probably better than the Windows equivalents.
> "Almost every dev tool included in the distributions (no need to go download some dodgy .exe of tucows or where ever)."
You don't need to get dodgy dev tools, there are plenty of useful dev tools from Microsoft and other well known software companies.
> "More example open source code to reference and work with makes coding similar ideas less error prone."
Are you familiar with MSDN? If you knew how easy Microsoft makes it to become a proficient Windows coder, you wouldn't be saying that.
To be fair to you, there's one advantage of Unix OSes I think you missed, and that better networking tools (such as firewall software).
The survey covered a whole ~40,000 people out of the ~11 million pro developers in the world. That's nothing.
The numbers also don't jive at all with the empirical evidence. Walk into any small, medium or large business IT shop that employs programmers and you'll find Windows more than any other OS. If they're running Linux it's in a VM on Windows.
Macs are still extremely rare for anyone outside of Mobile developers.
What do you think Enterprise devs use? Not Macs...
...in your very uninformed opinion...