Being able to fiddle with your schema on the fly isn't necessarily a good thing. Brainstorm, plan, execute. Or just slap another field on the end of the document. Whichever sounds better for the long term health of your data!
Being able to fiddle with your schema on the fly isn't necessarily a good thing. Brainstorm, plan, execute. Or just slap another field on the end of the document. Whichever sounds better for the long term health of your data!
However, once the app has shaped up, I make it a priority to harden up the data model and move the core data set to a relational database (Postgres is the likely choice now). One nice thing with this migration is I can do a lot of clean up and normalization with the originally unstructured data set.
I've found that when starting a new project I get paralyzed with coming up with the perfect data model that will also scale in the future, so I'd rather punt, take the tradeoffs that come with schemaless DBs, and keep a plan to migrate to something like Postgres in the backlog. Holding on to our tools too tightly is IMO how projects get crippled.
SQL schemas are not expressive enough to describe the data types which I want to store in the database. The argument SQL vs NoSQL becomes moot when any database is degraded to essentially pure blob storage. What matters to me is ease of use, administration/maintenance burden, and of course speed. How strictly the database verifies the schema isn't even on the picture.
For my own edification, can you expand on this?
Personally I like foreign keys, a lot, and couldnt live without them. I dont think Haskell's type system would be a replacement to that, and would rather use a hybrid of those two.
data User = Anonymous | Registered UserData
data UserData = UserData
{ userName :: UserName
, registredAt :: UTCTime
}
Foreign keys is something that can't be easily verified in Haskell code. So that's something where a proper database still is useful.I'm incredibly skeptical that SQL doesn't support your data types but JSON and RethinkDB do. A quick glance at the docs for PostgreSQL and RethinkDB, and the RethinkDB types are a small subset of the PostgreSQL types. It sounds to me like you just don't want to break out nested objects into separate tables.
It is not realistic to create a separate table for each type. You'd end up with hundreds of tables (and would have to replicate the types in two places). You can only break down the types so much, at some point you'll want to have columns with your own special types which you want to treat as primitives.
I know of a "Diablo clone" online RPG that uses SQL Server to store the character blobs for the people playing the game. They have a logically partitioned table, GUIDs for IDs (used for the partitioning), and just a big ass blob (somewhere between 32k and 128k) assigned to that id. I'm sure there are other items associated with the row, but that's the gist of it.
They handle all of the funky stuff in the application. I don't see why that wouldn't work in a document style database, and I'm not entirely sure what the design decision was to use SQL Server.
I'd like to see if they handle items in a relational manner or all of that is in the application. I wish more companies were a bit more open about their schemas and design decisions. I have done a lot of things a lot of different ways, and I'm always curious about how others have tackled similar problems.
What did you think about Domino replication?
The way Domino works at a high level is that you have a "data" and "design" template applied to a database. A database would be a single file on disk, and it's something like a table (or a set of related tables) and various views attached to that data.
We had an internal Domino server with a set of internally focused 'design' templates applied to them. We used the notes clients, and our analysts had their workflow applications there. The design of these databases was focused on serving content to the notes clients - not to a web browser.
We replicated the data from those databases to our staging and development machines. Those staging and development machines had their own designs - which was the web-centric, customer focused design. There was a minimum amount of information available to you in the notes clients - data at this point was meant to be seen in a browser.
Those staging machines pushed the data on to our final production cluster. Every 15 minutes something was pushed. We'd go from internal to dev/stage 15 after the hour. 30 minutes after the hour the data from stage would go to production.
We used normal build/release tools control the flow of design template replicas making their way out of their playgrounds, but sometimes it happened. I say it's brittle because if you accidentally push design to the wrong place you could grenade a lot of stuff. More than once we accidentally flipped the switch to push design changes to development.
Here's a link to a presentation of the last thing I was a part of in the Lotus community. I'm pretty proud of the site we put together. It had a pretty nice feature set for the time period. This was put together by a coworker. https://www.youtube.com/watch?list=PL6D93ED85F970BAE6&v=v9IU...
Data integrity is less of a technology problem than people make out.