HSBC moves from 65 relational databases into one global MongoDB database
diginomica.com
diginomica.com
This is ONE system at HSBC - out of literally HUNDREDS (maybe thousands) of applications. It's not as if HSBC is moving ALL of its applications to MongoDB. HSBC doesn't have a single tech stack. They have thousands of IT employees all using different tech, in different parts of the world, for different departments. This might be as inconsequential as the system that catalogues security camera feed URLs - or maybe the one that monitors remote employee company mobile data usage - who the F knows, because the article gives us zero information.
Account Management, Credit Cards, Mortgages, Security, Asset Management, HR, Legal, Compliance, Regulatory, Risk Management, Trading, Operations, Building management, Payroll, ATM comms, Clearance/settlement, Website, a million different reporting engines, etc, etc, etc... and then each region usually has its own application for each function - literally hundreds of tech stacks in every tech you can imagine from DB2 Mainframe to Oracle to MongoDB. All banks are like this.
This article is just vague PR and is just referring to one single group consolidating their regional instances. It does not deserve HN attention.
So in no way was the OP's article as catastrophic as it initially sounded.
> Historically, (...) HSBC did have an application core program environment, which had most of an application's core functionality. But it couldn't have a single programme environment running for all the countries, due to the differences in data models and databases.
> Another benefit is to use the same database for global data analytics and reporting. We don't need to translate into another data model or another database to run the analytics and reporting from that particular data.
So this "application core program" failed to provide an unified solution to their data warehousing needs, and now they believe they can solve this problem by unifying 65 databases into one MongoDB instance, "taking advantage" of its schema-less model.
It looks more like a global storage service than anything else to me.
Solution 1: refactor schema
Solution 2: abandon schema
The truth is that you can’t really abandon schema, it’s just spread through the code base now, and any violation of the schema will blow up not when you write malformed data but when you read it later.
Anyone know any good tools to manage the evolution in schemas to address this problem?
You still need something higher level to oversee the schema, but I didn’t look closely - my schema is small enough I can keep track of it in my head.
The community version (open-source/free) does have a few important limitations that may or may not matter to you.
As others have said, you can either validate the schema at WRITE time (RDBMS) or at READ time (NoSQL). Personally I hate the during READ time because then you have all sort of data integrity fun to work through... In either case, I am sure they have a ton of reference documents that they need to store, so easier to manage 1 Mongo than 64 different weird databases.
I've never wanted to short a company more in my life.
For any use case, SQLite is one of the best ways to persist structured data to local disk.
Its a little scary how much it can actually handle, and how lazy and sub optimal it is to reach for a "real" sql server before you need it.
Having the database as a simple file right next to yoyr application is incredibly convenient not to mention how brain dead things like backups are to grok.
VoltDB: https://www.voltdb.com/blog/2016/09/nosql-vs-newsql-whats-di...
Spanner: https://cloud.google.com/blog/products/gcp/from-nosql-to-new...
I'm fairly sure that this was driven by a similar sales push.
Mongo has always been a sales and marketing company first and technology company second.
I'd love to see a tear down of their sales and marketing strategy coz its clearly top notch.
Google Trends comparing "NewSQL" with "NoSQL" https://trends.google.com/trends/explore?date=all&geo=US&q=N... (Peaking with its introduction in 2011).
It doesn't seem like a more fashionable term than NoSQL.
"Local requirements for each country will be built into the application, but there's no need to maintain separate data models or separate databases anymore. We could easily design the global data model and database using the MongoDB JSON schema model. That brings data from all operating countries into one database and the application can run on just one database. Which is a lot of reduction in resource and maintenance cost. "
Are there any other schema-less databases out there that is better than MongoDB? If not I believe it's time for someone to build it.
Yes, PostgreSQL supports JSON and won't lose your data: https://www.postgresql.org/docs/current/datatype-json.html
Benchmarks:
* https://portavita.github.io/2018-10-31-blog_A_JSON_use_case_...
* https://www.postgresql.eu/events/fosdem2018/sessions/session...
Summary - PostgreSQL ● PostgreSQL has poor performance out of the box ○ Requires a decent amount of tuning to get good performance out of it ● Does not scale well with large number of connections ○ pgBouncer is a must ● Combines ACID compliance with schemaless JSON ● Queries not really intuitive
Summary - MongoDB ● MongoDB has decent performance out of the box. ● Unstable throughput and latency ● Scale well with large number of connections ● Strong horizontal scalability ● Throughput bug is annoying ● MongoDB rolling upgrades are ridiculously easy ● Developer friendly - easy to use!
ISHYGDDT
Oh and MongoDB Compass is a dumpster fire. Query takes too long? Too bad it will time out with no option to let it complete. Also I get to write my query in JSON in compass then I have to convert that to native Python objects if I want to use it from there. With SQL I copy my query from datagrip, inject my parameters and call it a day.
I forgot the best bit, if your query has a subtitle mistake most often you simply get no records back where a similar mistake in SQL throws a helpful exception.
This is actually a valid point. They should start distributing typical configs for typical AWS-alike machines, current default config is for some underpowered machine from 90s..
[1] https://www.theguardian.com/business/2012/dec/11/hsbc-bank-u...
Do you have any source that substantiates your assertion?
These mega-big banks and their practices, ranging from LIBOR manipulation, scamming retail banking practices, scamming their own IB/FX customers..). And they get away with it.
As much as I dislike religion (I love the notion of Faith, but I dislike the self-identified 'representatives of gods'), I believe that the most ethical type of Banking is the Islamic Banking (again, not about the religion)(gods that give cancer to little kids to "test their faith" are not my type of gods).
Uncontrolled data loss is most assuredly not.
In short order, I cut the card up. Still get sad emails (which they provide no way to stop) begging me to use it.
Also, interviewed at MongoDB once. Couldn't shake the feeling of a cult, where everyone was careful not to speak of the product's serious limitations.
My parents switched from Lloyds (UK) to HSBC and are looking to switch again.
You can't have multiple devices logged in. I.e. you can log in just fine on your "security key" device (say a phone), but if you want to login on your iPad to check something, you have to get a code from the app on your phone. Absolutely pointless.
- Banks have to put these double-checking and reporting systems in place because the law mandates them to. No amount of ACID compliance in your database is freeing you from that.
- These requirements are not stupid either: assuming you had a perfect system with perfect ACID compliance and 0 bugs ever (which we all know is impossible), a random hardware failure (or cosmic ray) can flip a bit somewhere and you are screwed.
- Your comment reads as extremely condescending, which is never good to get your point across.
- Not only that, but you dismissed the work of some of the best people money can buy in two sentences with 0 rationale. Next time you think you know better than an entire extremely well funded industry, think again, and again, and again. If you are still convinced you know better, please open shop because if you are right you'll not just make yourself the next self-made billionaire, but also probably improve our lives in the process.
Every time I read an article here I am put off from participating further because the comments are usually filled with toxic spew from peacock fan boys Who think they know better.
You have my up vote.
Basically each Banking transaction is an immutable log, which can easily reside in a DB with DB transactions feature. They are not related.
Can you please elaborate in technical terms what you mean?
Think of a account with $0 in it.
1. A -$25 event comes in
2. A +$50 event comes in
3. A -$25 event comes in
The first event would fail in a transactional/real time system, but in my bank it doesn't. This tells me the balance is only calculated every once in a while. The balance you get in a app is only a estimate.
This makes sense when you think how slow bank transfers and using paper checks is. At some point it is decided to run the events "for real" and that's where overdrafts and the like are applied.
Don't think of them as transactions, instead it's a transaction request that can be rejected.
In your example, a “good” bank would order the transactions: +$50, -$25, -$25 resulting in a zero balance with no penalties. What Wells Fargo did could result in the transactions being ordered as: -$25 (overdraft), -$25 (overdraft), +$50.
Each overdraft could charge a fee up to $35, leaving the account holder with a -$70 balance.
[1] https://www.latimes.com/nation/la-fi-court-bank-overdraft-fe...
Additionally, the number of folks that ask for spreadsheets (not encrypted) with your SSN... and then proceed to email that data around (sometimes via Gmail!)
1. Client initiates the transfer request.
2. Bank A decreases the client's balance.
3. Bank A sends the request to transfer money to bank B.
Now, I guess you would expect that steps 2 and 3 should form part of a DB transaction. This however does not really work because it is unclear what conclusion we should draw from step 3 failing. On one hand, it is possible that the request failed on its way to bank B, in which case step 2 should not be applied. On the other hand, it is possible that the request successfully reached bank B but the response got lost on its way back. In this case step 2 should be applied.
Therefore, not treating steps 2 and 3 as a single transaction makes sense if you want to take a conservative approach that does not allow for accidental transfer of infinite amounts of money to bank B in scenarios when the response fails to come back from an otherwise successful transfer.
I'm not informed enough to know if it'll affect the share price for a bank but if a tech company did onboard mongo that'd be a major red flag for me.
I would expect downtime from RDBMSS and most banks have scheduled downtime nearly every weekend they say when I log in.
Which is not his argument. His argument is "X uses Y, and X doesn't have problems with Y, so Y is good enough for X at least". Which is a completely valid thing to say.
It used to be lots of fun and highly optimal to write all sorts of procedures and triggers because it was the only fast and reliable way to do things, especially when your in-database-code was essential the only gatekeeper against a multiple owner database schema.
The idea that you have one database schema that you are going to share with different applications seems bonkers to me at this point in time. Compute and storage resources are far cheaper than development, DBA and potential stagnation that a schema that doesn't have a single owner brings.
Could be a great way to wash some crimes off the books, "oh our DB failed and our backups were stored in the database as well, oopsie".
In less than 5 years there will be an EU directive on how financial institutions need to have immutable db architectures with full provenance.
I'm a former Entreprise Architect from Banking Sector in Europe.
To be clear , this will never happen.
There has never been any directive that has had a "concrete" impact on IT Architecture of Financial Institutions.
Yet there has been dozens of regulations that urged banks to "simplify" their IT Systems.
90% of IT Staff are Baby Boomers with no background in Systems Design or Architecture , thus when new regulation comes in it follow this scenario 100% of the time :
- Find a vendor that sell a software that promise compliance with new regulation
- Find an integrator that promise integrations within the deadline
- Integrate the vendors software in Banks legacy stack of 3000+ monolithic apps
- Send report to regulator saying they have "redesigned" their architecture and made "investment" in order to take in account and that new regulation
Best examples of this is "PSD2" which has been the biggest fiasco of the industry , has of the last year only 18% of banks complied with the regulation.
France and UK said they would not "fine" anyone, because banks "aren't ready" and made "considerable" investment in it.
Regardless of what HN thinks , you won't solve Financial Institutions Multi Decade Legacy with a single Directive that would suddenly force them to use "ImMuTABLe Db ARChITectuRES"
It Directives have never worked and will never work.
The only way to enforce anything would be to remove their IT systems completely and have them use APIs provided by the regulators , and have the regulator become the sole provider of "Financial System".
They wont let that happen.
Banks have extremely strict rules when it comes to system "access".
When it comes to "System Design" they have very little.
This are two separates topic.
Isn't that's because of SOX compliance requirements?
So, the point above is that "a new set of requirements" could be added regarding "data store software integrity" (though probably named better) if it turns out to be needed. :)
Bank always had very strict rules , SOX just force them to formalize those rules/process.
Per say , prior to SOX they would not conceal the review of the logs , now with SOX they will will edit a PDF that say "We have review logs for Apps X and consider no suspicious activity had occur".
Apart from that , things didn't change much.
(I'm not trying to claim you don't somehow know this, I'm genuinely interested to read about this stuff!)
I have not seen regulators require a specific technology, but I have certainly seen them questioning technology choices. There may be some truth to echopom's claim (made regarding European regulators, nott US regulators) that regulators can be bamboozled with meaningless claims to have "redesigned" a system and "invested" in it... I couldn't say because both of the US banks I have worked for have taken even gentle hints from regulators EXTREMELY seriously and would not have attempted to bamboozle them.
I don't know that we'll ever get to 'move fast and break things', but I think that's OK.
- MongoDB: Yes
- Micro Services: Yep
- Some attempt at a grand unifying model: Of course
That's pretty much a 10/10 shitshow as far as I am concerned. Mix in banking regs & regional differences for afterburner on that money furnace.
But, in reality, I doubt that one of those 65 relational databases involves the core customer/account/transaction processing facilities. These are probably (hopefully) still running on an IBM system with proper ACID & uptime guarantees. Anything a live transaction flow (especially credit/debit processing) is going to touch could never be trusted on something like MongoDB in this current reality.
The dev cycle for releases and whatnot is incredibly long due to testing and, lets be honest, fear, in case something breaks but when it comes to battle-tested technology, mainframes and DB2 are right up there.
I do know of some code that's over 40 years old and still runs at the core of one of the big banks...
MongoDB, despite its capabilities, is... an odd choice imo! Not trying to second-guess but if I was in that meeting I would have certainly raised an eyebrow.
Good luck to them.
I expect you're right that this story is purely about using Mongo for downstream systems.
If they are intent on replacing the mainframe then, imo, MongoDB would be a bad choice but who knows... the new, challenger banks don't use mainframes.
Also, Monzo are heavy users of Kubernetes, so again, no mainframe. [0][1]
Starling use AWS and GCP [2]
In addition, the cost of entry of a mainframe system is a high bar when you can spin up a server on Azure in seconds and deploy code to it in a handful more seconds and then scale it to the heavens in yet just a few more seconds... all for less than the hourly cost of the IBM salesman.
I've yet to see any jobs that mention mainframes, Cobol or the like.
[0] - https://monzo.com/blog/2016/09/19/building-a-modern-bank-bac...
[1] - https://www.8bitmen.com/an-insight-into-the-backend-infrastr...
[2] - https://blog.container-solutions.com/starling-how-to-build-a...
The part of the bank that handles core processing is going to be a very small portion of their engineering team. Kube and DevOps ads don’t provide any hints about how much they might rely on DB2, COBOL, etc...
I will be tracking their technical reporting more closely over the next year.
Just curious, but what specifically do you mean? ie. I'd love to keep an eye on this also, but have no idea where to find these types of (public) reports.
What's a good example of this working out well?
I don't see why COBOL, mainframes, or whatnot would be any different. Perhaps the ramp-up time would be a bit higher, but most skilled programmers will manage. I actually wouldn't really mind working with any of that – I just never had the opportunity.
I wouldn't claim they are any good, just mentioning that grads are being trained in COBOL.
The cost of architecting relational model that would be cover for 65 mini databases in a bank will be astronomical. It is far easier to setup no-SQL entry point that would auto-conform for whatever requirements upstream applications might have.
It is a technical win for Mongo, but I don't believe this is an attempt by HSBC to get up to speed with database trends. HSBC are removing internal audit strikes, nothing more than that.
But they already have relational models, right? Can't you port the DDL over to centrally managed instances?
Not to sound pedantic or old or curmudgeonly (because while I'm 45, I was already curmudgeonly and reading Fabian Pascal when I was 28). But c'mon.
All data has a schema, it just takes effort to discover it. A relational database is a set of facts about the world. Lack of effort in finding that schema and normalizing data, "conforming" it... it means effort from very frustrated people who have to clean up the mess later and make it conform.
FWIW SQL can seem like a clunky and dated way to talk about data. It is. But it's not the relational model's fault, it's SQLs's. Going "NoSQL" simply makes data modeling problems worse. It pushes the problem til later, and pushes reasoning about the data into places where it shouldn't have to be. Even more so when the underyling data storage tech is dubious, like Mongo appears to be.
I’m aware you can keep references from one doc to another but it always ends up being a messy affair even at smallish scales, and it feels as if the cognitive burden of managing these relationships ends up falling on the dev, instead of being managed by the DB.
For those of you using MongoDB in large scale large production apps, is this not a problem? Is your underlying business domain really non-relational, or do you manage to comfortably run highly relational models on a doc-based DB? How?
Also in their blogpost [1], about $lookup they say:
We’re still concerned that $lookup can be misused to treat MongoDB like a relational database. But instead of limiting its availability, we’re going to help developers know when its use is appropriate, and when it’s an anti-pattern. In the coming months, we will go beyond the existing documentation to provide clear, strong guidance in this area.
Clearly they have some expertise, I'd love a more in depth look at how they arrived at the decision.
b) Dynamics and SAP are both ridiculously expensive especially since you would be need to also buy additional software e.g. Windows to run it on. Plus of course it's much harder to find specialist engineering talent who know these products.
c) All three are far too slow both on the ingestion side e.g. millions of mutations a second and on the read side e.g. feeding into a real-time ML model for an upsell prediction. MongoDB is specifically designed to cater to both of these technical requirements since you can mutate sub-documents and retrieve the full document extremely quickly.
Classic HN.
I'm not saying I'd choose Mongo specifically for this job (if for anything), but the lack of DB enforced schema might not be the biggest problem here.
Also, I don’t think I’d want my filesystem deciding what qualifies as a valid JPEG.
But yes, I can see the theoretical appeal.
Neither of those sound like great situations to be in.
And just because some parts of a schema need to be “in line” doesn’t mean all of it must be. Maybe I care that “Customer ID” is enforced, but not “Australian Tax File Number” or “Document revision annotations”.
Most of the legacy retail banks regularly take down their online banking for hours at a time to do 'planned maintenance'. I've also heard that this happens to things like payment gateway APIs that aren't directly visible to consumers. Sometimes a Faster Payment might take a suspiciously long time and that's because the API fell over or was down for maintenance.
At least the data integrity is pretty good.
Database alone isn't the deciding factor, how you used it in your application architecture matters the most. No database is perfect but some are more flexible than others.
Are you using ACID transactions?
In fact we replaced Kafka with Change Streams( little over 7k messages per second) and ES with text search (2K queries per second) in MongoDB on 16 CPU core nodes in our cluster.
I don't believe in any benchmark unless it matches the business use case and development practices I use at work. I recommend the same to others.
This is the video if anyone is wondering:
Geniuses!
Edit: In case you missed the story https://www.bbc.co.uk/news/business-18880269
All new tables had to go into dynamo, and if your service begins consuming an old DB2 table, it had the job of migrating that data to dynamo.
The end result of this was applications that half pointed to DB2, half dynamo, confusing new data with old, and tons of bugs relating to losing entries into dynamo tables under outdated keys, not being able to clearly match up customer data, and myriad problems that were never a consideration using a relational database. Needless to say, I pulled all of my assets out of this institution as soon as I had the chance.
Companies like IBM specialise in royally screwing companies through their licensing deals and with their legal enforcement teams. They are nothing like dealing with a normal startup vendor or with someone like AWS. They play hard ball.
So sure you can complain that they should've just stayed with DB2. But usually that means paying tens of millions which then increase each year since they know you won't migrate.
That's why companies take drastic decisions and move to the cloud even if it's technically not the best option.
Granted, there are plenty domains where you can get away with having a blurry schema with the occasional bug creating a new localized data anomaly, but banking, where such things lead to customer trust erosion - I just don't get it.
That "Micro Service Instance" is not a microservice, it's an Enterprise Service Bus (ESB), aka the Egregious Spaghetti Box.
That's aside from the fact that relational data belongs in a relational database.
I really can't imagine it being a single instance, that just does not make any sense.
Ps. I think they never heard of database per service when looking at their microservice graph. Lol
if they've got a decent QA process and they can't get the new system to actually work reliably there's a fair chance this project would get stuck in QA for a year or two, blocked from going live and cancelled, so it wouldn't cause catastrophic data loss, just reallocate a bunch of money from shareholders to contractors / employees / vendors.
> microservice graph
banks have graphs of macroservices.
https://www.warrants.hsbc.com.hk/en/tools/search/ucode/00005...
I'm not sure about their choice of Mongo but the initiative seems like it should be helpful for customers like myself who have accounts in a few countries. We can access them all from one app but the experience is inconsistent at the moment. I'm cautiously optimistic.
[1] https://www.theguardian.com/info/2018/nov/30/bye-bye-mongo-h...
I really have to pivot my business - new name will be „The Microservice-NoSql Consulting Company“ and we will sell tailored versions of the „Microservice-NoSql Strategy“! (Changing names and colors of boxes for 2000$ per day)
I simply don't get it. It doesn't make sense to me, but perhaps they know something we don't.
Or it could be a real-time fraud system where you need every bit of customer data in one payload to feed into an ML model.
You can't do any of that if you are having to do multiple joins across federated databases.
I also hope they’ve configured their mongo correctly since out of the box Config is not appropriate for many use cases
Would really be interested to know how something like Mongo ends up getting seriously considered for something like this, I mean the whole decision process that goes in to selecting something like this.
1) mongoDB does not offer the guarantees required by a bank.
2) Application developers using mongoDB rarely protect themselves against NoSQL injections. Most of them are unsuspecting that they even exist.
e.g.: What is your user id? {$ne: null}... whoops, now every user record is returned.
3) They better make sure they use the Decimal128 type for their currency fields, and make sure that the value is not being casted to float on the client. e.g.: on a JS client, every number is a float.
Unless they've taken the necessary precautions this is a disaster waiting to happen. Hopefully other banks do not do the same.
I don’t think you actually tried it. It returns user who’s id is “{$ne: null}” (string)
It is very infrequently updated, integrity is not at too much risk. In case it's lost, business is not impacted too hard as long as you can find contact details eslewhere.
But the replication story is good, and if they got tons of different schemas, it will make their life easier.
So yes, for transactions it would be bad, but maybe for this particular use case, it's ok.
It is interesting that you can use Mongo and keep the data in a geographical location and still be able to run a global query: https://docs.mongodb.com/manual/tutorial/sharding-segmenting...
First off, how would such an article even come about? Feels like more of a recruiting/marketing piece.
The old story at hsbc was that the ATM software hadn't been updated in twenty years, cause stability....
I remember specific flaws/hacks with bitcoin websites coming down to mongo race conditions/inconsistency issues.
_ It's now a one service environment, one database and one execution path for all the countries. This is made possible because of MongoDB's document model and the ability to map all the different table requirements for each country into a single collection, using sub-documents. Everything is simplified into one collection using country specific identifiers._
How they will address race conditions, eventual consistency and so on is not addressed at all though. I guess they trust Mongo's new transactions will take care of that (at which point, it would be nice to know whether that actually worked and what the performance will be when compared to the old relational system).
"Local requirements for each country will be built into the application, but there's no need to maintain separate data models or separate databases anymore. We could easily design the global data model and database using the MongoDB JSON schema model. That brings data from all operating countries into one database and the application can run on just one database. Which is a lot of reduction in resource and maintenance cost."
Is there any Database that can do this? other than MongoDB? I am hearing about data corruption in MogoDB, I believe that is due to bad, out of the box settings, which I believe they would have mitigated when using in cooperation with MongoDB company (which I am sure they are).
What if... those models map to sql database schema’s? wouldn’t that be magical? Not since 2008.
Whats left is maintaining multiple databases. That sucks indeed.
Sql databases support bson as well these days.
What they probably want (or did) is make a base framework, suitable for all countries. And have each country develop its own stuff on it. No need to use mongo for that though.
One use case i can imagine is that forms and input just change and that old information never will be compatible again with the newer forms. Mongodb serves as a giant more or less queryable data bin. Then one can ask the poor dba to lookup something for a client with “client_id”: xyz. Then return the raw bson/json output. Might be sufficient.
At least HSBC are 'web scale' now!
I don’t think I’ve seen a single HN article complaining about actual data loss. That would be something that would get upvoted immediately.
So what gives? Especially since earlier and older versions of Mongo apparently had far less data stability.
I’ve probably read far more complaints about Postgres in HN articles (difficulty setting up, poor defaults, etc). And Postgres may not even be as popular as Mongo. So what gives?
There was an article on it about it a few years ago.
Did they fix that?
Pretty common these days to just rely on Git and not DDLs.
(Sorry)
Reference: https://www.youtube.com/watch?v=_nVk25ZvTkU
Google has been running its entire infrastructure on a NoSQL database for years.
Perhaps folks need to keep some open mind and see if that works instead of dismissing a new trend?
NoSQL has proven to be hard to get right than was first assumed and at the same time RDBMS’es have improved and adopted many ideas from the NoSQL world faster than expected, often with better results.
What’s no clear from the article is why MongoDB was picked, so maybe there’s something that made it the obvious choice, it’s just hard to see what the might be.
SQL has been around since birth of computing, give NoSQL some time.
"NoSQL has proven to be hard to get right than was first assumed"
Nothing has been proven, again give it a chance, SQL had years of optimization.
Also, no need to downvote if you like SQL or your job depends on it, or that is all you know, just debate with reason and evidence, I've used both and open to change my mind with evidence and case studies which we don't have many yet and folks here predicting doomsday to HBSC without keeping an open mind.
I’m skeptical about this - Google’s entire infrastructure is so large that surely they must be using a wide variety of db technologies?
I took a databases course that did mention that part of the Google search infrastructure runs on a tailor-made NoSQL technology, but it wasn’t really a doc-based approach, it was more akin to a relational DB but with flexible columns (can’t remember the details sorry).
At any rate I don’t feel that this is any vindication for NoSQL, as Google builds much of its underlying tech from scratch and its hardly comparable to the out-of-the-box solutions that the rest of us, even massive corporations like HSBC, have to work with.
"Bigtable development began in 2004[3] and is now used by a number of Google applications, such as web indexing,[4] MapReduce, which is often used for generating and modifying data stored in Bigtable,[5] Google Maps,[6] Google Book Search, "My Search History", Google Earth, Blogger.com, Google Code hosting, YouTube,[7] and Gmail.[8]"
"Bigtable is one of the prototypical examples of a wide column store. It maps two arbitrary string values (row key and column key) and timestamp (hence three-dimensional mapping) into an associated arbitrary byte array. It is not a relational database and can be better defined as a sparse, distributed multi-dimensional sorted map."
Yes, keep an open mind, but please do not use a database that has a history of just losing data in mission critical software.
> there is no need to structure the data into tables and schema does not need to be enforced at the DB level.
There is no need to, just like your car doesn't need safety belts to function.
A database scheme is mostly a safety mechanism as it will prevent you from doing something stupid.
> Google has been running its entire infrastructure on a NoSQL database for years.
Do you have a source for this claim? I would be inclined to believe that the data for my Google account sits in a relational database somewhere.
The other statements are not even well reasoned, you read things like "this is madness" or "they up to disaster", "lol just lol", "they're gone", not explaining why. You would expect better arguments from the folks here.
Some pointed flaws in the design or implementation, and I'd argue with time and money, it will get sorted out, again SQL had years optimization and research.
I'd go further, I'd argue is that NoSQL is more flexible and easier to work with than SQL and I think it is a good thing that large organizations are given it a real try, so we have a large scale case studies, instead of completely dismiss the effort as a failure from the get-go.
You would? Why?
You don't think I should keep that expectation?