Thank you Microsoft, and so long
byterot.blogspot.com
byterot.blogspot.com
For me at least, big data is more about a form of non-relational data that updates very quickly, doesn't need joins, needs high availability and can be processed later to derive value out of it.
Relational databases on commodity hardware can easily handle hundreds if not thousands of updates per second these days.
What experience or expertise can you share as to why this doesn't work?
You can bring down trees with dynamite, but just because that "works" doesn't mean it's the optimal solution.
What kinds of problems fit this model? Specifically, why would I want to defer processing if I can process instantaneously?
NoSQL has its place for sure, but is always augmented by a relational database somewhere. The Lambda architecture that's so popular at the moment is cool and all, but I find a simpler scenario where you index data into a relational database before writing it to a document database is good enough. That way I can process immediately, and if needed get the document out of a NoSL database (usually only to display to a user).
A relational database that points at corresponding locations in a document database is useful in ways that document databases on their own are not. Relational databases on their own are useful in ways that document databases are not. I cannot say whether the same is true for document databases, as I've yet to encounter a real problem that can only be solved with a document database.
For example look at something like Google Analytics. As simple as page views with timestamps. If you have high traffic data, you can be writing to your database quite a bit. You could bucket that data per second per unique viewer. Then in your YUI you can select to see a line graph with varying degrees of detail. Per hour, per day, per month, per year.
No way would I want to store this in an SQL database. The processing that happens later are the per time period rollups. Creating buckets of minutes, days, etc per user. This can be done when asked for and cached (slow experience for the user), or scheduled and persisted (expensive - consumes more space). Again I would not want to store this in SQL. All the while I maintain my raw data for further processing later.
These kinds of solutions are great for monitoring data, tracking, any kind of high volume, real time statistical data.
That said, I'd do everything you mention in a relational database. SELECT COUNT(fieldname) FROM {tablename} and all that.
Are you ok with eventual consistency as a price for partition tolerance and availability? The data sizes I typically work with allow me to go with availability and consistency at the expense of partition tolerance. And that's ok because my data replicates, and my response time objectives and recovery point objectives are set to something appropriate. This is worth it to me if the benefit is an operational data store that I can query in real time without impacting writes.
Still haven't heard of a scenario which would compel me to go a NoSQL only route (a scenario that isn't Google or Facebook or similar).
You're talking about NoSQL actually. Interestingly, I find that big data is normally synonymous with the unstructured data that updates very frequently, doesn't need joins, etc that you speak of. Structured data in large businesses, such as general ledgers, sales orders, purchase orders, etc. is generally quite fast. The only problem in that area today is getting that data out into complex reporting. It requires a lot of processing (either routine or manual) to get it into an optimized format. Big data came about because companies like FB, Google, Twitter, pretty much any online analytics company, AirBnb, etc. are crunching large sets of behavioral events. (user clicked this button) Outside of those use cases, data just doesn't grow as fast as the marketeers of Big Data claim it is.
Folks who think they have big data problems really need to define the size of their data set -- not the size of the resulting BigDataDB, or the number of servers they have. The two are not the same.
What I mean are these efforts to build huge distributed compute clusters that do X by distributing the problem easily, but at huge overhead cost for each individual actually useful computation.
The distribution, tracking of the compute tasks, the coordination of it all and the frameworks to make it easy insert such a tremendous overhead that I suspect most of the problems that are approached this way would be better tackled with a single machine and some smart old-fashioned algorithm development and developer elbow grease.
Of course simple economics might decide if 20 machines doing something are cheaper than 1 developer working an extra 3 or 4 months so 1 machine can do the same thing resolves it. But I suspect that level of analysis just isn't done and the result is to just throw hardware at the problem.
I've had to remind my developers in the past that we're working with machines that are literally capable of doing billions of operations per second and can do something like this https://www.youtube.com/watch?v=MJfceF0syK8 in realtime.
Performance on individual machines is way more than sufficient to handle even pretty large datasets in the TBs, but we hamstring them with layers of garbage way too often.
One unfair benchmark I've heard before is "if 1 machine can do X in k time using this distributed framework, then 100 machines can do X in q<<k time using the same framework" but it's like also true that the same 1 machine could do it w time with w within spitting distance of q.
99% of startups and in house business apps will never need to scale at a level that an RDBMS can't support. We should all be so lucky to have the success that brings these horizontal scalability problems.
Also, I hope he lets us know what kind of incredible data he end up mining from his mouse movement logs. Surely every website and app on planet Earth has completely exhausted basic solutions like A/B testing and usability testing to make their apps better. The only solution now is to embark on a data mining expedition so that artificial intelligence will tell us how to make better software.
A bunch of statements that have little or nothing to do with the requirements for 99% of projects, then somehow links this to being the end of Microsoft? Wait, did I miss a step? And then he concludes with "I will carry on writing C#, ASP.NET Web API and read or write from SQL Server".
Totally bizarre.
You're absolutely right about Microsoft being behind the curve. There's a pattern of denial when it comes to new technologies, and a total ineptitude at spotting new trends and adapting to them. I watched with envy for years as all the rails kids played with their new toys until Microsoft got its shit together and came up with an MVC solution.
The whole world is using flash? Lets come up with a knock-off 5 years too late. Interactivity is a big deal? Let's try to convince them that webforms is MVC until its too late. Who needs javascript anyways? Distrubuted software you say? Let's not jump on the REST bandwagon, lets make WSDL defined web services and make a shitty API on top! Why have an ORM solution when putting all of your business logic in SQL is SO EASY with microsoft? See, I can drag and drop a database connection from Visual Studio!
I love C# as much as the next person. It's a great language. SQL Server is a pretty good relational database. But if you choose to devote all of your energy to Microsoft development and don't learn anything else, you have nobody to blame expect for yourself when you find your skill-set behind the curve.
.Net developers would never agree about the big-data part, because to be honest they haven't invested time into understanding what data science is. There are so many IT shops developing on the Microsoft stack that make less than say 50 million a year that have tons of data (billions of records they say!). Maybe they make CRM software, or school software, or an inventory system for a small grocery store. They have tons of test scores, purchase records, waste numbers, demographic info, yet if you ask them to develop something that gives you insight into what factors influence student performance, or to visualize perchasing trends, or to predict sales numbers so that the company doesn't over purchase perishables, you'll get a blank stare. The idea that data can be transformed into useful information, and that this isn't simply a matter of CRUD hasn't occurred to them. They probably don't even replicate their db for their half ass attempt at reporting, so they wouldn't begin to understand what something like hadoop even does.
You simply CANNOT explain to a .net zealot what big data is. They just don't get it. "It's stored in the database, and its a few terabytes...how big does it need to get? We aren't Google afterall!"
While the Windows/.NET stacks may not provide as many pre-built solutions to handle 'big data' (a term I'm growing to hate), I have worked with tons of competent developers who have been able to build solutions to answer exactly the kinds of questions you propose.
I've seen it done with SQL Server Reporting Services, Analysis Services and even straight custom C# code. I've seen it done with Hadoop + streaming API, StreamInsight (CEP rather than post processing) and custom Workflows in WF.
As for your other statements that are just built to jump on the 'hate Microsoft' bandwagon, just because a single large company doesn't iterate fast enough to keep up with the rest of the OSS landscape doesn't mean that .NET didn't have options available to it. Nancy provides great REST capabilities and has existed since 2010. Devs were building REST services with Microsoft's ASP.NET MVC during the same period. Others were using WCF REST (not great, but usable) at the same point as well. WCF Data Services was another option.
ORMs.. Theres the ever-expanding Entity Framework (which I don't love), NHibernate and many many others.
Lets not just jump in and bash MS because you've only worked with bad developers using their stuff.
"Big data"? Those are all just everyday, common, ordinary, vanilla-pure, old-fashioned exercises in statistics with maybe a little optimization thrown in, that is, the 'mathematical sciences' or, if you will, with a lot of overlap of 'operations research'.
But long the standard remarks were that "operations research is dead" and "statistics died long ago except for some of biomedical statistics", e.g., for doing FDA approved 'trials'.
Generally the amount of relevant data available and needed for such work is not very 'big'; yes, maybe the relevant data comes from a database of 10 TB, but a simple SQL query should be able to extract into, say, just a simple 'flat' file, maybe comma separated values or name-type-value triples, what is relevant.
From the OP
> 4. We need 10x or 100x more data scientists and AI specialists Again not a prediction. We are already seeing this. So perhaps I can sharpen some of my old skills in computer vision.
Ah, come ON! So, all of a sudden college courses in undergraduate, junior level in mathematical statistics -- weak/strong laws of large numbers, Lindeberg-Feller central limit theorem, Neyman-Pearson lemma, Cramer-Rao lower bound, minimum variance, unbiased estimation, confidence intervals, hypothesis testing, experimental design and analysis of variance, etc. are packed? I doubt it! And course in optimization -- unconstrained sufficient conditions, roles for convexity, Kuhn-Tucker conditions, the simplex algorithm, Lagrange multipliers and Lagrangian relaxation, nonlinear duality, the simplex algorithm and how to use it to solve nonlinear/integer problems, integer linear programming, etc. -- is also getting going, say, if only for maximum likelihood estimation in 'machine learning'? Why do I doubt it?
Guys, yes, for much of anything practical in statistics or optimization will require computing and software to do the data manipulations. Alas, that does not mean that .NET or computer science is what is central to the work. I may be that all that is needed from computing is SAS, R, some open source statistics/optimization code, etc. .NET has a big role in computing, but it's not intended for statistics/optimization.
Who the heck put the $10 into the hype machine?
>> You simply CANNOT explain to a .net zealot what big data is. They just don't get it
Could have not said it better.
We are looking at scaled DB backend using column stores, multiple compute instances feed through message queues running on Amazon or private cloud, and a mostly Javascript / HTML5 front-end. It is all tied together with C# and so far it doesn't look like it will be too bad.
Not surprisingly, the stickiest part of the puzzle is... Microsoft Excel, specifically Pivot Tables. Weird but true.
I thought it was just the atmosphere of my university, or maybe realities of the job market in my country[1]. But reading this made me wonder id there are really two parallel universes in computer-land.
[1] Based on the name of OP, I am 90% sure we have the same nationality.
Most problems are NOT big problems.
Data is getting bigger, sure. Movies in HD, umpteen-megapixel cameras, yadda yadda in the consumer space, but hardware is getting bigger too, for the same price. You can easily buy a terabyte drive now. My first HD was 20M and I thought I'd never need anywhere near that much!
But in the corporate space - where is this data actually coming from? If you have a company selling widgets, then by how many orders of magnitude do your sales have to increase before you have to worry about "big data"? Most companies could probably handle a tenfold increase in sales on their existing systems, even if those systems haven't been updated in a few years. Bread and butter software work is not going to change in the foreseeable future, it'll just have a different set of buzzwords.
In business I see lots of folks who do this: "Well, if I do what XYZ does, I'll be successful just like they are!" So they do what XYZ does without understanding the WHY part of what XYZ does, and invariable either fail, or succeed and making a very poor copy.
As problem spaces expand beyond our current hardwares ability, we start distributing it until a new wave of equipment becomes available and we start labeling our predecessors idiots for having ever built something so complex.
Remember "Can't hurt", not quite true. I am shifting focus, so I am spending time that could be spending on learning and mastering, blogging, speaking Microsoft technologies. If I am wrong, it will hurt me.
There was a fight? I must've missed it.
MS has two problems with "big data" (and let's just pretend lots of people have big data problems - I've seen people deploy 100-node Hadoop clusters to deal with 30M rows/month of data because they think Hadoop is some magic sauce.)
First, they want to squeeze as much money out of enterprise customers as possible. Look at the new SQL Server pricing and limitations, where the SQL team is doing things they promised they wouldn't and previously mocked Oracle for doing (charging by core). Their in-memory solution (Hekaton) comes a bit late, and only for Enterprise edition.
Second, they generally want to deliver solutions that the majority of their customers can figure out. Again, look at Hekaton. From what I've read, it'll have extremely high compatibility with T-SQL. Shipping a limited release that only had a small subset of SQL, even though it'd be useful for certain apps, probably was never a serious consideration. Hell, look at C# and how their customers are begging them to please not innovate too much, since learning is hard.
There's also the Azure push. SQL Azure has federations, making it easy to shard a traditional SQL schema across many physical instances. They also promote Hadoop on Azure. (As I understand, Hadoop's a dead-end; even inside Google, MapReduce was toasted by Dremel, right?)
If Microsoft just added Hadoop-style functionality to one of their server products, it would not be friendly. People capable of writing an OK SQL query for a report can't necessarily format that same query in an efficient map-reduce style.
And really, since this need is far less than it's played out to be, MS is probably just fine pushing and profiting off their traditional solutions. Azure helps them secure a few leading-edge needs, and eventually they'll roll out an easy-to-use commercial solution.
MS needs innovation and solid plan; not jumping on the band wagon of next hot topic.
"Cannot say the same thing for middleware technologies, such as BizTalk or NServiceBus. Databases? Out of question."
I'm curious about his view that middleware is not here to stay. Surely the multiple vendor cloud trend that enterprises are following (Google + salesforce + Microsoft + oracle) - because no one vendor does everything - means middleware is going to be more prevalent in the future.
Thoughts?
It basically a message based event driven transport layer. By its very nature it's extremely easy to scale horizontally.
For HA you simply run a distributor process for each logical event on windows cluster and add as many worker nodes as you see fit.
Because NServiceBus uses the store and forward pattern if a NServiceBus process/machine hosting it goes down, you are still guaranteed the eventual delivery of messages when the process/machine is resumed.
Fun fact is that DevDiv produces tons of things that allow .NET devs use existing OSS solutions for things that are not done by Microsoft themselves.