I.B.M. Unveils Real-Time Software to Find Trends in Vast Data Sets
nytimes.com
nytimes.com
Quick summary: It's for doing complex event processing of large streams of data. Think financial data or any other such sort of never ending flow of massive amounts of data that you would need to process very quickly to identify patterns and for other analysis such as determining when you hit certain thresholds or trends.
It's supposed to be very flexible so that it can act on a large amount of unstructured data as well as very scalable. It defines its own processing language too.
Any ideas how they compare? System S and kdb+ sound very similar, both offering a solution to development centered around streaming datasets with extremely low latencies, both with their own programming language.
Maybe I'm missing something fundamental (the article was vague) but I just don't see anything really new here. I read it as marketing.
For the technical research underlying it, see http://domino.research.ibm.com/comm/research_projects.nsf/pa...
Factor out those two buzzwords, and you have an article about IBM releasing some software that analyzes data really fast and helps you make decisions based on that data.
If newspapers are in such dire financial straits they really need to start charging for stuff like this.
What does "real-time" mean?
It means that data is processed at the same rate that it is produced. (Or as Wikipedia says, "originally it referred to a simulation that proceeded at a rate that matched that of the real process it was simulating.") A side effect of real-time is that RAM can be used for buffering rather than using disk for storage.
And "stream" means "the Cell marketing team encouraged us to use the word 'stream' a lot."
Wikipedia defines a stream as "a succession of data elements made available over time" and stream processing as "the quasi-continuous flow of data which is processed in a dataflow programming language as soon as the program state meets the starting condition of the stream". I don't think System S even runs on Cell.
While "real-time" and "stream" are too frequently used as buzzwords these days, IBM is actually using them correctly in this case.
"a succession of data elements made available over time" and stream processing as "the quasi-continuous flow of data which is processed in...
Basically automated batch processing where the "chunkiness" lies beneath a certain threshold of perception.
While "real-time" and "stream" are too frequently used as buzzwords these days, IBM is actually using them correctly in this case.
How can you tell? If you have a better link, with real information, post it!
a processor could make a hard commitment to have an answer in 3ms or whatever.
this is hard to guarantee on linux or whatever, because you never know when you'll get swapped out or whatever.