How To Get Experience Working With Large Datasets
bigfastblog.com
bigfastblog.com
https://github.com/rgarcia/generatedata http://smooth-frost-5744.herokuapp.com/
http://aws.amazon.com/datasets
http://www.kdnuggets.com/datasets/index.html
infochimps, datamarket.com,
reddit/r/opendata and datasets,
http://thedatahub.org/dataset,
NYT and Guardian,
UCI machine learning repository,
http://trec.nist.gov/data.html
https://sqlazureservices.com/browse/Data (expired Cert, but i'm reasonably sure it's a legit MS site)
Random data is god for the how and the with what (to an extent), but not for the when and the why.
Here are CSVs for several million names, addresses and voting patterns for the great state of Ohio.
http://www2.sos.state.oh.us/pls/voter/f?p=111:1:322955311706...
My personal experience is that I have no problem understanding data mining and machine learning in theory, but translating that in practice on my own is quite another thing. Is the market open for would-be data scientists with very little experience?
If you're not load testing with your real actual application code and real application load I wouldn't even bother testing. The numbers will be so misleading that it's mostly pointless.
What does it matter if your cassandra install can do 500,000 writes per second if your real app exhibits lock contention issues that brings that number down to 5,000 per second, or latency issues that bring the number down to 50,000.
Since you should be performance testing with real application code and load you'll need to add two things to your code:
1) Code to record the load (logs can work great for this) 2) Code to playback the load at a multiple
Then I'd add the parameters you want to tune for to your testing code and use a genetic algorithm to tune the parameters for your cassandra install.
So, yes, real load testing always involves custom code. If you're just looking for numbers to impress management then use whatever because it's not going to correlate to anything.