Show HN: Sense - A New Cloud Platform for Data Science and Big Data Analytics
senseplatform.com
senseplatform.com
Sense supports R, Python, JavaScript and SQL out of the box, but is fully extensible to new languages and tools:
https://github.com/SensePlatform/sense-engine
We have Julia, Hive, and Spark engines in development.
Do you have full support for numpy, scipy, matplotlib, pymc, and pandas?
E.g. would it be possible to run a query against Red Shift (step 1), then cleanse it in Python (step 2), run R scripts over the Python output (step 3) then dump the results back to Red Shift (step 4)?
If I then decide I need to change the Red Shift query (step 1), can I then re-run the whole pipeline?
Munging data between different tools is what I seem to spend most of my time doing, so anything that helps that would be a big productivity boost.
You're probably looking for something smoother, though. We definitely intend to have a good solution for workflows like you describe in the future.
You don't mention anything about RAM in your pricing - what are the restrictions? And what about I/O and storage?
This has some really interesting adjacencies to a project that we currently have in limited beta and getting to roll out widely very soon. I'd love to chat about some ideas I have to work together that could work out really nicely for both of us. If you're interested drop me a line: jason@applieddatalabs.com
Since all of your team are very high profile ( Stanford, Harvard) I am wondering how much all have kept to ground work after rising to such level ? Thanks for your answer.
ps: I am hopeful for Stanford MBA admission.
How well does the distributed filesystem perform and what size data sets an it handle?
How quickly can you ramp up 10, 100 or 1000 cores?
Improved performance in these areas are the big things that would get our group to adopt a new platform.
The distributed filesystem is meant for easily sharing code and medium sized data across containers. In the cloud, it is best to to use S3 directly for large data and local disks for high IO tasks. For on premise deployments, there are more options.
I'd be interested to hear about your use case. Feel free to drop me a line at tristan@senseplatform.com
We're not trying to replicate GitHub's features. The core of Sense is a better way to work with data: the compute infrastructure, engines, and analytics workflow. Advanced users using git will likely use Github in addition to Sense.
There is a difference though. In our experience, we've found that the notebook style development, with code inline, is awkward when doing serious analytics. It's harder to use version control, editors, etc. We have opted for the dual pane experience common in R and Matlab. The output however can be rich and interactive just like an IPython notebook and is always saved.