The Data Science of MathOverflow
blog.wolfram.com
blog.wolfram.com
I can not imagine actually using it with data and more complicated programs. That must be... really tedious.
It's great for solving equations and such, but I seriously wonder why anyone would do "data science" with it.
Maybe when matlab and mathematica were first created the existing dynamic languages were not very good?
It had less to do with the languages and more to do with the libraries. Would you use Python over Matlab if numpy and scikit learn did not exist?
Even with Python and R's mathematical ecosystem, they don't replicate the sheer breadth and depth of specialized tools like Mathematica and MATLAB.
I agree the front end struggles if you have a large dynamic object with a lot of data, it has been 32bit for a long time.
I found this just now http://www.wolfram.com/language/fast-introduction-for-progra...
And that page in particular is a good example of how not to write documentation for a language/environment.
Table[x^2, {x, 10}]
The page before introduced lists. And there was no mention of lists being able to do magic things like spanning values. I think that line up there makes a table and somehow that magic list goes from 1 to 10 ...
There is just too much hidden there. It is a poor introduction.
I have read this guide before ... and each time I shake my head and wonder why anyone would bother trying to get through the opaque/hidden syntax when there are way better choices of languages.
That's not true.
Statistic can analyze data in small quantities you can read up on it with nonparametric statistic. At least half of nonparametric statistic deal with small data (mostly through use of ranking). With Bayesian stat you can just assign a prior distribution.
I think the best way to describe statistic is that it uses data to infer the population. A subset data via sampling and infer a statistic (mean, median, whatever) about the population at large.
More often than not I'm just surprise as to why Data Science is so big when seemingly it seems like statistic does every freaking thing with data. From how to correctly sampling, designing experiment to collect the correct data to answer a hypothesis, making sure the data aren't bias, etc.. It deals with designing the data from inception to end, including either collecting the data or data given to you already. On top of dealing with problematic data such as missing data, imputation, etc..
Not for SEO purposes at any rate.
Some of the drawbacks are that it is closed source, costs money (although cheap as far as this kind of software generally goes), but most importantly the succinct benefits of the language can make it painful to deal with. Yes, I can probably write 3 lines to do something that would take a 1/2 page of Python, but I first need to know which of thousands of functions to use and the eccentricities of the language. I'm sure Wolfram employees are that skilled, but I'm not and will not be anytime soon.
With that being said, I spent a few hours writing a notebook demonstrating key fundamental and scientific formulas in my industry this weekend. It was easy and the resulting code, pictures, and graphs look fantastic. I exported it as a PDF for others to use. Even the console is pretty cool. I think it would be a very popular language and environment if it was free and open source. Another problem is running something in production. My solution is to not even bother and just write the final solution in Julia if I need it. I think Mathematica really makes sense though if you're at a lab where everyone else uses it and can pass around Notebooks. In short I really like having it around, but don't like dealing with licensing issues.
I'm always baffled by the expressiveness of the language. I think the platform is extremely underrated. How is it not the silver bullet we're all looking for?