Another Book on Data Science – Learn R and Python in Parallel
anotherbookondatascience.com
anotherbookondatascience.com
To be fair, the book advertises showing R and Python code side-by-side. And that’s what it does. But it does it unlike how the languages are most often used in industry.
As a quick example, I saw no tidyverse code, which is essentially the only thing keeping R in the game. Learning R from this book won’t prepare you for writing R in most R shops.
I don’t see the utility in knowing how to do the same thing in both python and R if you’re a beginner. This is even more true if you’re not taking advantage of the strengths/weaknesses of either language.
Instead, just learn one of the languages well, and then learn the other well. Shallow dives in both will make you weak in both.
Unfortunately, 90% of data science content seems to be geared at beginners.
In all benchmarks I've seen data.table is faster than dplyr on all tasks. Curious to see other results.
https://github.com/Rdatatable/data.table/wiki/Benchmarks-:-G...
List functionality within datatable is blazingly fast. Faster than anything else I've seen in python or R.
I think 90% of data science content is for beginners because anything more advanced isn't best described as data science. As soon as you get beyond the initial stages of data analysis (cleaning and processing data), you're doing something best described as some other word (statistics, machine learning, etc.) - although, granted, there isn't much content in these areas if you don't know _exactly_ what you're looking for.
How so?
For more sophisticated linear algebra algorithm, such as SVD, both will use typically LAPACK, and again, both will exhibit essential identical performance.
There is one important difference though: when R is compiled for 64-bit machines, it can only use 64-bit floats! While numpy can support 32 and even (through software emulation) 16 bit floats. This can halve memory usage, which in turn halves cache misses, which results in a significant speed up in cases where 64-bits of precision is not needed.
See the yellow benchmark (matrix multiply). I suspect it's memory-related.
Not an advantage if you ask me - exactly because data.frame is built in, people have been building their own versions (tibble, data.table) instead of improving it. That's how R ended up with 3 different structures that are similar but have inconsistent apis and behaviour.
> lots of domain-specific packages
That's true.
> more consistent interfaces for basic statistics and machine learning models
Can't disagree more - there is no one go-to library for ML in R (like sklearn in Python) and each package has it's own strange interface and implementation.
Python has utility. But R is far superior in its the quality of the packages, their documentation, their ability to behave predictably on a given data type.
I run a machine learning shop. Right now all of the training, application, and data management is handled via R. R is simply superior in too many ways for us to be bothered with python for the scale of work we are doing.
Since we're moving some big applications to keras/ TF we do use python and will be using more in the future. However, for almost all data management, munging, movement visualization, reporting, its an R world.
> ...
> Since we're moving some big applications to keras/ TF we do use python and will be using more in the future.
Not sure if I misunderstood, or you're contradicting yourself there.
> R is far superior in its the quality of the packages, their documentation, their ability to behave predictably on a given data type.
I not only disagree but I think that the exact opposite is true for each one of these points. But if things are working well in our shop, I'm not going to try to convince you otherwise.
The primary reason to moving these to python is due to convenience/ the community. Most new work is published in python. If we find a new/ interesting model we want to implement, its probably written in python. Rather than reskin the thing in its entirety, its easier here to work in python.
A couple disclaimers: my group works primarily in geospatial data, and principally in LiDAR and multispectral imagery.
The coarse division I see between R/ Python, is that if you come from a research/ academic background (non-engineering), you probably learned to program in R. If you were an engineer, you probably learned matlab. If you are self taught/ coursera/ youtube, you probably learned in python.
R libraries are generally more geared towards academic research, and specifically, working within existing frameworks (handling geospatial data as geospatial data rather then turning them into a numpy arrays). Working in python, there is far more re-invention of the wheel, and its always a pain the ass to get things back into the structures they came in as.
Python has huge utility and is an important tool for certain work. But its really really not faster than R (it def used to be, this isnt the case any more).
R has better support for more scientific programming than python.
Not as a point of argument, just additional information: R's support for keras and TF is a wrapper around the Python interface to those libraries.
> I not only disagree but I think that the exact opposite is true for each one of these points. But if things are working well in our shop, I'm not going to try to convince you otherwise.
I partially agree with you here. I'm extremely careful about what non-standard packages I use in R. Code quality varies wildly outside of these, likewise for documentation. But outside of neural networks, I've never found a package in Python that I felt better about in terms of code quality or documentation than its equivalent in R.
To make what I think your point is more explicit, people build their own things in R because R must maintain compatibility with S. So by and large, changes happen in packages and not the base language. This does lead to a proliferation of solutions for the same kinds of problems.
I've been fortunate to only work on projects that use built-in data frames, never encountered tibble or data.table in the wild.
> there is no one go-to library for ML in R (like sklearn in Python) and each package has it's own strange interface and implementation.
I still disagree here - one example being the unified interface for generalized linear models. Also, the vast majority of classifiers (RF, SVM, etc.) have similar or identical interfaces. Also, the unified `predict` interface as well. Granted, `sklearn` does have a consistent API as well.
That said, some of this is just a personal preference for the vaguely functional interface in R. The object-orientedness in Python feels a little forced for some tasks in `sklearn`.
From my experience this is not the case. In biomedicine and bioinformatics few people actually use tidyverse because the data is much better represented as a matrix, and not in the "tidy" form.
Outside of that corporations (well at least 2 I contracted with) used `data.table` explicitly. Join 3 ad-click dataframes matching by userID, sessionID and closest possible time-point - that's one line in `data.table`.
Tidyverse is well suited for learning and for managing (relatively) simple datasets. But becomes cumbersome for more complex data. It can be used for those data too of course, just that it will be adding ad-hoc solutions and maybe get in a way more than help.
I've only use base R for my medical data (subsetting dataframe and such). Very rarely do I need tidy and also I find the pipe operator makes debugging harder. If and when I need it I'll use it that's that.
I think R have much more packages in medical, especially statistical packages, where many fields within medical cares about inferences not just prediction/forecasting. So I disagree with the "essentially the only thing keeping R in the game". The breath of packages in R is one of the many things that keep R in the game.
The tribalism and highly bias comments makes it very toxic and harder to have an honest discord.
They are just tools, use what makes you happy and get the job done.
1. It's aimed at beginners.
2. If you're a beginner, you're best off picking one language and sticking with it for a while.
3. There are so many other beginner resources that are much better.
I'm currently following ISLR book & course
1. Which language?
2. Do you have any background in programming and if yes how much and in which language? Beginners can range from "I don't know what a loop is" to "I've been a front end developer for a few years and want to try something new". These people will obviously need to approach things very differently.
3. What's your math background? Do you know what a derivative is and how matrix multiplication works? Do you want to go in depth or do you need just a general understanding of the algorithms? Some people will need to start at precalculus if they really want a solid foundation. The same people also most likely won't have the patience or interest to stick with the theory for long enough.
4. How will you use what you're learning? Is there a specific goal related to this and does it have a time horizon?
...
Anyway, I'm rambling. Data science is just programming and statistics. Your general learning lanes are
1. Theory - calculus, linear algebra, statistics, ML algorithms.
2. Programming.
2a. Good software development principles - writing maintainable code, version control, testing, design patterns etc.
2b. Tooling - learning the language-specific ecosystem of libraries. This is what most beginner resources (including the OP) focus on and is also the one that constantly expires and has the least transferability to other skillsets and fields. Obviously, it's still necessary - using the right tools and knowing them well goes a long way.
You mentioned ISLR - it's a good beginner-ish theory book that helps you understand how the algorithms work without going too deep into the math.
1. Theory.
Math, calculus, linear algebra, probability, statistics, ML algorithms. ISLR is a very good beginner-ish resource that helps you understand the algorithms but doesn't go too deep into the math. As you go deeper, you may realise that you have gaps in your math knowledge and you need to cover a lot more probability, calculus and linear algebra.
2. Programming.
2.1. Good software engineering practices - writing maintainable code, design patters, version control, unit testing etc.
2.2. Tooling - knowing the language-specific ecosystem of libraries (the OP is an attempt to teach you this in two languages at the same time). This is what most beginner resources focus on; your knowledge here has the least transferability and tends to go out of date quickly. Still, using the right tools and knowing them well goes a long way.
#2: coding is my hobby and have been writing well designed apps for a long time, so thats not an issue
#3: ISLR is teaching how to do ML algo in R, so there goes that point
#1 is what I'd like more information. Good important is the maths to work as a data scientist? I'm planning ISLR and then maybe ESL or some advance course
A few of my friends work on Data science and they said that maths isn't that important as in, one needs to know the formulas and why things work the way they do as in not rote learning the math.
It'd be great if you can list down intermediate courses, the learning market has drowned good tutorials and books with not so good guides!
> A few of my friends work on Data science and they said that maths isn't that important
Without the math you won't understand anything in ESL.
Which might be okay if the job doesn't require you to go into that much depth - some data science jobs are more focused on research (very math-heavy), some on ETL and/or engineering, others on business understanding and communication; it's a really broad title.
Core Python is fine. But pandas is an atrocious mess of object orientedness and other weird stuff.
SAS lives on borrowed time.
But the more Data Science focused roles (vs. Pure actuarial roles) are going more and more with R or Python.
I've seen pharma, nonprofit cancer organization, kaiser permanente hospital, State government (epidemiology), etc...
Also it's not just about the numbers.
Source for this, please?
If you're really going down to the FFI, it's hard to think it wouldn't be more productive, but that's not what a true beginner like the target of this book would do. Though it's quite nice to quickly extend some tool for your purpose without compromising anything or to understand how something works thanks to being written in high level code.
Syntax-wise, Julia's Common Lisp-like feature set gives the language a lot of power, but normal use will probably be just on par with Python in terms of productivity.
I work in finance doing data science-y things and have yet to meet anyone who doesn’t think that Python is a pile of garbage.
People used to make the easy to learn argument, but Julia is even easier. And more elegant, extensible, and faster.
However my point was that with Julia, those libraries would have been written in Julia.
All the Python libraries that one throws around for these use cases are C, C++ and Fortran libraries, that happen to have Python wrappers.
Any programming language can have wrappers for them, there is nothing written in Python per se.
Python itself is written in C. The Julia github repo shows Julia 68.2%, C 16.3%, C++ 10.4%, Scheme 3.2%. R is a mix of C, C++, R and some Fortran I think...
And how many Python libraries are just plain wrappers, not really written in Python.
I use TensorFlow from .NET ML and C++ API.
Can you give tell us which language and industrial grade NN library you're using?
Because from where I'm sitting I see that Python is the only language that gets first grade support for both Tensorflow and Pytorch. It's so ahead for working with NN that it's not even close.
And this is no exception.
The website is decent. Read the standard library https://github.com/JuliaLang/julia/tree/master/stdlib
search github for cool projects https://github.com/JuliaInterop/RCall.jl
Nowadays I program in python because it has all I need. If something comes out in Julia that makes the cost of picking up another language worth it then I'll do it without a second thought. Until then, why bother? The don't waste your time argument can cut both ways, you know.
To clarify: I see no inherent reason to not program in Julia or any other language. But you work with whatever gets your job done efficiently and right now that's neither of those languages for a lot of people.
Not really, syntax and semantics are adjoints.
In a great many cases, they aren't a first-order issue - which is where my objection to a blanket "don't waste your time" claim comes from.
R, Python, Matlab, and C++ are the big dogs in scientific programming, and the inertia behind having a large community behind than will continue to drive adoption.