An Introduction to Stock Market Data Analysis with R – Part 1
ntguardian.wordpress.com
ntguardian.wordpress.com
I've seen people use python (or R) a lot in online tutorials and courses (there's one on coursera[1] check it out). It's understandable that python is preferred since it's an easy language to get started with, but I've never really seen C/C++ used by people online (Although I've heard about the use of FPGAs by the hardcore guys)
Is the only reason C++ isn't recommended because it is difficult? The performance difference seems rather large to me. And especially if you're doing lots of number crunching, it would really benefit using C/C++ wouldn't it?
Do people use them but I don't hear about them because anyone above 'tutorial level' doesn't share their code or talk about it?
Or are the R and Python Computational libraries close to the performance of C/C++ that it's a viable compromise?
Conversely I don't think C++ has this same type of library for reading in data in a single line.
While C++ is faster than Python overall, for most financial simulations Python is definitely fast enough.
It's mostly Python with a bit of Cython, and pull requests that are not pure Python are more likely to be rejected. There's basically zero C++ in Pandas itself.
That combined with it's ease-of-use working with tabular data, and the myriad of packages available make it a very popular choice.
In that sense it is much faster to do these analyses in Python or R than in C++.
I think Python is a great tool for finance, particularly with libraries like Pandas, Numpy, etc. Most of the work in those libraries is C/Fortan/etc anyway. Great community. For more 'academic' style work/exploration, I think R might have an edge, but I personally find it more difficult to write and maintain high quality code in R.
I'd like to just throw a quick plug in for an open source project I'm working on [1]. It's an event-driven trading library written in Python. We're pretty close to having live trade execution ready in an alpha state.
The motivation is that you can put together a pretty simply strategy in ~40 lines of code. The codebase is designed to be pretty flexible & modular. Mike from http://quantstart.com has some great resources & many of his recent articles use the same library.
Source: https://ipnetwork.bgtmo.ip.att.net/pws/network_delay.html
I'd guess that low effort threshold to try and backtest ideas is very attractive even if you're going to have to rewrite fast/better/stronger in other languages to trade in real time?
That's a big part of it. Python's data analysis tooling is generally written on top of Numpy, which is insanely optimised code that an average C/C++ developer couldn't compete with in terms of perf. Numpy will completely smoke naive C code to do common data manipulation tasks. So it's a double win, really. You get very good perf out of the box, and you get a high level dynamic typed language which lets you focus on high level logic and iterate quickly.
The other part is the ecosystem. For a variety of reasons, Python has become the premier language for Big Data™, and over the past 5 years has accreted a huge collection of libraries for analytics, visualization, ML, distributed compute, etc. A single developer can now click these libraries together to achieve stuff that could only be done with huge teams before. You can build a distributed training cluster for a deep learning algorithm and deploy it to Amazon in maybe 2k lines of code. C++? I don't even know where to begin.
R could be just as capable as Python, but I think Python has largely won the race to be the most popular language for data analysis which in turn encourage more developers to commit to it, cementing Python's advantage.
R still has solid lead in statistics and a good mindshare amongst academics.
Your comparing Apples and Oranges. R is a domain specific language and will never be a general purpose language.
It is not true that Python won any race in statistics. http://www.kdnuggets.com/2015/05/r-vs-python-data-science.ht...
Let alone in industry investment coming from Microsoft and other major players.
R is above Python in Statistics in momentum and numbers. Python is a good choice but Python is still playing catch up to R due to the speed at which R is developing. R with data.table and Hadleyverse (https://www.r-bloggers.com/welcome-to-the-hadleyverse/) and RStudio the momentum has been clearly on the side of R.
R just 5 years ago was a fraction of the users it has today.
Python and R are both good choices with equal speed but the difference is that R is a domain specific language that has a lot of positive ecco system.
Python is a good choice but R is amazingly good language that doesn't deserve to be pushed back negatively.
It's not easier to learn Python then R. They both have plus and minuses but knowing both if I was to teach someone statistical programming I absolutely would teach R over Python.
This disclaimer is a good idea, but I'd have gone as far as saying that getting into HFT as a hobbyist programmer is going to be throwing your money into a giant hole.
i will say as an advent stock guy i think most of this statistical stuff is bs.
i think we like to find patterns and only the ones that work in the past dont really work in the future.
stocks are run off news and articles and weather the big bank sells today or buys among a myriad of other things.
i believe hft are only really meant for big companies that have to trade 300 positions and maintain certain standards.
Not as a average guy running one in his basement to make $200 dollar a day like clock work, only if he worked a little harder and studied more phd statistics and found this "amazing unbeatable forumla".
by the way would you like to make $5242 a month from working at home? insertscamwebsitehere,com
Any advice?
Also remember that this market (or any market with big money) is brutal and will eat you alive. I don't expect the average person and his capital to last for a long time.
In fact, many still are. Not much has changed in the field of economics. How data science isn't a mandatory requirement for such a data driven field just shows to me how immature the field is.
They should not trade.
have you guys been on https://www.quantopian.com/
very informative
very good blog statistic articles on top of backtesting
I had a friend who used to run a financial exchange. That was his favorite sentence.
The house always wins.