Numerai – A hedge fund built by a global community of anonymous data scientists
numer.ai
numer.ai
- You need to know something about the domain in order to make sensible predictions. Is the data daily? Is it per second? Is it ticks? You can't build a sensible model if you don't know that, even if you have good predictions. Relative cost will vary a lot between timescales.
- It matters what the features are. Maybe there's some clever reason why it doesn't, but until I hear why I'm going to take the ordinary view that some features are different in nature to others. For instance maybe one feature is volatility, a thing we typically model with GARCH, while another is some fundamental like P/E, which we'd incorporate some other way.
- How are you executing the trades? It matters a lot whether you're click-trading through some broker API, automating via Excel, or running your own network of colo servers. Some things just aren't possible if you're too slow.
- If you make the data encrypted, you'd better know very well what it represents. For instance, you might take all the closing prices of the LSE stocks on a given day as inputs. You can make analyses that are valid with that, and ones that aren't, because the data you've collected do not represent a snapshot of the market at a specific time. It might sound like it does, but it doesn't on deeper inspection (market opens and closes are not simultaneous).
Does anyone know how it's going for them?
Im not optimistic...
While prediction based on data may be valuable in some cases, it isn't robust enough to scale up in any meaningful way. Context matters, like you state above, and most quantitative traders start by taking their contextual knowledge of the markets, and then collecting data on features, and THEN they fit a model to it.
Skipping these steps is only going to lead to a bunch of blowups. I doubt they have any meaningful sharpe that they could scale up or publicly defend with the approach they have taken so far. I'd guess they are paying people VC dollars, not actual market profits right now.
I do think its cool that they have been able to use homomorphic encryption to solve the problem of wanting anonymize data, but I'm not sure it actually helps in this case.
Let's give many people some [encrypted] data to play with. May be someone will find a curious and unexpected pattern we didn't looked at before. May be this idea won't be workable straight away but we [hedge fund] are going to investigate this idea further and probably use this in our trading strategy.
In short, they are probably doing ideas mining using power of crowd.
This is not a pure math problem. Eventually the outcome of all these models and predictions affects the stock prices and - if it becomes as successful as you hope - the economy as a whole. And the physical world: people, animals, plants, pollution, CO2 and so on.
I would much rather see work in that area - than this data juggling deep learning bullshit which results in profits being paid out to a bunch of intelligent, greedy and unwise people.
I just hope that when the intelligent people finally become wise, it won't be too late.
Holy anti-intellectualism, Batman!
If you describe machine learning as "data juggling bullshit", this is very strong evidence that you simply don't understand what it is. This is an indictment of you, not of machine learning. Machine learning would more accurately be called "applied computational statistics" in 99% of cases.
What makes you think that the people using applied statistics to make money are "unwise"? Based on your tone, I would guess that it's because what they're doing doesn't agree with your folk definition of "an honest day's work" or something like that. This isn't really a good criticism; it just means you don't see the utility of what they're doing, which requires some degree of abstract thinking about the market.
I believe he's suggesting that there are a lot of other areas where their knowledge and skills could be applied which would create some societal benefit as a byproduct.
It does both, the latter by more efficiently communicating price information and incorporating future events into asset prices. More accurate AI traders absorb risk, which is good for the production economy.
Particularly the problem of this website, is it's huge focus on anonymity. You don't need to expose your code or methods, or even your name. It's an ongoing competition so it encourages secrecy so you can make more money next week.
This means even if someone does invent a new super ML method, they have every incentive to not share it and keep it a secret. So this website could actually have a net negative impact on the world.
Why spend 80 hours a week chasing grant money when you can spend 60 moving numbers around?
In fact I'd say this job is quite uniquely, extremely, only about the money. Even when compared to other kinds sweet/lucrative, perhaps "stupid lucky" jobs you may find, that may require 60 hours mindlessly pushing papers, yet paid extremely well. I'd be hard-pressed to come up with any hypothetical kind of job could be more sharply, singularly focussed on the "do extremely-intelligent-monkey-dance, receive ample bio-survival-tickets" to the exclusion of any other meaning or fulfilment than this Encrypted Machine Lottery/Learning/Sudoku.
Second, I could be wrong (or too optimistic) but I'd hope that most "very intelligent" people demand more nourishing reward than just money to regularly put in more than 40h/wk. You can do this for research and advancement of science, or maybe because one considers the work itself a net positive to society, or maybe you chose to sacrifice those hours of your life to support loved ones, maybe there's some deep personal reward in it, or maybe it's temporary and you're saving for something rewarding. But to pretend just money is enough to waste your life on, hey it's a choice, but I'm going to have to see you turn in your "very intelligent"-card.
On the other hand, there will be exceptions. Some people don't care, just want to maximize $$$ for the least amount of effort/hours of one's life, regardless of side-effects. This website does not fill a niche for these people. If you are extremely intelligent, just want money and don't care about moral aspects, there's always been plenty "business opportunities" to fulfil these particular needs.
Finally, if one would argue, but this job isn't quite explicitly "badwrong" like those others. I'd suggest to think it through: On the one hand, a (possibly) well-paying job that is designed from the bottom up to have no way of determining whether it has a net positive or negative external effect besides the money it pays you. On the other hand, literally anything else you could do with your time.
I can see people trying this for a short time, as a funny puzzle, at most.
[0] I don't mean this one, but because I still think it's funny, I'll point out the way-too-easy retort here: this says something (not terribly good) about you. (j/k)
Actually, it does. Increasing liquidity, communicating price information, more accurately accounting for predictable future changes in price, and absorbing risk allow for increased agricultural and industrial production. Like I said, you have to think through a few layers of abstraction.
A nice rule of thumb is that if someone is getting paid to do something you personally think is useless, it's either a government job or you just don't get what they're doing, but the market does.
The stock market is one of the mechanisms that allows the market to be right by communicating price information.
But the opportunity cost is what I'm getting at. So many very intelligent people could be producing immense value in other areas, but instead drained away doing this garbage. And they are incentivized to keep anything useful they invent a secret.
http://www.investmentreview.com/files/2010/07/The-value-of-l...
https://fp7.portals.mbs.ac.uk/Portals/59/docs/Finger,%20Mark...
http://people.stern.nyu.edu/adamodar/pdfiles/papers/liquidit...
https://www.macquarie.com.au/dafiles/Internet/mgl/au/mfg/mim...
https://en.m.wikipedia.org/wiki/Liquidity_crisis
The stock market is absolutely connected to "real value" (which doesn't make sense as a concept in the way you're trying to use it). Besides allowing companies to raise capital for future endeavors, it also allows the market to communicate price information, which is critical for reducing the risk associated with a given transaction.
> at best any benefits just benefit rich investors buying stocks.
It benefits everyone; investors risk their capital by selling it to companies in exchange for partial ownership, which allows the company to grow, and if the investment pays off the person who risked their money is rewarded. Without the stock market, companies would have to get loans for growth, which is risky and too expensive to be practical.
> It certainly doesn't increase production.
Yes, it very much does, for the reasons I outlined above. This is like Econ 101; why are you commenting so strongly when you clearly do not have any background here?
> but instead drained away doing this garbage.
A small efficiency increase expressed over 100 trillion dollars of trades every year is a huge productivity increase for society.
Look, for instance, at high frequency trading. Traders spend millions of dollars trying to get signals across the earth a millisecond faster than the competitors. Does it actually benefit anyone that information travels a millisecond faster? Who benefits from the high frequency traders at all? Yet they eat up millions, maybe even billions, of the economy's resources doing this nonsense.
Look at these trader spending millions of dollars to fly drones over oil tankers or parking lots, to get slightly more accurate information faster than anyone else. Look at these hedge funds spending tons of money on bribes to people in companies to get insider information before the public does. These things clearly provide zero benefit to the world, it's just a massive waste of resources.
This hedge fund is a bit different, in that they seem to be actually doing fundamentals and longer term bets. That's not so bad, but it's still pretty disconnected from any real world benefits. At best they predict a company will increase in value and buy it's stock. How does that benefit anyone? It's almost a zero sum game. The people that buy the stock that week lose, because the price is higher than it otherwise would have been. The people that sold the stock lose, because it was really worth more than they actually sold it for. Any money the hedge fund makes necessarily comes from someone else losing that money.
The stock market itself is not really connected to the real world. Sure sometimes companies sell stock to raise capital. But most stocks being sold and bought are not by the companies themselves. It's a huge indirect chain of traders and investors and speculators, trading these stocks many, many years after the company sold them. And the company itself would still be able to sell stocks if there were fewer traders. At worst the prices would be slightly less accurate.
Price information of stocks is also a horribly inefficient way of getting this information. These companies are often involved in many separate projects. So even if someone has a model that can exactly predict the success of products, it's very difficult to use that information. You also need to determine how much the company is worth given all of it's projects and businesses and property it owns. It doesn't directly give that information to the company itself - they see their stock went down, but they may have no idea why. Not to mention Keynesian beauty contest problems, where you have to predict not what the company is worth, but what other people think it will be worth, and how much they know currently, and what the interest rate is, etc.
But you didn't address my main concern, which is the bad incentives of secrecy, and the opportunity cost of employing smart people doing this, instead of doing something else. Even if it does produce some value, it doesn't mean it's worth the cost.
Because it is an incredible waste of talent.
I know this is in contradiction with all the 'values' that have been drilled into our minds since we were born, but it's about time we wake up and reorder our priorities.
>Based on your tone, I would guess that it's because what they're doing doesn't agree with your folk definition of "an honest day's work" or something like that.
If you look at the state of the biosphere / atmosphere / oceans - the data - and if you have children - then it should be quite obvious.
Also there are the socio-economic challenges that the world faces right now - in fact it is unclear if we're going to make it to the next century as a species.
It really don't matter how much "money" you have in your account when your city sinks under the ocean...
You just move. That's what money enables. No city is going to sink so rapidly people just drown.
The fact that these people chose to spend their time doing numerai probably means that they think that's the best use of their time.
The only way we're going to wipe ourselves out is with nukes, nanotech, or engineered viruses. Global warming is not an insurmountable concern. It will be expensive for coastal cities, though.
There's no reason we can't focus on multiple things. Global warming isn't so pressing that we need to dedicate all intellectual output to it. Every 10 or 20 years, the environmental alarmists come up with something new to obsess over (I think it was landfills previously), and they're doing it now with global warming.
Ironically, it's also not entirely a math problem - the underlying system doesn't exactly follow the expected rules of math.
They are in effect crowd sourcing curve fitting because the number of possible models for this maths problem is so large.
Think of crowd sourcing here as a pruning heauristic in a search problem.
I just downloaded the training set [1] and plotted some of its descriptive statistics [2]. It looks that all features are uniform distributions and the response variable is Bernoulli coin-flipping. In layman's terms, you can't really come up with a good predictive model with this training set.
I give them the benefit of the doubt that they wanted to have something in place to push the website live, but I cannot imagine any serious data scientist not noticing this.
[1] http://datasets.numer.ai/57feb95/numerai_training_data.csv
I would love to see peer reviewed articles from numerai with some of their behind the scenes results.
I can't imagine any serious data scientist not knowing that.
>any serious data scientist __ - I mean, any human - data scientist or not - would see at least that the website wasn't just pushed live; or at least have noticed the film is from like August... ^are you a serious data scientist? If you are (which I assume you are - just also an eager beaver to download the datasets without learning about the company sufficiently) -- then you'd do well to join the tournament and have a real go before you [attempt] to knock it as you have here -- and then come back and give some real feedback -- I'm sure everyone would appreciate to hear the 'after' report from you -- since you're so kind as already to dish out benefits of doubt :) Looking forward to catching you above controlling capital!
[0] https://en.wikipedia.org/wiki/Multiple_comparisons_problem
edit: Also, by hoeffding's inequality the number of training examples needed for a given level of confidence is only logarithmic in the number of models (even assuming they are independent). See page 6 here: http://cs229.stanford.edu/notes/cs229-notes4.pdf
That doesn't matter though because you only a need small amount of good signal to be profitable in the money management biz. The larger issue is whether the aggregate signal quality is high enough to be able to pay for development and trading costs.
What does this mean? What do they do?
This brings to mind the Buffett hedge fund wager, where he invested in a vanguard s&p 500 tracking fund (VFIAX) and a hedge fund actively managed an equal amount, and Mr. Buffett ended up winning handily.
Most people don't realize that the "markets" are 49% random, 48% sentiment driven, and 3% fundamentals. If you approach the problem with that assumption held true, monkeys throwing darts isn't such a horrible mechanism for investing.
See: "Monkeys Are Better Stockpickers Than You'd Think: Why dart-throwing primates demolish S&P 500 returns and most active fund managers don't even come close." http://www.barrons.com/articles/SB50001424053111903927604579...
Maybe I haven't understood that part of data science but I never got why throwing more unpredictability on an already unpredictable data source would somehow make it more predictable.
I guess another way of saying it is that your mess is starting to look more like their mess.
Anything that deals with the future is inherently non-predictable. Using chaos theory as a framework, we say it is unpredictable because we are unable (and will always be unable) to model the currently system completely. There will always be data that was not captured hiding between the data that was captured. Follow the arrow-of-time far enough out into the future and that non-captured data will manifest itself in the captured data, thereby (usually) creating a deviation from the modeled future.
To get around this we use statistics and probability. We rely on the law of large numbers and regression to the mean. In other words, we hope that the future won't get too weird and will be similar enough to the past, within some confidence interval.
So, we're not really predicting a specific outcome, we're predicting that the outcome will be some point within some confidence interval.
The reason we can get better at this, is as we capture more data we can better guess the inputs and assumptions we use to create the model. We throw out stuff that didn't happen to be relevant. We discover stuff that we should have considered relevant. If we're lucky the model closely follows the physical laws of our reality and we can apply the frameworks so arduously worked out by chemists, physicists, biologists, etc. If we're not so lucky we're dealing with sentiment, conjecture, or any of the other human inputs of the financial markets, and we are forced to make up formulas that work until they fail spectacularly.
I am skeptical of Numerai for different reasons. If someone can consistently churn out profitable and novel equity pricing insights, it would be more rational for them to work for a more well established hedge fund in quantitative research. Perhaps more importantly, I'm skeptical of how they judge accuracy in their participant-volunteered insights.
A former tournament winner did well on both the public and private leaderboard. It is very difficult to do this by luck. He was also a student from Bangladesh who got the opportunity to play with hedge fund data with just the cost of an internet connection and zero risk for messing it up. Should he start working for a hedge fund now, without any finance experience? Could be a good bet. Numerai would still beat him, because they can aggregate all the top models into an ensemble. It is hard to beat 50 individuals, but near impossible to beat a team of 50 competitors. Compare the variance of a single decision tree with a Random Forest.
Perhaps they live in the wrong location -- it's hard/impossible to get a quant job if you don't live in a major market center. Not everyone is 23 years old and unattached and prepared to move around the world for a job.
Perhaps they can do the predictions but they didn't go to a high end university and have no track record -- just try even getting an interview at a hedge fund without one of those.
The bet isn't over until December 31, 2017: http://longbets.org/362/
So in a sense, the "active" manager is just as much a monkey throwing darts as the "inactive", yet one of them is better than the other.
Why can't a third monkey be even better?
When you take a group of hedge funds, they end up being a very good proxy for the market. At about 20, they will be almost indiscernible.
So - a passive fund, buying the market, with lower fees, should outperform.
The returns are far from indiscernible from S&P 500. At the end of 2015, the S&P 500 portfolio had seen 3x the profits of the Protégé portfolio.
You're also forgetting how much the compounding of the fees can cost, and you have two layers of them. This could easily make the S&P at least 3x higher in returns after 10 years.
In CFA Institute's estimate, the difference in fees only made up about half of the underperformance at the end of 2014, especially considering that 2008 was an easy win for hedge funds:
https://blogs.cfainstitute.org/investor/2015/02/12/betting-w...
Note that at the end of 2015, eight years had elapsed, not ten.
The short portion of the hedge fund will have been losing as we are in a bull market. The full length of a cycle should see this effect reverse (and indeed 2008 shows this).
So Buffet did have to pick a bull market in making this bet, but he still had better than 50/50 odds regardless.
You also have other costs specific to hedge funds over active funds, such as higher brokerage from more frequent trades, and short interest. These have the same drag as the management fees, and I should have included that also in the description.
Participants with the most accurate [1] insights are paid a sum in exchange for the insights being incorporated into the hedge fund's trading system, which in turn manages the fund's capital.
_________________________________
1. "Accurate" is a tricky word and I can't comment here on how much effort Numerai puts into ensuring that insights are accurate over meaningful time scales instead of being the result of e.g. overfitting. The insights could presumably be accurate on ranges of anywhere between "once" and "months."
I don't know why you're harping on about sound encryption, the point of this is to keep the statistical information intact in the cipher, without giving away the underlying market data.
I think it's either an abuse of the word 'encryption'. That, or they really have done something weird to this dataset. Which will probably make it useless for statistical algorithms. Even normalization destroys a lot of useful information.
It doesn't have to have an exponential time complexity on decryption to qualify as 'encryption'. Multiplying by 2 could be considered homomorphic encryption.
You might think encryption means something else and that it's an abuse of the word but unlike the spy novella that you derive this impression from, these guys actually are ex-spies.
Thanks for explaining this, I was struggling to figure out the difference between this and https://www.quantopian.com/
That's actually pretty clever.
[0] https://en.wikipedia.org/wiki/Computational_indistinguishabi... [1] https://en.wikipedia.org/wiki/Ciphertext_indistinguishabilit... [2] http://cs.au.dk/~stm/local-cache/gentry-thesis.pdf [3] https://arxiv.org/ftp/arxiv/papers/1305/1305.5886.pdf
Also, I think the big benefit from a data scientist working on this is you can test methods with generic features, submit the results and get paid if they are good, but not submit any part of the methodology to any third party.
If you can kill it on numerai then maybe you would consider buying data sources and apply your methods to your own data. Although you still don't know what the features are.
It's the polar opposite of open source.
The owners don't have to trust the data scientists. They evaluate their results against additional data.
Hell, you don't even need to be the owner of the company. This would be a great way to obtain large amounts of political sway/power. Like a company/want it to succeed for some agenda? Make it look better as an investment opportunity. Dislike a company? Well that stock is going to do horrible next quarter. It's also a self fulfilling prophesy.
For all of you finance people out there, is what I am saying impossible or stupid? I hope I'm wrong otherwise this is a horrible idea.
It's probably against their TOS though.
For 100k per "tweak" you could make a lot of money and still give people rather significant influence over the world (especially if a LOT of people used this service).
> pump and dump, which is considered fraud
Yea it is definitely a pump and dump but if it makes you rich and you're in a country with no extradition treaty who cares? Remember the golden rule of the "elite class": laws are for poor people. As long as you get rich before anyone sees what you aredoing, and you "cant" be found, you're all good.
Maybe they have some way of preventing people from gaming the system that way?
In today's news anonymous scientists form human genetics laboratory to improve the human species.
* How does a logloss relate to earnings?
* They only receive predictions based on old data (by definition) and not the models, how do they just the predictions them to make trading decisions?
* How can I invest in this hedge fund?
Found this company interesting, read a bunch of blogs from them and tweetstormed: https://twitter.com/Royal_Arse/status/787725301908242432
It seems that only very few nerds are taking the theoretical impossibility of predicting the future seriously and we are missing the opportunity to get funding for some crappy models from the greater fools.