HNHacker News
TopNewBestAskShowJobs

2-tpg

154 karma · joined November 1, 2020

submissionscomments
2-tpg··on Jane Street Market Prediction ($100k Kaggle competition)
The winner model will likely outperform anything that Jane Street could come up with this 130 feature set. With 3000+ competitors, the top 10 will likely be superior to what Jane Street can do in-house. Then an ensemble of the top solutions will be the best possible model anyone can come up with.
2-tpg··on Jane Street Market Prediction ($100k Kaggle competition)
Using purely historical price data it is harrowingly difficult. There are 130 anonymized features, so that's unlikely to be only price data. It could include information on the order book, correlated assets, fundamentals, vectorized/embedded text, etc.

Besides, I bet you can train monkeys to do (slightly) better than blindfolded random throwing. Even with public data (replace satellite images with Youtube mentions, or number of links moving into a company website) it is very possible to do better than average guessing on quite a lot of assets (especially smaller and newer markets).

Most hedge funds, even with specialized expensive non-public data, are not magical unicorns. Their quants really may just run a gradient boosting machine and leave it at that. Some hedge funds even prefer linear methods, because this lowers risk through lower variance. Such models can be beaten by experienced Kagglers for sure. For one, I did.

2-tpg··on Jane Street Market Prediction ($100k Kaggle competition)
Posting notebooks can get you upvotes, which contribute towards becoming a Kaggle (Grand)Master. It is also a good way to "win" some attention and goodwill, without spending months trying to actually win the competition itself. Publishing Notebooks also helps you improve your coding/presentation skills, for a popular notebook needs to be useful for a wide audience (or fairly competitive).

The best techniques, certainly coming from teams, are hardly ever published as Notebooks. But yes, many winning teams will eventually incorporate some of the information in the Notebooks, if only to hedge against the others doing the same.

2-tpg··on Pfizer submits Covid vaccine to FDA for approval, to distribute in December
It is my understanding that adversaries of Western nations have already started disinformation campaigns to mess up the vaccination effort. The entire Bill Gates / 5G causes COVID discussion from a few months back was fabricated; their current playbooks should be more sophisticated and target also outside the big social networks, perhaps even on this very site.

Also, that governments will show little patience for anti-vax talk of its citizens. They'll have to switch talking points from avoiding panic (It's harmless for young people!) to accepting the vaccine (There are many long-term effects of infection!). A Herculean effort in balancing propaganda with allowing free speech.

If nothing else, it is going to be interesting watching the next months unfold. Saying you are refusing the vaccine may simply put you on a list, but openly detracting from the vaccination efforts, should see some actual (and scary) pushback.

Public health will be an increasingly difficult thing to manage, when individuals get the information for their decisions online (however wrong).

2-tpg··on Yann LeCun on GPT-3
I used GPT-2 to create a health website. One sentence was enough to get a full page of authoritatively sounding lists of symptoms and treatments. Very diverse, unlike all other sites, because the articles it generated only looked and sounded like a health encyclopedia. Of course it is going to spit back decent diagnosis, when it is in the training data, but what do you trust? An expert system that logically and interpretable explains its predictions, linking the original source. Or a language model that uses a temperature to stay on track, and randomizes its output on every new run?

Generating data with a possible high impact on lives sounds like a recipe for disaster and frankly, irresponsible. And Google would have to really solve it, to detect false or questionable information, when its not possible to rely on spam signals (like when a legit site is transferred to a malicious spammer).

Aside, I bet LeCun would be more favorable of GPT-3 had it been a deep CNN and they had adopted his self-supervised learning paradigm :).

← PreviousPage 3 of 3