I have wanted to do something like this for a while, purely for learning. The thing which puts me off is that there is a huge amount of knowledge needed in understanding the features vs the ML.
Could you recommend a base system / reference one could use to get started which explains or bakes in some of the feature / signals engineering work?
Also would this approach work with crypto?
Some of it works on crypto. TBH I've stayed away from the asset class, but only because I find it difficult to build mental models and think about features (in my mind, it's a mix of commodity factors and currency factors, but I'd have to test it out).
I seem to remember coming across papers that have tested momentum factors at larger time-frames (e.g. weeklies).
> Could you recommend a base system / reference one could use to get started which explains or bakes in some of the feature / signals engineering work?
The references I put in at the end of the post will really help with this! I might actually write out a separate blog post about starting out in this space from an ML perspective. Thanks for the idea!
Across any market area (eg: mineral resources) there are thousands of documents released daily across multiple exchanges (via EDGAR, SEDAR, etc) ranging from two line advisories, to 4,000 page technical reports on projects | acquisitions, alongside the usual quarterly | yearly annual reports, etc.
There's plenty to do parsing common forms for generic changes (board members, board member share changes, etc) and market regime specifics (exploration property aquisition) and trends (series of related aquisitions) for those that like the weeds.
Some might argue that 'understanding' these patterns lead the changes in stock price movements, and give insight wrt weathering short term changes for longer term returns.
I focused on 10-Qs for the EDGAR filings module as you rightly pointed out - it seemed to be a good balance between implicit information and usefulness of the data. TBH I didn't actually investigate the other (many) patterns.
Having said that, I have really enjoyed Kai Wu's research from Sparkline Capital (https://www.sparklinecapital.com/), especially his extraction of the innovation factor from EDGAR filing texts. He's appeared in numerous podcasts, and they have all been super useful to listen to. Maybe someday when I re-investigate EDGAR filings and go further, I might target these signals you talk about here.
14+ years back a small group of West Australians put together what became
https://www.spglobal.com/marketintelligence/en/campaigns/met...
which was based upon integrating (GIS and regular DB) every daily mineral lease record across the accessable globe together with every publicly filed document across the relevant stock markets (AU, TSX, South Africa, London, etc) using (and updating|refing) templated patterns that appear in various classes of forms .. I would assume that territory has been revisited with better ML techniques.
A few questions:
1) Were these stock picks for major stocks / ETFs? Or small market cap stocks?
2) How many people were subscribed to your newsletter?
3) What do you estimate the impact was of creating a “self-fulfilling prophecy” of entering a position and then recommending your subscribers take the same position?
4) Do you think your asset mix outperformed the market by picking high risk / high reward stocks in a bull market? Or picking safe stocks in a bear market? In other words, how do you know the engine wasn’t biased towards the market trend that happened to play out? For example, if I have a basket of tech stocks that would typically outperform in a bull market and flip a coin to buy or short them and guess right, I could outperform the market by chance. How did you account for this?
5) Have you backtested the engine to see what it would have returned in previous years? (Obviously on unseen data, rather than data it used for training)