249 karma · joined January 17, 2015
feedback from af2 folding confidence + structural scoring
LLM (foundational papers)
* Attention is all you need - transformers + self attention
* BERT - first masked LM using transformers + self attention
* GPT3 - big LLM decoder (Basis of gpt4 and most LLM)
* Instruct GPT or TKInstruct (instruction tuning enables improved zero shot learning)
* Chain of Thought (improve performance via prompting)
some other papers which are become trendy depending on your interest
* RLHF - RL using human feedback
* Lora - make models smaller
* MoE - kind of ensembling
* self instruct - self label data
* constitutional ai - self alignment
* tree of thought - like CoT but a tree
* FastAttention,Longformer - optimized attention mechanisms
* React - agents
If you're interested in adversarial NLP, I also recommend reading this blog post on adversarial attacks on GPT2 with universal triggers (e.g. adding "nobody" as prefix for all inputs causes all entailments to be predicted as contradiction).
10 day windows were used to reduce the amount of volatility/noise in a time frame
Return horizons for 1 years was used because price targets are for one year.
Theres only 15 or so analysts I looked at.
I was doing this as an exploratory data analysis and didn't want to pull out my old stats textbook.
Cutoffs were chosen to reduce volatility of measurements since I was looking at percentages. A stock going from $1.5 to $2.0 is a 33% increase whereas the movement of $100 to $133 is significantly more impactful. Stock with lower market cap have more volatility. The minimum analyst rating was chosen to eliminate analysts with very small number of ratings as they would be unreliable.
My blogpost shows that stock price predictions also show a terrible track record. They are wildly off and on average higher than actual results.
The top analysts were determined by the average performance from one year after their ratings have been made. This isn't the top analysts out of 50,000 it's the top out of 50 or so analyst-rating pairs. There were only 16 or so analysts in total that I looked at. This isn't an instance of survivor bias as your example states. If I were to be more rigorous I could give a statistical test for this.
The analysis was more about measuring the performance of analysts which is why the price data for before and after the recommendation. For practical purposes of using this strategy, you are right that the price data from days after release would be better.
If the top 10 stocks you picked beat the market and have consistent earnings and dividends over a period of time, would this not be a repeatable strategy?
2. A more in-depth analysis could be done on analyst releases' effect on prices but assuming that this does occur, then the performances are understated and provides further evidence to the conclusion that outperform ratings can do better than the market.
3. Not sure how narrowing down the top analysts is a flaw here.
This blogpost is probably not as mathematically rigorous as it could be as I just wrote it as an exploratory analysis for fun and out of curiosity.
I am the person who posted the blog entry. It's true that I don't know much about how NLP parsing works but I do know how the parse trees are structured. I believe matching parse trees is scalable. The examples in my post were for short imperative commands but it is relatively simple to create rules for more complex sentences. It might not be perfect in every case but I would say for a majority of cases it works well.
I'm glad you were able to understand the blogpost and I agree that the material on tregex is not clear and would be difficult to pick up. I hope the library I wrote will let programers start using the Stanford parsing libraries more easily.