H3 – Outperforming GPT-Neo-2.7B with only 2 attention layerstwitter.com10 points·m00x··0 commentsOpen articleSaveView on HN