Full threadkp1197·Does performing gradient descent on token input embeddings lead to interpretable results? And if not, why?View on HN