ParentFull threadtextcortex·Most probably they are doing something called prompt tuning. This creates a small ai model that adds virtual tokens to prompt before passing to original model: https://developer.nvidia.com/blog/an-introduction-to-large-l...View on HN