Full threadamatic·This sounds amazing! Are there any metrics on how often different models pass tests? Has someone used a similar process to finetune an LLM?View on HN