Automating eval design and hillclimbing with Claude
claude.dev
claude.dev
Recently two bad lines were approved and then published and weren't the judge's fault. As the older changes never went through the panel, an earlier analysis had labeled each change and I then approved those by hand. Two were labeled as no change in meaning and only added one line summaries as notes for the reviewer and not text for the public. I only caught it the next day.
Running the grader once more wouldn't flag it because there was no inconsistency. The fail was human. So my question is when the skill asks whether you'd score a case differently, does it show the text it will publish or only score and reasoning?