Claude 3.5 Sonnet Reproduces BIG-Bench Canary String
lesswrong.com
lesswrong.com
Take the set of benchmarks you care about, and then as you build a training dataset (by scraping, or whatever), you scrub each new item for benchmark questions (or just discard that entire item/webpage/whatever for being contaminated).
Otherwise you're negligently and willingly inflating your benchmark performance and defrauding new investors / users / customers, I think.
In other words, if that string is present, the benchmark results for that model are a lie.
However, let's assume they should care because it's a major benchmark from an industry leader.
The entire point of the canary string is that LLMs are supposed to discard / ignore it and data found on the same document.
After all, the documentation literally says
""Do not edit the canary comments. These are to prevent BIG-bench tasks from leaking into web-scraped training data.""
Anthropic did not do that (they obviously HAVE scraped data containing the GUID), therefore it is demonstrably a gotcha.
e.g. it should ignore both
https://github.com/google/BIG-bench/blob/main/bigbench/bench...
AND
https://github.com/google/BIG-bench/tree/main
Even though the latter is the readme, it has the guid, and there's no reason not to ignore every document containing the GUID.
So, if Anthropic wants to ignore it, fine, but it still feels a little fishy, doesn't it?
I disagree with this. There are plenty of websites out there that talk about LLM training in general, and have sections dedicated to canary strings. This page for example has the GUID in it: https://ravinkumar.com/GenAiGuidebook/deepdive/BigBench.html
I'd argue that it is something that LLMs should train on. Having context of how LLMs work is something that isn't related to the benchmark data at all. Just because the GUID shows up as an example doesn't mean the benchmark data is present on the page.
The LLM will still know how LLMs work without having trained on the handful of documents containing that specific canary string, because other documents will mention the concept of a canary string without that exact GUID.
Better to do that and be on the safe side and look honest than have people believe your company not really competitive.
Anything else is a risk for no gain in an industry theoretically worth trillions.
I expect people more clever than me have spent more than the three minutes I did thinking on it, so I'm genuinely curious as to how the two scenarios above are protected against. From my limited understanding, though, it feels like there's just an honor system involved.
Still, I think it's less than clear that the problem here isn't just the idea that you can publish benchmark data on the open web at all without making the benchmark outdated. Expecting secret (to the model) benchmark data to be reliably detected and filtered from training data seems... unlikely. Especially in this day and age where lots of people dislike AI enough that they would be interested in actively sabotaging efforts...