What are is this problem from? What areas in general did you find useful to create such benchmarks?
May be instead of sharing (and leaking) these prompts, we can share methods to create one.
May be instead of sharing (and leaking) these prompts, we can share methods to create one.