I think I've discovered that the target of the prefetch is irrelevant, and that what matters is that the store-to-load forwarding does not try to execute in the same cycle. For Sandy Bridge, I'm finding that I can get slightly better performance with a dummy load of an unrelated volatile variable:
volatile unsigned dummy = 0;
void loop_dummy_read() {
IACA_START;
unsigned j;
unsigned dummy_read;
for (j = 0; j < N; ++j) {
dummy_read = dummy;
counter += j;
}
IACA_END;
}