- https://news.ycombinator.com/item?id=26887670
1 = https://cse.umn.edu/cs/statement-cse-linux-kernel-research-a...
- https://cse.umn.edu/cs/open-letter-linux-community-april-24-...
- https://cse.umn.edu/cs/statement-computer-science-engineerin...
Imagine if it was taboo to independently test the integrity of bitcoin for example.
The sibling mentioned the linux kernel case. I admit that one felt wrong. It was a legitimate waste of contributor time and energy, with the potential to open real security holes.
I don't pretend to have reconciled why one seems right to me and the other wrong.
> The sibling mentioned the linux kernel case. I admit that one felt wrong.
> I don't pretend to have reconciled why one seems right to me and the other wrong.
The "how" is what matters here, not just the "what". "Testing the integrity of Bitcoin" by breaking the hash on your own machine (and publishing the results, or not) is one thing. "Testing" it by sending transactions that might drain someone else's wallet is quite another. Similarly with Linux, hacking it on your own machine and publishing the result is one thing. Introducing a potential security hole on others' machines is another. Similarly with water: messing with your own drinking water is one thing. Messing with someone else's water is quite another.
Playing devils advocate for a moment. How else do you test the robustness of the human process to prevent bad actors? Don’t you need someone to attempt to introduce a security hole to know that you are robust to this kind of attack?
It's a hard problem having fully editable storage by anyone, while maintaining integrity.
You sift through the edit log to find edits correcting factual errors.
Then you find the edit where the error was introduced.
You can probably let an LLM do the first pass to identify likely candidates. With maybe 20 hours of work you could probably identify hundreds of factual errors. (Number is drawn from a hat.)
In general what I'm saying is, this is a fertile ground for natural experiments. We don't need to manufacture factual errors in Wikipedia. They occur naturally.
But more importantly we probably won't catch that sort of error by introducing fabrications, either. Fabrications might replicate a class of error we're interested in, but if we just throw it onto Wikipedia, it's not going to be a longstanding misunderstanding which is immune to fact checking (at least without giving it a lot of time to develop into a citogenesis event, but that's exactly the kind of externality we're trying to avoid).
(Of course, "how many times do we need to replicate it?" remains unanswered. I think maybe after we have several replications and have data on false negatives by our teams of experts, we could come up with an estimate.)
How do you test that the White House perimeters are secure, or that the president is adequately protected by the Secret Service?
I've asked the author about ethical review and processes on the Fediverse.
That said, both Wikipedia and the Linux kernel (mentioned in another response to this subthread) should anticipate and defend against either research-based or purely malicious attacks.
Sometimes it will be worth it anyway, and I don't have an opinion about this Wikipedia example, but I think it's pretty uncontroversial that the Linux example was out of line.
* users are misled about facts * trust is lost in Wikipedia * other users/organizations use this as a blueprint to insert false information
Harm 3 seems to be the most serious, but I suspect it has happened/will happen irrespective of this research. As opposed to the water reservoir example, these harms seem quite small by contrast. I would have liked to see a section discussing this in the blog post, but perhaps that's included in the original paper.