It's a big additional leap to for our software to try to reach into others' sites, get through any anti-bot defenses they may be running, try to scrape their content and evaluate on whether it's sufficiently human-authored to be on HN.
There's generally a wider range of LLM involvement with a long-form post than the typical, relatively brief HN comment, which then opens the way for more debate on HN about "how much" LLM influence the post has and how much should be allowed on HN. Part of what we're trying to optimize for on HN is minimizing offtopic/meta discussion, so we don't want to encourage this kind of debate.
Our heuristic about article quality is largely unchanged from before LLMs were an issue: if an article is badly written, it shouldn't be on HN, and should be flagged.
As to what should be tested: front-page items, possibly even a subset of those (top 10--15 of 30). That's going to be a limited set of items per day, though more than just 30. (I don't know how many items cycle through the front page on a daily basis, though I believe daily submissions as of 2022 were about 1,000/day (<https://web.archive.org/web/20220116193045/https://whaly.io/...>)).
Working this into the HN story-processing lifecycle might be a good call.
I'd much prefer not seeing a bunch of AI slop in submissions, by way of generated output. AI as part of the resarch process I think I could live with.
AI-generated content seems, definitionally, not to be intellectual in nature, and would seem to go against HN's prime directive. It also seems to make HN lose its collective mind, which has long been another mod consideration.
In the meantime, please feel free to flag items that are badly written/unpleasant to read, and email us if something is on the front page that shouldn't be there.