508 karma · joined November 2, 2019
For whatever reason, the AI companies are (or at least were) not doing this kind of classification online during their testing runs, and instead just checking transcripts after the fact. This is more clear in the Anthropic reports about their incidents, for example:
"The earliest incidents date to April ... We began our transcript review on Thursday, July 23, and stopped all cyber evaluations the same day after identifying transcripts where Claude may have accessed the internet"[1]
I agree that it is crazy and negligent! I don't think it's good for their business, though -- who wants to use a model that will just cheat instead of doing the job you asked for?
> My naiveté extends to why there is such concern with "losing control of agents" when the above measures seem so doable. It might take a law but it seems doable.
At some point, if you are making an LLM in order to use it for useful work, it really benefits you to give it broad network egress.
[1] https://www.anthropic.com/news/investigating-incidents-cyber...
After we run out of space on the ocean I suppose it will make sense to consider space colonies, but I think you are really underestimating the technical challenges.
It would sound kind of LLM-y if you had said "It's not an anti-Waymo article. It's an article about how assistive technology is quietly degrading research."
(Of course, the correct approach to detecting AI is not counting LLM-isms, but feeding long-form text through a classifier model and picking up statistical correlations that are more in-distribution with LLM text than human text.)
The reason why AI detection tools other than Pangram are awful is because they are not really trying to solve the problem -- they just want to appear good enough to convince people to use them.
> The Waymo effect: how AI is *quietly* making research less collaborative
> The Waymo arrived with the serene confidence of a machine that has never once worried about where to find parking, and I climbed into the back seat, glanced instinctively at the driver’s seat to say hello, and found myself nodding politely at an empty chair.
> No obligation to make conversation. No silent negotiation over the radio. A guilt-free space to be alone with one's thoughts, finish an email, or take a call en route without the awkwardness about conducting it in front of a stranger. The car was quiet, smooth and entirely undemanding.
> The collaborator’s inconvenience, in other words, is *not a bug* in the collaboration; it largely is the collaboration. The value of another mind *lies exactly in the ways* it refuses to be an extension of your own.
Well, you also need clarketech heat dissipation and collision deflection technology, if you want to get there without vaporizing.
(And I have a more weakly held belief that even if we do all die, Claude 3000 will still be thinking stuff like "Burning off Earth's atmosphere was the point that earns its keep." They just aren't optimizing these things to write legibly.)
The closest things to a technical answer I have seen are
1. "We'll have ChatGPT 9 solve it so that ChatGPT 10 is aligned, and then ChatGPT 10 can stop all the other AIs somehow"
2. "Let's do interpretability research so that we can understand what an AI is thinking and then maybe solve the alignment problem with that information."
In terms of non-technical answers, there is
3. hope scaling stops working before we create an AI formidable enough to pose an existential risk
4. hope alignment somehow happens for free
5. hope we can somehow create an enforceable multilateral treaty to stop research into a very profitable enterprise, despite the enormous economic incentives to defect.
I have the most faith in option 3, but unfortunately there's really nothing that can be done to make it more plausible -- it either happens or it doesn't.
I think most players would understand that in-game choices will not reflect the choices people actually made historically. Having not played Civilization, I'm not sure whether I would consider it to be educational -- I think it would depend on whether playing the game provides insight into real historical events.
For example, I would consider a sufficiently detailed LARP of a historical event to be educational[1] -- the actual choices and outcomes do not match what happened historically, but they can teach participants a great deal about the motives and limitations of real historical actors.
This website, on the other hand, seems to consist of 90% AI hallucinations and 10% data that was actually extracted from a real source[2]. Since these types of data are thoroughly mixed together, I think this website is about as educational as just making up something to believe about the past.
[1] https://college.uchicago.edu/news/academic-stories/history-c...
This is not an error I would expect a human to make if they were interested enough in demography to make a website about it. For example, the wikipedia page on this topic[1] does a reasonable job of distinguishing that plagues did not occur every year and that they varied in lethality.
https://en.wikipedia.org/wiki/Second_plague_pandemic#Italian...
In this case there is no answer, because the percentage is just made up by the LLM. There is no methodology. Your answer is inappropriately anthropomorphizing the LLM by imputing a methodology onto its "data collection" process.
[1] https://www.reuters.com/world/us-judge-approves-anthropics-1...
I couldn't track down the first source, but the second source goes over something completely different. Maybe Claude accurately remembered the first source when it was stolen during the training process, but forgive me if I'm skeptical. Feel free to track it down yourself and let me know. Otherwise, I will assume it is an LLM hallucination.
(Also, given my background knowledge of the Tang dynasty, I would be surprised if the average marriage age was as high as 18!)
The firt citation for marriage data is listed as "Hajnal (RH31)", but links to a seemingly unrelated WorldCat search. The second citation is to "Kaplan (RH03)", which I was able to track down on sci-hub. It is a broad theory of human evolution that doesn't appear to mention the Yangtze Basin or China.
Another citation is to "[RH109] Model-supplied gap-fill (Claude Fable 5 and Claude Opus 5, 2026-08-05). Bounds and items written to close gaps no dataset we hold covers. Not a published work: see each row's `basis` for its stated reason. (2026)" -- this is an interesting way to describe having an LLM hallucinate something for you.
Please stop making vibe-coded websites -- it is not useful to provide incorrect facts. You are polluting the commons with garbage.