And that's for just basic detection. There have been plenty of examples of 100% human written content (either old, or unpublished) that gets falsely flagged as AI.
And these are just technical aspects. The main issue is that pangram and other solutions are being used to summarily judge students work, and that is orders of magnitude more fucked up. Accusing someone of cheating can have devastating effects on their education/career/etc. and they're doing it with snake-oil closed boxes, at scale. We really really shouldn't support this, especially here on a technical site.
I only tested it with the two pieces, but was not impressed.
Edit: It was removed. Here is a gist: https://gist.github.com/jedberg/e24124e577c0da1c8445c22798eb...
For example they have a corpus of older pre-llm text and use that as the example their current model doesn't misclassify human written text. It shouldn't take much thinking to realize why this is a fucking stupid benchmark.
Every day humans use LLMs and read LLM content Panagram becomes more useless because it forces languages to have a stopping point sometime around 2020. If you adopt any LLMism or are one of those unlucky people that already talked like an LLM before LLMs then all your shit is getting marked even though it was created by the human mind and written by human hands.
I posted here on HN but it got removed.
Are you absolutely sure you wrote this by hand? If so it's kind of remarkable how close to an LLM you write like.
I’ve been accused of being an LLM multiple times here on HN too. I know you have no way to know for sure other than trusting that I’m not using an LLM to write. But it’s pretty frustrating that people jump right to LLM accusations.
I think discounting Pangram’s accuracy based on that sample isn’t very reasonable. Really nobody writes like that other than LLMs!
Today Pangram says it is 100% human, which is correct. But yet I got multiple DMs when I posted it 8 months ago saying "stop posting AI slop!" in response to that comment. At the time, Pangram marked it as 50% AI.
So to Pangram's credit, they got better.
People say it's in the training data, but I haven't seen any good examples of clearly dated pre 2023 text that sounds like this.
The ones that sound like LLMs tend to be the ones that were well researched and spent more time on, not the off the cuff stuff, which is most of what I write. So it would take me a while to find something like that.
But you're welcome to dive into my reddit and HN history, or all my blog posts on the wayback machine if you want to look for one. :)