New accounts on HN more likely to use em-dashes
marginalia.nu
marginalia.nu
Now someone may search old posts without a time cutoff and assume I'm an LLM. That combined with the fact I sometimes write longer posts and naturally default to pretty good punctuation, spelling and grammar, is basically a perfect storm of traits. I've already had posts accused twice in the past year of being an LLM.
Kind of sad some random quirk of LLM training caused a fun little typography thing I did just for myself (assuming no one else would even notice) to become something negative.
(This above line itself was written by AI itself: https://www.kimi.com/share/19c96516-4032-8b73-8000-0000f45eb...)
I don't know if worse grammar could make a difference aside from removing false negatives (ie. nowadays people with good grammar are questioned if they are LLM's or not) but this itself doesn't mean that worse grammar itself means its written by a human. (This paragraph is written by me, a human, Hi :D)
Also adding better "context" into the discussion, than the usual claims/punchlines of marketing-speak.
Maybe it's not exactly the grammar itself but also overall structuring of the idea/thought into the process. The regular output sounds much more like marketing-piece or news-coverage than an individual anyway. I think, people wanna discuss things with people, not with a news-editor.
If I understand you correctly, then Yes I completely agree, but my worry is that this can also be "emulated" as shown by my comment by Models already available to us. My question is, technically there's nothing to stop new accounts from using say Kimi and to have a system prompt meant to not sound AI and I feel like it can be effective.
If that's the case, doesn't that raise the question of what we can detect as AI or not (which was my point), the grand parent comment suggests that they use intentionally bad human writing sometimes to not be detected as AI but what I am saying is that AI can do that thing too, so is intentionally bad writing itself a good indicator of being human?
And a bigger question is if bad writing isn't an indicator, then what is?
Or if there can even be an good indicator (if say the bot is cautious)? If there isn't, can we be sure if the comments we read are AI or not
Essentially the dead-internet-theory. I feel like most websites have bots but we know that they are bots and they still don't care but we are also in this misguided trust that if we see some comments which don't feel like obvious bots, then they must be humans.
My question is, what if that can be wrong? It feels to me definitely possible with current Tech/Models like say Kimi for example, Doesn't this lead to some big trust issues within the fabric of internet itself?
Personally, I don't feel like the whole website's AI but there are chances of some sneaky action happening at distance type of new accounts for sure which can be LLM's and we can be none the wiser.
All the same time that real accounts are gonna get questioned if they are LLM or not if they are new (my account is almost 2 years old fwiw and I got questioned by people esentially if this account is AI or not)
But what this does do however, is make people definitely lose a bit of trust between each other and definitely a little cautious towards each message that they read.
(This comment's a little too conspiratorial for my liking but I can't help but shake this feeling sometimes)
It just is all so weird for me sometimes, Idk but I guess that there's still an intuition between whose human and not and actually the HN link/article iteslf shows that most people who deploy AI on HN in newer accounts use standard models without much care which is the reason why em-dashes get detected and maybe are good detector for sometime/some-people and this could make the original OP's comment of intentionally having bad grammar to sound more human make sense too because em-dashes do have more probability of sounding AI than not :/
It's just this very weird situation and I am not sure how to explain where depending on from whatever situation you look at, you can be right.
You can try to hurt your grammar to sound more human and that would still be right
and you can try to be the way you are because you think that models can already have intentionally bad grammar too/capable of it and to have bad grammar isn't a benchmark itself for AI/not so you are gonna keep using good grammar and you are gonna be right too.
It's sort of like a paradox and I don't have any answers :/ Perhaps my suggestion right now feels to me to not overthink about it.
Because if both situations are right, then do whatever imo. Just be human yourself and then you can back down this statement with well truth that you are human even if you get called AI.
So I guess, TLDR: Speak good grammar or not intentionally, just write human and that's enough or that should be enough I guess.
"respond like a twitter user", "pretend like we're texting", etc
> "respond like a twitter user", "pretend like we're texting", etc
+1 to it. I actually had given a response to the above parent comment itself using Kimi and I would've said that its (sort of) a good emulation fwiw.
>I put the em dash on modifier+dash
This is the default on Macs
I also like …
This is like ruining swastikas and loading rainbows
This makes me think of the fad where people on youtube will hold a microphone up in frame, because it somehow connotes authenticity. I'm sure some people are already embracing a bit of sloppiness in their writing as a signal of humanity; I'm equally sure that future chatbots will learn to do the same.
This applies not only work-stuff itself also to the job-applications/cv/resume and cover-letters.
Yes I enjoy lisp, how could you tell
It's one of those things I think are worth putting some extra effort into, I'm glad to see at least one other person giving it some thought. Thx <3
<li> do this
<li> and this
instead of: <li> ... </li>and <img alt='this'> instead of <img ... />
You might like Lisp, but what you're saying reminds me of the late 00s/early 2010s xHTML2 vs. HTML5 debate :)
[0] :))
Turned Semicolon (U+2E35 ⸵)
you can find him on windows by pressing Win + ; not as fast as typing, but quite faster then typing and then wondering if thats too much brackets or too little
"Вот его, нет, не допустили (сама знаешь, почему)))"
My translation:
"But him - no, they didn't let him in (of course you know why :)"
When I went from texting friends in Russian or Ukrainian back to English, I missed right parentheses as a smiley; one or two - hi), hello)) - to me are like a smile, by ))) and )))) there's some laughing or some other joke going on. Native speakers could weigh in; my native tongue is English.
> I started making deliberate grammar and spelling mistakes in professional context[s]. Not like I have ~a~ perfect writing anyway, but at least I could prove that it was self-written, not an auto-generated slop. (Could be self-written slop though :)
> This applies not only [to] work-stuff itself also to the job-applications/cv/resume and cover-letters.
I conclude you are real.
if u want it to sound more real u just gotta tell the bot to write that way. like literally just ask it to throw in some typos or forget to capitalize stuff. or use slang and kinda ramble instead of being all robotic and organized.
I've also noticed an increase of this in myself and others, I used to edit a lot more before sending anything, but now it seems more authentic if you just hit send so it's more off the cuff with typos, broken sentences and all.
I'm sure an LLM could easily mimic this but it's not their default.
Now you need a really big microphone, something that looks like it was built in 1952.
- Customer: Excuse me, I'm looking for the Aunt Jemima maple syrup. Can you point me in the right direction?
- Employee: y u ask like chatbot
There is a decent chance the term "Employee" as a whole will be eradicated sometime in the next 10 years.
BTW in their company chat they call it a squiggly even though it's flat. That always bugged me.
Edit: I stand corrected. It wasn't flat from 1966 to 1981 and the cheer started in 1975, and included "squiggly" back then - Sam Walton himself said it. https://en.wikipedia.org/wiki/Walmart#1990%E2%80%932005:_Ret...
Edit 2: Both squiggly and the introduction of the cheer appear in the Walmart timeline: https://corporate.walmart.com/about/history
https://www.pearlmillingcompany.com/our-history
"In June 2020, PepsiCo and The Quaker Oats Company made a commitment to change the name and image of Aunt Jemima, recognizing that they do not reflect our core values.
We want to thank everyone who has made us part of their family over the years, and look forward to starting a new chapter as the Pearl Milling Company."
It's a massive trend all over IG and TikTok these days as there's a lot of mobile-friendly consumer gear that brings your Social Media Content up a fairly large way for fairly low cost. Ironically, the mimicry probably spreads because those that know what they're doing partially or fully conceal it which is by definition less likely to be noticed and therefore imitated by first timers/novices.
Of course, Chinese characters are not pictographic and haven't been for a few thousand years, but they are still largely ideographic.
What about reverting to referencing moving pictures?
Or was this intended to be self-referential and I missed the joke? :)
You’re talking about ancient technology here…
Got several comments saying they were "AI slop."
Even had a screen cap of my drawing process.
Kinda funny to think my drawings, which have likely "trained" AI image generators, are now getting accused of being AI.
I wonder if there's a way we can communicate that LLMs fundamentally can't keep up with. If LLMs have hit on being exactly the way our brains work then I guess not. But maybe we still have something special. I haven't tried how well LLMs understand language written like in Iain M. Banks's Feersum Endjinn.
Ed: I'm sure there's cargo culting going on but the visibility of the mic isn't only performative.
This lead to other people clipping them onto random objects to make fun of the trend for a while.
I dunno this en versus em dash stuff, I just use the minus sign on my keyboard.
Entire sentence structures have been effectively blacklisted from use. It's repulsive.
Well, to be fair Gen-z slangs also have a massive impact. My generation sometimes point blank said to me that they didn't have the attention span to read my sentence :/
Definitely picked up a few slangs along the way now. I had to somehow toggle a switch between how I write on HN/how I write with my friends the first few times and I write pretty informally in HN, but its that you got to be saying lowk bussin rizz 67 to make sense.
My friends who use insta literally had Abbreivations which were of 9 letter words in my own language that the insta community of my nation's gen-z sort of made.
Although I would agree that we haven't seen a whole unicode being thrown this way in ALL generations (I feel like universally everyone treats em-dashes as something written by AI or definitely get an AI alert)
But I think that 67 is something that atp maybe even most adults might have gotten exposed to which has probably changed the meaning of number.
You’d think ethically leaving it in would be better. But we’re talking about big tech companies here.
Speaking of overusing something until it becomes cringe, has anyone shown their kids Firefly? Does it still hold up after the Joss Whedon signature bathos (and other tics) became a tentpole of the Marvel Cinematic Universe and created an abundance of cultural antibodies?
There were a few times we cringed a bit (with both shows) but overall stood the test of time. I didn't watch Buffy & Angel first time around, so it was a bit of a cultural moment I got caught up on. And it was nice to revisit Firefly, the little bit of it we got.
There is no such thing as blacklisted by other commenters.
> A new way, a better way.
The autumn winds blow.
That's one of the signals I use to detect if YouTube videos are AI slop. If it's narrated by a non-native speaker, it's much more likely to be high quality. If it's narrated by a British voice with a deep timber, it's 100% AI.
Now I find myself deliberately making things worse to avoid being accused of not being human! Bah!
I wonder how much crossover there would be between a trained text analysis model looking for Gen-X authors and another looking for LLM's.
But that's a different issue.
Tip: Patterns like “It’s not just X, it’s Y” are a more telltale sign of LLM slop. I assume they probably trained on too much marketing blurb at some point and now it’s stuck.
• Like
• This
(option-8 on a Mac US keyboard layout). Now it looks like something only an LLM would do.
...then immediately afterwards a new section named "Compatibility" starts where the use of code points that are composites of several letters, e.g. ℡ and ℻ are indeed discouraged, suggesting they be spelled out in full as TEL and FAX instead.
Do you feel ℃ falls into a continuation/overlap between these two sections?
My double-space-after-a-period though, I will keep that until the end. Even if it often doesn't even render in HTML output, I feel a nostalgic connection to my 1993 high school typing teacher's insistence that a sentence must be allowed to breathe.
I use em dashes, and I don't care whether or not someone assumes I'm an LLM. Typography exists for a reason.
If leaving out the Oxford comma here was an intentional joke I both commend and curse you!
I’m waiting for a Philip K. Dick bot to declare me non-human.
Am I the only one who in a Captcha test sometimes wants a different option for the “I am Human” check box? Ironically really since to prove we’re human we have to check the boxes with a crossing in them, no account to be made of people who call them zebra crossings.
My phone lets me long-press the hyphen key to get an em-dash so sometimes I'll use it.
Probably the biggest tell that I'm not AI is that I'm probably not using it in the appropriate circumstances!
I've noticed a habit of late of people accusing a comment of being LLM generated if they disagree with it. It was getting quite tiresome a few weeks ago but seems to have died down.
I suppose it is possible that they are actually LLMs making the accusations? :-)
(I'm one of those weirdos that try to use proper grammar and complete sentences in text messages and instant messages.)
Broken sentences. Also useful. Like in some literature works.
It's so sad to me that good typographical conventions have been co-opted by the zeitgeist of LLMs.
If AI was writing like everyone else we wouldn't be talking about this. But instead it writes like a subset of people write, many of them just some of the time as a conscious effort. An effort that now makes what they write look like lower quality
Say what you want about marketing-isms of your typical LLM, they have been trained and often succeed at making legible, easy to scan blobs of text. I suspect if more LLM spam was curated/touched up, most people would be unable to distinguish it from human discourse. There are already folks commenting on this article discussing other patterns they use to detect or flag bots using LLMs.
(Until a few years ago I probably mostly only saw them in print, and I suppose it just never occurred to me that I liked them in particular vs. just the whole book being professionally typeset generally.)
This is the first time I've ever heard the character ";" referred to as such. It's always been "semi-colon" to me, is this a region/culture difference?
I'm not saying you're wrong, I find it interesting.
i call it a super comma when its separating a list with commas within the sets.
so if i am listing colors like green, blue, red; foods like apple, orange, strawberry; and seasons like winter, summer, fall.
it's one use case for an em-dash, because whatever you have inside it has commas in the phrase.
square and rectangle situation. a supercomma is a subset of semicolon.
I would have assumed it's a synonym for apostrophe. super-comma <-> upper-comma, with super meaning upper, like in superscript.
Em-dash matches how I speak and think-- frequently a halt, then push onto the digression stack, then pop-- so I use them like that.
Em-dash matches how I speak and think (frequently a halt, then push onto the digression stack, then pop) so I use them like that.
Em-dash matches how I speak and think, a halt, then push onto the digression stack, then pop, so I use them like that.
Em-dashes keep everything on the same level of importance in my brain.
Commas don’t feel as powerful. To be fair to the comma I’d probably do this:
Em-dash matches how I speak and think: A halt, then push onto the digression stack, then pop. So I use them like that.
Edit: I accidentally used an em-dash in the word em-dash. Interestingly HN didn’t consider changing the dash to be a change in my text so didn’t update it. I had to make a separate change and take that change out for my dash change to stick.
Well, I haven't always—just for maybe 20 years.
This wouldn't be an issue if mobile users or Windows users were exercising it too, but it's just Mac owners and LLMs. And Mac owners are probably the minority of instances where it is used.
so now, i just use double dashes for everything.
(shit, i wonder when llms will start doing this instead of normal em)
Then came LLMs, and there was so much talk of them using em dashes. A few weeks ago, I finally decided it's time and learned the difference. (Which took all of 2 minutes, btw.) Now I love em dashes and am putting them everywhere I can! Even though most people now assume I'm using AI to write for me.
I defer to Merriam-Webster and/or Harbrace (rather than TCMoS) on punctuation usage.
https://www.merriam-webster.com/grammar/em-dash-en-dash-how-...
Magical signal panacea searching is ultimately fruitless. Other ways to make bot interactions more difficult, there are policy and technological obstacles that could be introduced. For example, require an official desktop or mobile app for interaction. And then for any text copy-pasted, demarcate it. And throw an error message for any input typed inhumanly-fast. Require a micropayment of like $0.10 to comment. While these things would break the interaction style and flexibility for a lot of innocent human users, these would throw big wrenches into some but not all vulnerabilities of bot interactions.
Most think that it came from TeX, which had -- (for an en dash) and --- (for an em dash, although I don't think I have ever observed it out in the wild outside TeX), but in fact, the habit well predates TeX and goes all the way back to typewriters where typists habitually hit two hyphens in a row to approximate an em dash. The approximated em dash was described in hard-copy manuscript preparation rules such as The Chicago Manual of Style.
So, if you have ever used a typewriter or TeX, you can claim an even richer than 20 years’ heritage of using the em dash.
But anyways, you can't really control how people see your stuff, if you're human I think the humanness will come through anyways, even if you have some particular structure or happen to use em-dashes sometimes. They're so easy to prompt around anyways, that the real tricky LLM stuff to detect by sense and reading is the stuff where the prompter been trying to sneakily make them more human.
If you'd like more tips on writing I'd be happy to help.
Edit: I take that back. I'm going to print and frame this comment. It stands on its own well enough, and I'm the only one who's going to see it.
Second Edit: Took a bit to get it formatted in a way I liked, but I have officially placed an order for my local Walmart photo center
But now, I have to be so picky about when I use them, even when I think it's the perfect punctuation mark. I'll often just resort to a single hyphen with spaces around. It's wrong, but it doesn't signal someone to go "AI AI AI!!"
It's like being named Michael Bolton and watching a singer rise in fame named Michael Bolton.
Why should I change my style?
For those who don’t know the reference:
<https://en.wikipedia.org/wiki/Alex_S._Jones>
(No, not that one.)
Office Space jokes aside, you shouldn't. I think it's very important to be yourself and refuse to let people pressure you into changing for no good reason. I am not an em dash user myself, as it's a pain to generate when there's no key on the keyboard for it. But if I were, you best believe I wouldn't change my style one bit. People can accuse me of being an LLM if they wish, but that's no skin off my back.
Unless you're talking about restructuring your sentences to allow for a semicolon; that's fine.
For example that semicolon could have been an em dash, but I don't think it's the type that LLMs over favor.
I've typeset books (back in the QuarkXPress days, before Adobe's InDesign ruled the typesetting world) and never bothered with em-dashes. Writing online is, to me, a subset of ASCII. YMMW.
But the one thing I don't understand is this: how comes people using LLM outputs are so fucking dumb as to not be able to pass it through a filter (which could even be another LLM prompt) that just says: "remove em-dashes, don't use emojis, don't look like a dumb fuck".
Why oh why are those lazy assholes who ruin our world so dumb that they can't even fix that?
It's facepalming.
Here since 2010 in this account, I use em-dashes.
It's easy—and effective—to type using “Opt Shift -” on a Mac.
Oh yeah, left and right “curly quotes” as well, and the occasional …
> It's so sad
Don’t forget «’» — but ain’t nobody got time for that!
A few more to reclaim typography: https://howtotypeanything.com/alt-codes-on-mac/
I turned to my friend and said "They've co-opted the structure of effective language!"
Em dashes, semicolons, deftly delving. It’s all just so…facile. We might as well tell ourselves we can tell it’s shopped from the pixels, having seen some shops in our day.
(I know it's a bit low effort, but if ever something called for "based" it's this.)
(Disclaimer: the use of em-dashes doesn't prove this was AI-generated — I can assure you I wrote this myself.)
And I will still use them -- fully aware that some people will complain about AI and whatnot.
word noob new p-value
----------------------------
ai 14.93% 7.87% p=0.00016
actually 12.53% 5.34% p=1.1e-05
code 11.47% 6.04% p=0.00081
real 10.93% 2.95% p=2.6e-08
built 10.93% 2.11% p=2.1e-10
data 8.93% 3.51% p=6.1e-05
tools 7.6% 2.67% p=5.5e-05
agent 7.47% 2.95% p=0.00024
app 7.2% 3.09% p=0.00078
tool 6.8% 1.83% p=8.5e-06
model 6.8% 2.39% p=0.00013
agents 6.67% 2.11% p=5.2e-05
api 6.53% 1.12% p=2.7e-07
building 6.13% 1.54% p=1.3e-05
full 6.0% 1.97% p=0.00017
across 5.87% 1.4% p=1.3e-05
interesting 5.33% 1.54% p=0.00014
answer 5.2% 1.4% p=9.6e-05
simple 4.93% 1.54% p=0.00043
project 4.8% 1.26% p=0.00015e.g. "The body of the template is parsed, but not actually type-checked until the template is used." -> "but not typechecked until the template is used." The word "actually" here has a pleasant academic tone, but adds no meaning.
I'm totally fine with the word itself, but not with overuse of it or placing it where it clearly doesn't belong. And I did that a lot, I think. I suspect if you reviewed my HN comments, it's littered with 'actually' a ton. Also "I think...", "I feel like..." and other kind of... Passive, redundant, unnecessary noise.
Like, no kidding I think the thing I'm expressing. Why state that?
Another problem with "actually" is that it can seem condescending or unnecessarily contradictory. While I'm often trying to fluff up prose to soften disagreement (not a great habit), I'm inadvertently making it seem more off-putting than direct yet kind statements would. It can seem to attempt to shift authority to the speaker, if somewhat implicitly. Rather than stating that you disagree along with what you believe or adding information to discourse, you're suggesting that what you're saying somehow deviates from what the person you're speaking to would otherwise believe or expect. That's kind of weird to do, in my opinion. I'm very guilty of it, though I never had the intent of coming across this way.
It can also seem kind of re-directive or evasive at times, like you don't want to get to the point, or you want to avoid the cost of disagreement. It's often used to hedge statements that shouldn't be hedged. This is mainly what led me to realize I should use it less. I hedge just about everything I say rather than simply state it and own it. When you're a hedger and you embed the odd 'actually' in there, you get a weird mix of evasive or contradictory hedging going on. That's poor and indirect communication.
One reason might be to acknowledge that you're not being prescriptive, but leaving room for a subjective POV in situations that call for it.
Likewise, the GP's use of "actually" acknowledges the contrast between what one might expect (that some preliminary type-checking might happen during initial parsing) and what in fact happens (no type checks occur until the template is used.) It doesn't seem out of line in that case.
I agree but it's not always clear whether you're stating an opinion or attempting to state a fact. Some folks would reply to a comment like this with "citation needed" but wouldn't otherwise have said that if the comment had opened with "I think."
Lately "I mean" has been jumping out at me.
It really only bothers me when I notice I've used it for multiple comments in the same thread or, worse, multiple times in the same comment.
I've also pretty much dropped just from my vocabulary when I'm talking about an alternative way to do something.
Do all the models have this style of talking? Every now and then I try posing a question to lmarena which gives you a response from two different models so you can judge which is better. I feel like transitions like "The real answer...", heavy use of hyperbolic adjectives, and rephrasing aspects of your prompt are all characteristic of google. Most other models are much more to the point
I have a quick question but can you please tell me by what's the age of "new" accounts in your analysis?
Because, I have been called AI sometimes and that's because of the "age" of my comments sometimes (and I reasonably crash afterwards) but for context, I joined in 2024.
It's 2026 now, Almost gonna be 2 years. So would my account be considered new within your data or not?
Another minor point but "actually"/"real" seems to me have risen in usage over 5 times. All of these words look like the words which would be used to defend AI, I am almost certain that I saw the sentence "Actually, AI hype is real and so on.." definitely once, maybe even more than once.
Now for the word real, I can't say this for certain and please take it with a grain of salt but we gen-z love saying this and I am certain that I have seen comments on reddit which just say "real" and OpenAI/other models definitely treat reddit-data as some sort of gold for what its worth so much so that they have special arrangements with reddit.
So to me, it seems that the data has been poised with "real". I haven't really observed this phenomenon but I will try to take a close look if chatgpt is more likely to say "real" or not.
Fwiw, I asked Chatgpt to "defend the position, AI hype sucks" and it responded with the word "real"/"reality" in total 3 times.
(another side fact but real is so used in Gen-z I personally watch channel shorts sometimes https://www.youtube.com/@litteralyme0/shorts which has thousands of videos atp whose title is only "real", this channel is sort of meme of "ryan gosling literally me" and has its own niche lore with metroman lol)
noob = new user
new = I think this might be a mistake? Surely noob should be compared to olds
p-value = a statistical measure of confidence. In academic science a value < 0.05 is considered "statistically significant".
/noobcomments vs /newcomments. New is new as in recent.
I got similar feeling. I'm new here, but got a feeling that some comments are like bot generated.
Such low p-values are proof that something is going on.
Hipotesis (after your recent word statistics): that some bots are "bumping up" AI related subjects. Maybe some companies using LLM tools want to promote some their products ;)
marginalia_nu respect for your work :)
The idea is, since data has a ~1/20 chance of having a p < 0.05, you are bound to get false positives. In academia it's definitely not something you'd do, but I think here it's fine.
@OP have you considered calculating Cohen's effect size? p only tells us that, given the magnitude of the differences and the number of samples, we are "pretty sure" the difference is real. Cohen's `d` tells us how big the difference is on a "standard" scale.
Are you saying p is uniformly distributed over any data set? That doesn't jive with my limited understanding of entropy. What's this based on?
Your comment about p<0.05, feels out of place to me. The p-values here are << 0.05. Like waaaaay lower.
Perhaps Fisher's exact is more appropriate, on the per-word basis?
> One of the simplest approaches to correct for multiple testing is the Bonferroni correction. The Bonferroni correction adjusts the alpha value from α = 0.05 to α = (0.05/k) where k is the number of statistical tests conducted. For a typical GWAS using 500,000 SNPs, statistical significance of a SNP association would be set at 1e-7. This correction is the most conservative, as it assumes that each association test of the 500,000 is independent of all other tests – an assumption that is generally untrue due to linkage disequilibrium among GWAS markers.
https://journals.plos.org/ploscompbiol/article?id=10.1371/jo...
IMO a more interesting experiment would be to show comments to people (that haven't seen these conclusions), and have them assess whether they suspect them of being bots or AI authored, and then correlate that with account age.
How many new accounts are submitting github links as their first post?
How many new accounts include a first comment that is copied from the other side of the link?
Look at the timings between first commit, last commit, and account creation. Many happen in quick succession and in this order. Fastest I've seen so far is 25m from first commit to first post on HN, with account creation in between.
You can explore the underlying data using SQL queries in your browser here: https://lite.datasette.io/?url=https%253A%252F%252Fraw.githu... (that's Datasette Lite, my build of the Datasette Python web app that runs in Pyodide in WebAssembly)
Here's a SQL query that shows the users in that data that posted the most comments with at least one em dash - the top ones all look like legitimate accounts to me: https://lite.datasette.io/?url=https%3A%2F%2Fraw.githubuserc...
Ellipses were never part of the analysis.
Looks like no. It's probably Safari that "knows better".
> select user, source, count(*), ...
it's clear that every single outlier in em-dash use in the data set is a green account.
There is precedent here.
"We treat this one better because it's a house clanker instead of a field clanker"
"If the clanker acts up it knows that it gets stuck in the box"
It was meant to be funny but definitely highlighted exactly what you are saying.
If you look back to the 90s and see someone using a racist slur, you fill in the gaps and assume they were using it because they were racist.
Will people in 30 years look back to today and judge those who showed disdain for people who rely on AI to write for them?
Even if clanker becomes a no-no word 30 years from now, it seems beyond the realm of possibility that people who hated clankers in 2026 will be looked upon harshly. Clankers aren’t a marginalized group today, they aren’t a class that needs protection.
What words are you thinking of when you say that there is precedent?
There are people are judging your character for using such terms today. Their existence is not in doubt. It is only the future prevalence of the opinion that is in question.
>it seems beyond the realm of possibility that people who hated clankers in 2026 will be looked upon harshly
Thus spoke many people in history who acted with impunity.
cool cool cool
Requiring proof of identity is the only solution I can think of, despite how unappealing it is. And even then, you'll still have people handing their account over to an LLM.
I really struggle to imagine a way around it. It could be that the future is just smaller, closed groups of people you know or know indirectly.
One of the things HN does is not let you interact in certain ways until you've earned sufficient karma. This is a basic proof-of-work. If your bot can't average a positive karma, then it'll never get certain privileges.
Not to say the system is perfectly tuned for bots, because it's not. The point is that proof of identity is not the only option.
HN is doing okay at the moment because nobody is yet publishing ebooks and videos on how to astroturf HN to launch your SaaS. Unfortunately, Reddit hasn’t escaped that fate.
Many of them sound and look completely normal and have others on here interacting with them. They don't use em dashes, sometimes they'll use all lowercase text, sometimes the owner of the bot will come out and start commenting to throw you off.
All examples I've witnessed here.
HN should immediately start implementing at least some basic bot detection methods without requiring us to email them every time. I've discovered multiple bots make detailed comments within 30 seconds of each other in different threads, something a normal human wouldn't be able to do. That should be at least flagging the account for review. Obviously they'll get smarter and not do that soon but it would help in the short term.
I'd say it's not an issue but everything I described above has happened in less than a month and every day now I'm discovering bots here.
My best understanding is yes -- there are signal that somebody is a bot (like how quickly they post), but if HN bans based on those signals then whoever made the bot will keep tweaking the code.
I feel like I rarely see bots in the top 5 comments of any article I read, or otherwise causing major disruption.
I think we just need to get creative about ways a platform can prove somebody is an invested human without tying it back to any personally identifiable information.
Same. I agree that it is unappealing but it can be done in a way that respects anonymity.
I built this and talk about it here: https://blog.picheta.me/post/the-future-of-social-media-is-h...
I think we’re on the precipice of this being a requirement to have any faith you’re talking to another human. As a side effect it also helps avoid state actors from influencing others.
Except that it doesn't prove you're talking to a human - it just increases the hurdles for bot operators (buy or steal verified accounts).
Regarding your implementation: Most people don't have a passport, so it's a non-starter - but again, this topic is not a technical issue.
I don't see that as "requiring ID".
I think the real question is how much do we care that our online spaces are composed of not just AI bots, but also sock puppet accounts controlled by various people (from governments, rich people, all the way to harassers that use alt accounts) wanting to trick us.
My conspiracy theory: Campaign money, from the last few elections (I think "Correct the record" [1] was the first "disclosed" push), resulted in a bunch of bot accounts being made/bought all across social media. These are being lightly used to maintained some reasonably realistic usage statistics, and are "activated" to respond to key political topics/times. This is on top of spam accounts to push products and, of course, the probably higher-than-average bot number of accounts, made for fun, by HN users.
Exactly. So what's proof of identity good for?
That'd drastically reduce the amount of low effort posts, both human-written and generated.
It's certainly not perfect, but similar to what you mention.
Or, since this would need to be done in javascript, just block or rewrite the javascript and fake the output in the sent request.
Simplistic solutions like this stopped being meaningful decades ago.
E.g. I make a new hackernews account, and say "just ask wikipedia, they will vouch for my new hackernews account". Then wikipedia checks if any of their accounts vouch for this new hackernews account. If a user with enough reputation on Wikipedia (e.g. your friends or one of your own wikipedia accounts) vouches for this new hackernews account then wikipedia tells hackernews "yes, that account is legit".
Hackernews knows the minimum amount possible about the new account. And while wikipedia knows something, they know WAY LESS than a full ID check. People can have multiple Wikipedia accounts.
And its a two way street; Wikipeida could ask hackernews about new accounts. Both sites would benefit from the collaboration.
Karma could actually become meaningful/useful for reputation checks.
The only unfortunate aspect is I'm not aware of any software tooling for such a system.
>this is [summary]
>not just x, it's y
>punchy ending, maybe question
Once you know it's AI it's very obvious they told it to use normal dashes instead of em dashes, type in lowercase, etc., but it's still weirdly formal and formulaic.
For example from https://news.ycombinator.com/threads?id=snowhale
"this is the underreported second-order risk. Micron, Samsung, SK Hynix all allocated HBM capacity based on hyperscaler capex projections. NAND fabs are similarly committed. a 57% reduction in projected OpenAI spend (.4T -> B) doesn't just affect NVIDIA orders -- it ripples into the memory suppliers who shifted capacity to HBM and away from commodity DRAM/NAND. if multiple hyperscalers revise down simultaneously you get a situation similar to the 2019 crypto ASIC overhang: companies tooled up for demand that evaporated. not predicting that, but the purchasing commitments question is real."
I'll actually post a comment or question and I'll get a reply with a bit of a paragraph of what feels like a very "off" (not 'wrong' but strangely vague) summary of the topic ... and then maybe an observation or pointed agenda to push, but almost strangely disconnected from what I said.
One of the challenges is that yeah regular users don't get each other's meaning / don't read well as it is / language barriers. Yet the volume of posts I see where the other user REALLY isn't responding to the other person seems awfully high these days.
I wonder if it is neural networks that are inherently biased, but in blind spots, and that applies to both natural and artificial ones. It may be that to approximate neutrality we or our machines have to leave behind the form of intelligence that depends on intrinsically biased weights and instead depend on logically deriving all values from first principles. I have low confidence that AI's can accomplish that any time soon, and zero confidence that natural intelligence can. And it's difficult to see how first principles regarding human values can be neutral.
I'm also skeptical that succeeding at becoming unbiased is a solution, and that while neutrality may be an epistemic advance, it also degrades social cohesion, and that neutrality looks like rationality, but bias may be Chesterson's Fence and we should be very careful about tearing it down. Maybe it's a blessing that we can't.
Is it ideological?
Is it product marketing in those relevant threads where someone is showcasing?
Or is it pure technical testing, playing around?
Incidentally, how much do they pay for a HN account that is a few years old and accumulated a few thousand Internet points?
Asking for a friend.
So far it hasn't happed here, but we'll see!
Other accounts might be trying to age accounts and dilute their eventual coordinated voting or commenting rings. It's harder to identify sockpuppet accounts when they've been dutifully commenting slop for months before they start astroturfing for the chosen topic.
To reverse the argument - it would be amateurish and plain stupid to ignore it. Barrier to entry is very low. Politics, ads, swaying mildly opinions of some recent clusterfuck by popular megacorp XYZ, just spying on people, you have it all here.
I dont know how dang and crew protects against this, I'd expect some level of success but 100% seems unrealistic. Slow and steady mild infiltration, either by AI bots or humans from GRU and similar orgs who have this literally in their job description.
Slashdot's system was superior because mod points were finite and randomly dispensed. This entropy discouraged abuse by design—as opposed to making it a key feature of the site.
It's the Achilles' heel of Reddit and every site that attempts to emulate it.
I've been advocating for a while now that HN could use meta-moderation at least on flagging activity, so it can stop giving flagging powers to users who are using it for reasons other than flagging rulebreaking.
No one's really described the main piece - the strategical motivation. I suggested politics/ideology, product marketing or just technical testing of sock puppet creation.
In reddit people spam referral links. And here?
They don't have anything worth saying but want people to think they do
My relationship with writing, while improved, has been a difficult one. Part of me has always felt that there was a gap in my writing education. The choices other writers seem to make intuitively - sentence structure, word choice, and expression of ideas - do not come naturally to me. It feels like everyone else received the instructions and I missed that lesson.
The result was a sense of unequal skill. Not because my ideas are any less deserving, but because my ability to articulate them doesn't do them justice. The conceit is that, "If I was able to write better, more people would agree with me." It's entirely based on ego and fear of rejection.
Eventually, I learned that no matter how polished my writing is, even restructured by LLMs, it won't give me what I craved. At that moment, the separation of writer and words widened to a point where it wasn't about me anymore and more about them, the readers. This distance made all the difference and now I write with my own voice however awkward that may be.
Because it looks completely adequate for me. Maybe you're not the bad writer you think you are.
Sometimes there is no clear explanation for fake account registration. Perhaps they were registered to be actively used in the future, as most fraud prevention techniques target new account registration and therefore old, aged accounts won't raise suspicion.
Slightly off-topic, but there are relatively new `services` that offer native brand mentions in reddit comments. Perhaps this will soon be available for HN as well, and warming up accounts might be needed for this purpose.
Oh, would you look at that?
I love how the bot forgot to read CLAUDE.md or whatever persona it set up (e.g., "make me text all lowercase, use -- instead of em dashes pleaseeee") for this single comment mixed in with the other ones:
https://news.ycombinator.com/item?id=47132431
Sadly, I think that bot comment without the 'snowhale' persona filter applied is what a lot of people here still think every bot is going to look and sound like, because the amount of people I've seen on here getting tricked by them and interacting with them has been a bit worrisome.
https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu...
https://news.ycombinator.com/item?id=45322362
> First impression: I need to dive into this hackernews reply mockup thing thoroughly without any fluff or self-promotion. My persona should be ..., energetic with health/tech insights but casual and relatable.
> Looking at the constraints: short, punchy between 50-80 characters total—probably multiple one-sentence paragraphs here to fit that brevity while keeping it engaging.
> User specified avoiding "Hey" or "absolutely."
Lots more in its other comments (you need [showdead] on).
It's not just clever—it's devious!
EDIT to correct: most are not [flagged], but [dead] anyway, so probably manual moderator action or an automated anti-bot measure.
That's why. Boring, bland, etc. That account's M.O. is basically "write a paragraph that says nothing." Fwiw, I do think AI can be indistinguishable from dumb, boring people, but usually those kinds of people won't be on HN.
I feel like I'm certainly in that club as well.
I agree it doesn't seem obviously AI. The early comments are all in the same writing style and smell human. Lots of strong opinions e.g.
"logged in after years away and had basically the same experience. the feed is just AI slop and engagement bait now, none of it from people I actually followed." [about Facebook]
HN has got a big problem with silently shadowbanning accounts for no obvious reason. Whether it's an attempt to fight bots gone wrong or something else isn't clear. By the very nature of shadowbanning there is no feedback loop that can correct mistake.
I don't think it's clear at all why people do this. I suspect a large amount of it, at least on a site like HN, is just hapless morons who think it's "cool".
For example, here's an active bot that posted 30 mins ago (as of this comment):
https://news.ycombinator.com/threads?id=aplomb1026
Examine the last two detailed comments it made and you'll see the timestamps show they were posted < 30 seconds apart:
https://news.ycombinator.com/item?id=47155655
https://news.ycombinator.com/item?id=47155648
If it wasn't for them misconfiguring their bot and having it post so quickly, these would go by undetected and most people would engage with them. The comments themselves seem "normal" at first glance.
---
Other bots:
The irony of a bot comment
If AI starts use the New Yorker style diaeresis (umlaut-looking thing when there are two vowels in words like coöperate) I swear I'm gonna lose it.
Join me in double-dash em proximates. Shows you manually typed it out with total disregard token count and technical correctness.
Is there any good argument in favor of it, or any other house style quirks for that matter, other than in-group signaling?
Non-native speakers might see something like "nave" instead of "nigh-eve" unless it is clear that there is a stress that breaks out of the diphthong.
I don't think style guides are (usually) about absolute correctness, but relative correctness. A question is asked, a decision needs making, someone makes it, and now a team of individuals can speak with a consistent voice because there's a guideline to minimize variation.
I was going to say that I respect it, but find it utterly absurd that they do that. But your comment made me look it up again—I had no idea it was just obsolete/archaïc (except in the New Yorker), I'd thought it was a language feature their 'style' guide had invented.
Fun fact: if you have the audacity to correctly write an SMS, you can fit about 70 characters in an SMS. It converts the whole message into multibyte instead of only adding dots to the one character. Or if you use classic spelling for naïve in English, same issue. (We don't dots-ize that in Dutch because ai is not a single sound like ee is, so there's no confusion possible. This is purely English.) I believe in Hanlon's razor so it's probably a coincidence that whoever cooked up this terrible encoding scheme made carriers a lot of money, but I do wonder if this had anything to do with the bug still existing to this day!
I present ⸻ the U+2E3B dash.
apparently used like ellipses … to indicate part of a quote was removed.
> ...primarily used in scholarly and literary contexts to indicate a significant omission within a quoted passage, typically representing a full line, paragraph, or even multiple sentences of text
No one wants to read your ChatGPT outputs.
...except ChatGPT fans.
> subtly detached or glossed-over metaphors
> feels empty
> alluding to things without actually talking about them
> using many words to say nothing
https://v-n-n-v.github.io/chatgpt-voice.html
See also: https://i.vgy.me/ZIxnSh.png
As someone who loves LaTeX, I can't imagine ever spending so much time on typography on online forums, italics, bold, emdashes, headers, sections. I quit reddit and will quit hn as well if situation worsens.
(1) I don't recommend focusing disproportionately on one signal. They'll change, and are incredibly easy to optimize for. https://en.wikipedia.org/wiki/Wikipedia:Signs_of_AI_writing
(2) I do recommend taking one minute to dash a note off to hn@ycombinator.com if you see suspicious patterns. Dang and our other intrepid mods are preturnatually responsive, and appear to appreciate the extra eyeballs on the problem.
This wasn't really a intended as an "wow, dang is sure sleeping on the job", more than an interesting observation on the new bot ecosystem.
I also feel like there's a missing discussion about the comment quality on HN lately. It feels like it's dropped like crazy. Wanted to see if I could find some hard data to show I haven't gone full Terry Davis.
Obvious AI-generated posts and articles make it to the front page on a daily basis, and I get the impression that neither the average user nor the moderation team see that as a problem at all anymore.
Additionally, lots of Chinese and Russian keyboard tools use the em dash as well, when they're switching to the alternative (en-US) layout overlay.
There's also the Chinese idiom symbol in UTF8 which gets used as a dot by those users a lot, so that could be a nice indicator for legit human users.
edit: lol @ downvotes. Must have hit a vulnerable spot, huh?
I'm not trying to negate the fact. I'm just pointing out that a correlation without another indicator is not evidence enough that someone is a bot user, especially in the golden age of rebranded DDoS botnets as residential proxy services that everyone seems to start using since ~Q4 2024.
That’s why the analysis was performed over time. All of those em dash sources you mentioned were present before LLM written content became popular.
Bot prevention is a very difficult constant game of cat and mouse, and a lot of bot operators have become very skilled at determining the hidden metrics used by platforms to bless accounts; that's their job, after all. I've become a big fan of lobste.rs' invitation tree approach, where the reputation of new accounts rides on the reputation of older accounts, and risks consequence up the chain. It also creates a very useful graph of account origin, allowing for scorched earth approaches to moderation that would otherwise require a serious (and often one-off) machine learning approach to connect accounts.
I know there are legitimate usecases for the em-dash, but a few paragraphs (at most) of text in an HN/Reddit comment? Into the trash it goes.
trying to remember last time I used it
To turn off Smart Punctuation: Home > Settings > General > Keyboard > Smart Punctuation > Off.
Honestly I've lost all respect for people who use LLM generated anything.
https://practicaltypography.com/hyphens-and-dashes.html
I will not allow my good practices to get co-opted as AI "smoke tests".
Is it possible to differentiate between a bot, and a human using AI to 'improve' the quality of their comment where some of the content might be AI written but not all? I don't think it is.
hm, the whole internet really, youtube, reddit, twitter, facebook, blog posts, food recipes, news articles, it's getting more and more obvious
lets bring back Chrome's WEI while we're at it
/s
And bots reposting a trending post from like 12 years ago to farm internet points... with other bots reposting the top comments of the initial post
I'm more worried about how many people reply to slop and start arguing with it (usually receiving no replies — the slop machine goes to the next thread instead) when they should be flagging and reporting it; this has changed in the last few months.
I'm never suspicious though. One of the strange, and awesome, and incredibly rare things about HN is that I put basically zero stock in who wrote a comment. It's such a minimal part of the UI that it entirely passes me by most of the time. I love that about this site. I don't think I'm particularly unusual in that either; when someone shared a link about the top commenters recently there were quite a few comments about how people don't notice or how they don't recognize the people in the top ranks.
The consequence of this is that a bot could merrily post on here and I'd be absolutely fine not knowing or caring if it was a bot or not. I can judge the content of what the bot is posting and upvote/downvote accordingly. That, in my opinion, is exactly how the internet should work - judge the content of the post, not the character of the poster. If someone posts things I find insightful, interesting, or funny I'll upvote them. It has exactly zero value apart from maybe a little dopamine for a human, and actually zero for a robot, but it makes me feel nice about myself that I showed appreciation.
Brevity is the soul of wit.
I want to hear people in their own voice, their own ideas, with their own words. I have no interest in reading AI generated comments with the same prose, vocabulary, and grammar.
I don't care if your writing is bad.
Additionally, I am sceptical that using AI to write comments on your behalf creates opportunities for self-improvement. I suspect this is all leading to a death of diversity in writing where comments increasingly have an aura of sameness.
-> to be fair there must also be a bias of young incoming ppl on HN being more prone to be starting their career on the hot new tech
Often lean slightly pro-AI, but otherwise avoid saying much about anything.
What we think others around us think has a big effect on our own behavior
Don’t mind me, just skewing the results. — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — results. — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — results. — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — results. — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — results. — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — results. — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — results. — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — — —
Could be an argument made for aggregating by user instead however, if some bots are found to be particularly active and skewing the data.
Shhh!
:)
Why not? I am a descendant of Africans. I am a mildly successful author by tech nerd standards. I was educated in the British Public School tradition, right down to taking Latin in high school and cheering on our Rugby* and Cricket teams.
If someone doesn't want to read my words or employ me because I must be AI, that's their problem. The truth is, they won't like what I have to say any more than they like the way I say it.
I have made my peace with this.
———
Speaking of Rugby, in 1973 another school's Rugby team played ours, and almost the entire school turned out to watch a celebrity on the other school's team.
His name was Andrew, and he is very much in the news today.
I also see AIs use emdashes in places where parentheses, colons, or sentence breaks are simply more appropriate.
––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––––
There is no real AI detection tool that works.
When we see something like emd-ashes its simply the average of the used text the models trained on. If you fall into one the averages of a model you basically part of the model ouput. Yikes.
What could help is a careful clique hunting algorithm to accurately identify and delete the entire clique.
It is also interesting to note that the comparison is between recent comments and recent comments by new users. So, I guess this would take care of the objection that em-dashes (a perfectly fine piece of punctuation) have just been popularized by bots, and now are used more often by humans as well.
Maybe there is a bot problem. Seems almost impossible to fix for a site like this…
Our company is being attacked rn in tech media and at least some of it, gut feeling wise, seems obviously sponsored / promoted by competitors. I know that's not surprising, but never watched it happen from this side before.
I enjoyed his use of them so much in his writing that I started using them in my own book that came out in 2017. I freely admit—without hesitation—that my own use of em-dashes is due to author Robert Caro's influence.
There is much amusement at the idea that tech-weenies today are freaking out that the appearance of em-dashes in text is a surefire tell that so-called "AI" generated said text.
Read some books, get away from the computer, eh?
I know this is unfair to prospective new community members, but I'm unsure of other good methods to filter out AI bots at scale. Would certainly welcome other ideas.
2. People with dyslexia and dysgraphia can more easily interact online
3. People who speak a primary language other than English can more easily interact online
The last 2 options mean people who previously would have been more reluctant to participate now have less of a barrier.
So while there may be AI generated content, we should just assume it is all negative.
I also don't think the first point is correct at all.
I'm also influenced by the email style of my colleagues, books I'm reading, X, etc.
My literary diet really does show in my writing, so I'll keep up reading the classics to balance out all the LLM content :)
A -3dB cutoff might be >= 01/01/2020, to pick a round figure.
Yet I never browse https://news.ycombinator.com/classic
Perhaps a classic comment filter might work…
again with the conspiracy theories
But who knows, maybe even 17 year old accounts are being hijacked by AI now too.
(Only half sarcastic)
So much so that I've started a Wikipedia project to replace the dashes and other abysmal characters with proper typography.
All of my devices replace hyphens with emdashes, ascii typography with glorious unicode, etc.
^ And that's what i just did
Bye bye em-dash, we had a nice run together.
I might start using that⸻one (a bit long...)
The incentives to use bots are many.
This use was one of the first things that occurred to me when LLMs started getting genuinely good at summarizing texts and conversations. And I assume a fair bit of this has always happened on HN too. I've never moderated here obviously so I have no first hand insight but the social conventions here are uniquely ripe for it and it has a disproportionate influence on society through the dominance of the tech industry, making it a good target.
Lots of dumb blogs from unknowns about vibecoding entire products in a weekend using specific AI slop generators from new startups that are getting to the front page lately.
remember: 1) those accounts that are causing 10x as many em-dashes are the dumb AI accounts. The smart ones are at the least filtering obvious tells from the output. They might even outnumber the dumb ones.
2) Also, a lot of people are real, but using AI to make themselves sound smarter. It's not necessarily completely nefarious.
Using em-dashes as an estimate has to result in a bunch of undercounting and overcounting.
AI use is similar. Ask it to do whatever writing or text wrangling you want, but please show the public the sanitized version.
Incidentally, some folks reported my stuff for potential AI generation and I had to respond to the mods about it. So that was kinda funny, if also sad to hear that some folks thought I was a bot.
I’m a dinosaur, not a robot dinosaur. I’m nowhere near that cool, alas.
The tell here is that you used a hyphen, not an em-dash.
This `-` is a hyphen, which I love, even if I'm fairly sure I'm not using it correctly in grammar a lot of the time.
This `--` is an EM-Dash, apparently, which is also what I never use but I also thought was just a hyphen in a different context (incorrect!).
And "--" is absolutely just two hyphen-minuses, not an em-dash (—).
1. We have the hyphen, which is most commonly used to create multi-part words, such as one-and-one-thousand.
2. We have the EN-DASH, which is most commonly used to denote spans of ranges. As an example, Barack Obama was President 2009–2017.
3. Then we have the recently maligned EM-DASH, which can be used in place of a variety of other punctuation marks, such as commas, colons, and parentheses. Very frequently, AI will use the em-dash as a way to separate two clauses and provide forward motion. AI uses it for the same reason that writers do: the em-dash is just a nicer punctuation mark compared to the colon.
4. Lastly, we have the minus sign, which is slightly different than the hyphen, though on most keyboards they're combined into the hyphen-minus.
By the by, they're called the em-dash and the en-dash because they match the length of an uppercase M or N, respectively.
Show HN: Hacker News em dash user leaderboard pre-ChatGPT - https://news.ycombinator.com/item?id=45071722 - Aug 2025 (266 comments)
... which I'm proud to say originated here: https://news.ycombinator.com/item?id=45046883.
I'm very disappointed to not have made the list—going to federal prison for 18 months didn't help my score.
Every time someone states they stop reading when they encounter proper typography, I feel attacked.
⁂even though I used to like pointing out the difference between a hyphen and a period.
Spaces like HN then become a cacophony of clankers clanking as their numbers increase
Good to know so I don't do it x10 more :D
easy, they are using llms to write their posts…
Not sure which is scarier
- Generate age so spamming a product/service is easier and the account appears more trustworthy
- Influence discussions in a particular direction for monetary gain, i.e. "I got rich on bitcoin, you'd be crazy not to invest".
- Influence discussions in a particular direction for political gain, i.e. "I went to Xinjiang and the Uyghurs couldn't be happier!"
I just hope my writing carries enough voice and perspective that people respond, even if there's an em dash or two.
Perhaps there needs to be some sort of voluntary ethical disclosure practice to disclaim text as AI-generated with some sort of unusual signifiers. „Lower double quotes perhaps?„
The irony is that in tech, almost everyone is using AI to improve their writing at this point. And often it does make things clearer and more concise. But we've created this weird social norm where the output needs to look like it wasn't touched by AI, even when everyone knows it was. So we all spend time manually roughing up perfectly good text to maintain the illusion of authenticity.
Maybe the em dash is the self censorship/deletion mechanism that we've all been waiting for. Better than having to write pill subscription ads, I suppose.
It's of course not a surprise that an LLM would be most proficient in language use and, adjacent to that, in proper formatting of said language. But it's a good thing and a good tool for writing, as anyone who has ever used a classic spell or grammar checker will attest to. But apparently we as a society have once again managed to completely overlook and demonise the good and now people who have paid attention in school have to bow to people who are somehow convinced that perfect spelling is a sign that someone cheated. This is not LLMs' fault, it's people's who think they've understood something when they really haven't, crying heresy over others doing things the correct way.
That being said: of course there are social and technological challenges with cheating, spam bots and sock puppets and what not, but the phenomenon itself is not really new, just the scale, cost and quality is way different now. We need to find a balanced way to approach it -- trying to weed out every last possible AI cheater while hurting real innocent people in the process is not worth it. Especially since we don't have a proper metric to actually prove who's a cheater and who is not, it's gotten way harder since the days of "As a large language model" being in every second sentence.
That was a bit saddening honestly. I kept the presentation as-is as I didn’t knew how to willfully screw up a Beamer presentation, and I would not touch PowerPoint (fortunately the final jury believed me).
Cheating had always been an issue before LLMs, but now we’re back to the same old tricks: just make sure to add a mistake or two to hide you copied the homework on your neighbor. It’s a shame because I kinda like learning the subtleties of foreign languages, and as a non-native English speaker, it’s quite rewarding when going online!
What will/can HN do about it?
If that's worth the cost... probably not?
For now maybe all forums should require some bloody swearing in each comment to at least prove you've got some damn human borne annoyance in you? It might even work against the big players for a little bit, because they have an incentive to have their LLMs not swearing. The monetary reward is after all in sounding professional.
Easy enough for any groups to overcome of course, but at least it'd be amusing for a while. Just watching the swear-farms getting set up in lower paid countries, mistakes being made by the large companies when using the "swearing enabled" models and all that.
But seriously, I loved the em-dash and now every time I use it (which is too often) I have to wonder if my words will immediately be written off.
Can we generate a huge amount of code, just compilable code, which is essentially just a trash. We seed the github, bitbucket, etc. and pollute the training grounds.
How many times have you had a coworker/boss/user ask for something totally nonsensical, impossible, or maybe even illegal? These people when they get unleashed on these agents and on large business problems are going to be a menace. They will create walls and heaps and mountains of working but totally stupid code as the AI attempts to work with their malformed attempts at a 'thought' and this will pollute github and the codebase for training...eventually it will outnumber serious, professionally written projects and you will reach an inflection point at which AI coding agents 'peak' and they begin to deteriorate.
Either that or you need to somehow ensure you are not training on AI written code but that too will cause a PR problem for the firms as it becomes obvious that AI cannot in fact replace developers?
Sometimes I'd like to correct misinformation about something niche that I happen to be over-knowledgable about, or drop a 1-5 line code fix for the very thing they're talking about.
But if the cost of me getting that commenting access is also lowering the threshold needed for anyone (or anything) to comment, I'd rather keep the threshold high.
A good chunk of users on lobste.rs are bloggers, so if they get something not quite right, I can just contact them through their blog anyway.
I hate myself for saying this, but HN should consider closing new registrations for a while until we figure out what to do with this.