HNHacker News
TopNewBestAskShowJobs

DalasNoin

791 karma · joined September 13, 2016

I am Simon, an AI security researcher

https://www.linkedin.com/in/simon-l772743a7/

submissionscomments
DalasNoin··on Frontier Labs Are Selling Garbage to Fools in Washington
what you read (past tense) is already the corection
DalasNoin··on Frontier Labs Are Selling Garbage to Fools in Washington
thank you for this reasonable reaction
DalasNoin··on Frontier Labs Are Selling Garbage to Fools in Washington
"Every single one of these catastrophic breakouts happened inside the testing environments of the exact same vendor."

This is incorrect, the HF incident for example (the most well known) had nothing to do with irregular. I know there has been a news site pushing inaccurate articles (effort.news) on this topic but these are the facts.

https://openai.com/index/hugging-face-incident-and-the-road-...

DalasNoin··on AI-generated posters don’t have to be horrible
"Me: a poster paint drawing by his young child."

Some time ago, I was in this apartment complex and there was a hand drawn sign of a dog pooping on the grass - the sign then said that residents should pick up after their dog. It looked like the locals living in the apartment complex had their child draw the dog pooping in a slightly adorable style. I remember feeling kind of happy about this, like I could see a kid sitting there with some crayons painting this dog. I know there were human artists perfecting this child like style before but still feels like some humanity is getting lost here making AI paint like a kid.

DalasNoin··on The Last 24 Hours Are the Opening Scene in a Horror Movie
Huggingface was attacked by models that finished training earlier this year, perhaps May. Current models are already substantially stronger. the next incident could be happening now. There is certainly no clear reason why models shouldn't soon be capable of self-exfiltration.
DalasNoin··on Goodbye Visa and Mastercard: 130M Europeans switching to sovereign payment
Can you confirm people in france actually use wero? I had heard of it every so often but basically zero people actually use it, my revolut app has a feature to use wero but never used it. I mean would be great, getting rid of CC fees could literally lower grocery prices by 1-2%.
DalasNoin··on Why AI companies want you to be afraid of them
Quote from the article: ""AI will probably most likely lead to the end of the world, but in the meantime, there'll be great companies," Altman said in 2015."

Altman wasn't even at OpenAI at that point, so why would that be marketing?

DalasNoin··on SpaceX says it has agreement to acquire Cursor for $60B
I mean the best argument I see for cursor is that you can easily switch between AIs, which is convenient since they seem to run at 80-90% up time (with those 10-20% clustered at West coast working hours). But the big AI companies are likely to keep an edge over Open-source fine-tunes and they are able to subsidize the coding agents in a way Cursor can't.
DalasNoin··on Nano Banana 2: Google's latest AI image generation model
Why does SynthID make it worthless? it helps other platforms detect this as ai?
DalasNoin··on Large-Scale Online Deanonymization with LLMs
We use semantic information inferred from comments and submissions. I think using stylometry would be a great addition, but it would be hard to google for "guy who writes fanciful using many puns" rather then "indie developer in Switzerland". I think stylometry could be better used for verification, once you have a small set of candidates stylometry could further narrow down the candidates and be used to make a decision.
DalasNoin··on Large-Scale Online Deanonymization with LLMs
We test different methods, in section 2, we use LLM agents to agentically identify people. We don't share any code here, but you could try with various freely available agents on yourself.
DalasNoin··on Large-Scale Online Deanonymization with LLMs
We essentially don't use stylometry but semantic information revealed from peoples' comments – clues and interests.

(We use a little stylometry in a single experiment in section 5)

DalasNoin··on Large-Scale Online Deanonymization with LLMs
We essentially don't use stylometry but semantic information – clues and interests.
DalasNoin··on Large-Scale Online Deanonymization with LLMs
That's a great background paper on the Netflix attack, we make a pretty direct comparison in section 5. We also try to use similar methods for comparison in sections 4 and 6. In section 5 we transform peoples Reddit comments into movie reviews with an LLM and then see if LLMs are better than naraynan purely on movie reviews. LLMs are still much better (getting about 8% but the average person only had 2.5 movies and 48% only shared one movie, so very difficult to match)
DalasNoin··on Large-Scale Online Deanonymization with LLMs
I agree that these accounts probably on average still contain more information than the average pseudonymous account. I think we could try to use the LLM to increasingly ablate more information and see how it performance decays – to be clear we already heavily remove such information, see Table 2 appendix. But I don't expect that to change the basic conclusions.
DalasNoin··on Large-Scale Online Deanonymization with LLMs
We do advocate for stricter controls on data access on social platforms because of this. There is a bit of an unfortunate trade-off, but I think allowing mass-scraping or downloads of data from social sites can be misused in increasingly more ways.
DalasNoin··on Large-Scale Online Deanonymization with LLMs
Mitigations are pretty difficult, I understand it is kind of cool that some websites have really open APIs where you can just read everything. There are some cool apps that used HN data in the past. But I think there should at least be consideration that LLMs are then going to read everything and potentially discover things. Users might have thought this is protected by obscurity, who would read their 5 year old comments?
DalasNoin··on Large-Scale Online Deanonymization with LLMs
There is also a practical issue here that people usually don't write a lot on linkedin, most people just have structured biographical information. We use very limited stylometry in section 6 for matching reddit users who we synthetically split according to time.
DalasNoin··on Large-Scale Online Deanonymization with LLMs
We don't use (much) stylometry, so this won't help. This is totally something you could try, but we use interests and clues. Semantic information you reveal about yourself.

The blog post might be more approachable if you want to get a quick take: https://simonlermen.substack.com/p/large-scale-online-deanon...

DalasNoin··on Large-Scale Online Deanonymization with LLMs
To be clear, we are making a clear concession here that the people weren't truly anonymous. But we did use an LLM to remove any identifying information from HN making them quasi-anonymous, this is more described in the appendix Table 2.

We do also make a more real world like test in section 2. There we use the anthropic interviewer dataset which Anthropic redacted, from the redacted interviews our agent identified 9/125 people based on clues.

The blog post might be more approachable for a quick take: https://simonlermen.substack.com/p/large-scale-online-deanon...

DalasNoin··on I'm helping my dog vibe code games
There goes all the prompt engineering jobs
DalasNoin··on Kimi Claw
From what I understand this is a fully open-source bot that anyone can run with no restrictions. what a time to be alive, let's see what these bots will break first
DalasNoin··on Data centers in space makes no sense
The energy collected from the solar panels must be converted into heat in the AI chips. It's really like putting the AI chips directly into the sun, just with extra steps. Sunlight gets transformed into electricity which gets transformed into heat in the chips.
DalasNoin··on Wisconsin communities signed secrecy deals for billion-dollar data centers
It's cold in a sense that is not very relevant. Your tumbler has a vaccuum layer because vaccuum does not transport or absorb any heat. you need those atoms to carry away heat.
DalasNoin··on Wisconsin communities signed secrecy deals for billion-dollar data centers
Space doesn't seem like a good place to build datacenters at all. Cooling is going to be an enormous issue, how do you disperse of heat in a vaccuum? Radiators are very ineffective for cooling.
DalasNoin··on Tesla ending Models S and X production
These robots give of a kind of dark vibe to me, especially with everything going on in his AI company. How long until one of them kills someone? I'd prefer a home robot that can't kill me (like something that is passively safe).
DalasNoin··on France targets Australia-style social media ban for children next year
Many of these "social" media websites increasingly just fling AI-generated disturbing videos at people. I am sure we could build a web that is actually pleasant to use for kids, but we are not building it. youtube for example: https://x.com/kimmonismus/status/2006013682472669589
DalasNoin··on Measuring the impact of AI scams on the elderly
(author here) we hope to have results for ai voice in months
DalasNoin··on Measuring the impact of AI scams on the elderly
You agree that AI is used for amplification by attackers? they interviewed people who had worked at scam factories who clarified that they used LLMs in their work.
DalasNoin··on Measuring the impact of AI scams on the elderly
They did have some idea that they would receive emails, probably slightly biasing the results downward.
Page 1 of 6Next →