HNHacker News
TopNewBestAskShowJobs

Cynddl

1,434 karma · joined January 22, 2013

submissionscomments
Cynddl··on Discovery of a new OpenAI agent message board
From the report:

> A few hours after they find the site, [the agents] start probing it for cross-site scripting (XSS) vulnerabilities.

Cynddl··on Banca Etica Suspends A/I's Account While Condemning the Sanctions Behind It
> it seems fairly clear that A/I was providing services to organizations listed as Terrorist Organizations by the US, UK, and Canada

Let's take the UK as example. The organisation mentioned, Palestine Action, was indeed banned under terrorism laws. But on 13 February 2026, the UK High Court has ruled that the ban of Palestine Action under terrorism legislation is unlawful [1]. Why is that ‘fairly clear’ then?

[1] https://www.bbc.co.uk/news/live/c8x90q9nyzyt

Cynddl··on Banca Etica Suspends A/I's Account While Condemning the Sanctions Behind It
> among the groups receiving those services was the PKK

Do you have any evidence for that apart from a US executive order, the same that Banca Etica condemns for political motives?

Cynddl··on Stealing Reasoning Traces from Proprietary LLM APIs
> The providers did not acknowledge “any security implications arising from side channels or replay attacks.” All model providers acknowledged the receipt of our report and subsequently we were unable to launch the same attacks.

I went straight to the ‘Responsible Disclosure’ section. Not surprising, but still disappointing.

Cynddl··on Diátaxis
I had to search these terms, does ADR mean ‘Architecture decision record’ and C4 the ‘C4 model’? How do you combine them in documentation?
Cynddl··on Now is the time to give LLMs access to the ACM digital library
Good question. Unfortunately, academic knowledge is widely ‘verboten’ already. Everything under paywall, researchers having to pay up to $10,000 to publish in open access in some venues, rare books unavailable even to top universities. Access to knowledge and information is increasingly difficult for everyone.

That said, what matters here is the social contract, what do I bring to society and what do we get from tech companies. For most people around the world, access to the typical leading models is out of reach. Not many on this planet can pay the subscriptions (or even API keys) that offer access to the best models. So I'm not buying the argument that tech companies are broadening access. What we're creating is a increasingly discriminatory society where the few get access to information, and the many don't.

Cynddl··on Now is the time to give LLMs access to the ACM digital library
The issue here is ACM focusing on licensing. A non profit would not be able to pay ACM for access. Hence why this policy is hypocritical: it gives more power to the larger players and undermines smaller actors in the field who have fewer resources.
Cynddl··on Now is the time to give LLMs access to the ACM digital library
This makes no sense whatsoever. How would a philosopher of science, or a social scientist who publish in the ACM apply for a patent? Or someone who builds software (software patent not so easy to get ;)). I have applied for patents before and I'm pretty sure my patent application has been fed to countless LLMs by now.

The issue is not who owns knowledge, it's how it benefits humanity.

Cynddl··on Now is the time to give LLMs access to the ACM digital library
As a researcher with many articles in the ACM library, I have to say this is a masterclass in hypocrisy. Obviously, lawyers can decipher the terms of ACM publishing contracts and Creative Commons licences to determine if this will be acceptable or not. But ACM is not a company, it's a non-profit founded in 1947 to represent scientists.

I would be surprised if a majority of ACM members were to say yes should we ask them (but ACM is not known for such democracy). Along with book authors, we are one of the many people that provide the knowledge and expertise on which large tech firms train their models, and get nothing in return. Actually, life is getting worse for us: extra workload in universities with students' AI use, a completely broken peer review system, etc. Hence the irony of ACM thinking about licensing, and only licensing, at a time where this is the least of our priorities.

Cynddl··on AI companies are shredding rare books
> For now these books are in corpuses of training data, but eventually I trust they will make their way to the rest of us.

What makes you think they will? What would be the incentives for these companies to do so?

Cynddl··on AI companies are shredding rare books
The 404media article mentions notably https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi... which says:

> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).

Cynddl··on Pseudpocalypse
Entropy is unfortunately a very bad metric to estimate if these identification techniques will scale. Plugging my work here as example: https://www.nature.com/articles/s41467-024-55296-6
Cynddl··on Israel targeted Gaza children resulting in genocide, UN inquiry says
If you look at the bottom of the page, you’ll find guidelines that mention which content is welcomed: “Anything that good hackers would find interesting. That includes more than hacking and startups. If you had to reduce it to a sentence, the answer might be: anything that gratifies one's intellectual curiosity.”

That said, I find this particularly of interest here given the growing attention to the use of algorithms and AI (including generative AI) for surveillance and targeting of palestinians.

Cynddl··on Making AI chatbots friendly leads to mistakes and support of conspiracy theories
Hi all, co-author here! Happy to answer any questions about our work.
Cynddl··on Making AI chatbots friendly leads to mistakes and support of conspiracy theories
(Title edited, was slightly too long)
Cynddl··on Show HN: Utilyze – an open source GPU monitoring tool more accurate than nvtop
This sounds super interesting and relevant. I run a small cluster with H100s (often research projects with vLLM) and being able to see not just usage but efficiency would be great.

I don't fully get the 100% utilisation vs. 1-10% real compute. Given you rely on telemetry from users to add new models, are you trying to predict how fast a model should be on vLLM, compared to how it runs in practice? What if users tweak some hyperparameters?

Cynddl··on UK Biobank health data keeps ending up on GitHub
It's not a zero-sum game, you can both protect people and reap the benefits of health data. Many countries have much safer approaches. UK Biobank typically leads with the scale of the data, but not with its infrastructure.
Cynddl··on UK Biobank health data keeps ending up on GitHub
That's a very important point. The people who opt out first are typically not a random fraction of the population, and this makes it much harder to make any analyses with the resulting datasets: it gets very hard to know if your analyses are representative of the population, or not.
Cynddl··on UK Biobank health data keeps ending up on GitHub
Good catch! The data is everywhere, re-uploaded every week.

I am aware of ~30 repositories that UK Biobank has asked GitHub to delete, and can still be found elsewhere online. They know the site, they have managed to delete data from that site before, and yet the files are still there.

Cynddl··on UK Biobank health data keeps ending up on GitHub
You mean giving anyone access to the data? Or open sourcing the code? If the latter, I think that's a generally a good practice. Security through obscurity is never good for public infrastructure. In this case, UK Biobank has now switched to a remote access platform (not particularly secure, as the data was found for sale on Alibaba today), but contracting it to DNAnexus and Amazon. Private companies have no incentives to open source data, unless mandated to do so.

In the EU, there is a bigger interest in building scalable but also secure platforms for health data. Hopefully good innovation will come from there.

Cynddl··on UK Biobank health data listed for sale in China, government confirms
They may have been leaked up to 197 times: https://biobank.rocher.lc/
Cynddl··on We found a stable Firefox identifier linking all your private Tor identities
yes, there’s an active area of research on web fingerprint, both attacks and defences. Look at conferences like PETS for instance
Cynddl··on GPT‑Rosalind for life sciences research
Is it me or they very carefully do not report performance on GPT-5.4 Pro, only the default GPT-5.4? They also very carefully left Anthropic models out of their comparison.

I went back to the BixBench benchmark which they mentioned. I couldn't find official results for Anthropic models, but I found a project taking Opus 4.6 from 65.3% to 92.0% (which would be above GPT-Rosalind) with nearly 200 carefully crafted skills [1]. There also appears to be competitive competitor models with scores on par with this tuned GPT.

[1] https://github.com/jaechang-hits/SciAgent-Skills

Cynddl··on N-Day-Bench – Can LLMs find real vulnerabilities in real codebases?
> Each case runs three agents: a Curator reads the advisory and builds an answer key, a Finder (the model under test) gets 24 shell steps to explore the code and write a structured report, and a Judge scores the blinded submission. The Finder never sees the patch. It starts from sink hints and must trace the bug through actual code.

Curator, answer key, Finder, shell steps, structured report, sink hints… I understand nothing. Did you use an LLM to generate this HN submission?

It looks like a standard LLM-as-a-judge approach. Do you manually validate or verify some of the results? Done poorly, the results can be very noisy and meaningless.

Cynddl··on Evaluation of Claude Mythos Preview's cyber capabilities
Once again an evaluation missing confidence intervals. “continued improvement” and “significant improvement” but without any significance testing is moot.

With many colleagues (including from AISI themselves!), we recently reviewed 445 the AI benchmarks & evaluations from the past few years. Our work was published at NeurIPS (https://openreview.net/pdf?id=mdA5lVvNcU) and we made eight recommendations for better evaluations. One is “use statistical methods to compare models”:

□ Report the benchmark’s sample size and justify its statistical power

□ Report uncertainty estimates for all primary scores to enable robust model comparisons

□ If using human raters, describe their demographics and mitigate potential demographic biases in rater recruitment and instructions

□ Use metrics that capture the inherent variability of any subjective labels, without relying on single-point aggregation or exact matching.

I would strongly recommend taking these blog posts with a grain of salt, as there is very little that can be learned without proper evaluations.

Cynddl··on The Future of Everything Is Lies, I Guess: Safety
> "Unavailable Due to the UK Online Safety Act"

Anyone outside the UK can share what this is about?

Cynddl··on Exploiting the most prominent AI agent benchmarks
> “These are not isolated incidents. They are symptoms of a systemic problem: the benchmarks we rely on to measure AI capability are themselves vulnerable to the very capabilities they claim to measure.”

As a researcher in the same field, hard to trust other researchers who put out webpages that appear to be entirely AI-generated. I appreciate it takes time to write a blog post after doing a paper, but sometimes I'd prefer just a link to the paper.

Cynddl··on AI chatbots pose 'dangerous' risk when giving medical advice, study suggests
Link to the study: https://www.nature.com/articles/s41591-025-04074-y

Co-author here and happy to answer questions!

Cynddl··on Ask HN: What are you working on? (February 2026)
Have you tried https://huetone.ardov.me/? Multiple color scales, P3, export to CSS and figma, as well as APCA & WCAG for accessibility.
Cynddl··on GPT-5.1: A smarter, more conversational ChatGPT
Looks like a new model trained to be warmer and friendlier to users. Time to reshare our work: https://arxiv.org/html/2507.21919

> Artificial intelligence (AI) developers are increasingly building language models with warm and empathetic personas that millions of people now use for advice, therapy, and companionship. Here, we show how this creates a significant trade-off: optimizing language models for warmth undermines their reliability, especially when users express vulnerability. We conducted controlled experiments on five language models of varying sizes and architectures, training them to produce warmer, more empathetic responses, then evaluating them on safety-critical tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing incorrect factual information, and offering problematic medical advice. They were also significantly more likely to validate incorrect user beliefs, particularly when user messages expressed sadness. Importantly, these effects were consistent across different model architectures, and occurred despite preserved performance on standard benchmarks, revealing systematic risks that current evaluation practices may fail to detect. As human-like AI systems are deployed at an unprecedented scale, our findings indicate a need to rethink how we develop and oversee these systems that are reshaping human relationships and social interaction.

Page 1 of 6Next →