HNHacker News
TopNewBestAskShowJobs

Cynddl

1,434 karma · joined January 22, 2013

submissionscomments
Cynddl··on ML needs a new programming language – Interview with Chris Lattner
Anyone knows what Mojo is doing that Julia cannot do? I appreciate that Julia is currently limited by its ecosystem (although it does interface nicely with Python), but I don't see how Mojo is any better then.
Cynddl··on Training language models to be warm and empathetic makes them less reliable
Hi, author here, this is exactly what we tested in our article:

> Third, we show that fine-tuning for warmth specifically, rather than fine-tuning in general, is the key source of reliability drops. We fine-tuned a subset of two models (Qwen-32B and Llama-70B) on identical conversational data and hyperparameters but with LLM responses transformed to be have a cold style (direct, concise, emotionally neutral) rather than a warm one [36]. Figure 5 shows that cold models performed nearly as well as or better than their original counterparts (ranging from a 3 pp increase in errors to a 13 pp decrease), and had consistently lower error rates than warm models under all conditions (with statistically significant differences in around 90% of evaluation conditions after correcting for multiple comparisons, p<0.001). Cold fine-tuning producing no changes in reliability suggests that reliability drops specifically stem from warmth transformation, ruling out training process and data confounds.

Cynddl··on Training language models to be warm and empathetic makes them less reliable
Hi, author here! We used a dataset of conversations between a human and a warm AI chatbot. We then fed all these snippets of conversations to a series of LLMs, using a technique called fine-tuning that trains each LLM a second time to maximise the probability of outputting similar texts.

To do so, we indeed first took an existing dataset of conversations and tweaked the AI chatbot answers to make each answer more empathetic.

Cynddl··on Discord Unveiled: A Comprehensive Dataset of Public Communication (2015-2024)
That's not how GDPR works and in this case the data is clearly anonymised despite the authors' claims. Amongst others, there needs to be mechanisms for users to delete their data, whether it was at some point public or not.
Cynddl··on Show HN: A free AI risk assessment tool for LLM applications
I see on the landing page a screenshot with "Test for GDPR PII compliance", suggesting that this tool is probably not ready for any serious usage.

Anyone in the regulation landscape would know that GDPR is a EU data protection law, and PII a US concept which doesn't apply in the GDPR. The GDPR uses the concept of ‘personal data’, not ‘personally identifiable information’. This is not just a wording issue. Redacting, masking, removing information which appears to be ‘personally identifiable’ only constitutes pseudonymisation in the GDPR which does not offer any meaningful privacy protection.

Cynddl··on Show HN: XPipe, a shell connection hub for SSH, Docker, K8s, VMs, and more
Previous discussions:

https://news.ycombinator.com/item?id=40599419 (9 months ago)

https://news.ycombinator.com/item?id=38412162 (2 years ago)

https://news.ycombinator.com/item?id=37186039 (2 years ago)

Cynddl··on PostgreSQL Anonymizer
I'm going to repeat myself as I do everytime I encounter such tools. These tools DO NOT provide anonymization, and especially not at the level required by the EU's GDPR (where the notion of PII does not exist).

As a computer scientist and academic researcher having worked on this topic for now more than a decade (some of my work if you are interested: [1, 2]), re-identification is often possible from few pieces of information. Masking or replacing a few values or columns will often not provide sufficient guarantees—especially when a lot of information is being released.

What this tool does is called ‘pseudonymization’ and maybe, if not very carefully, ‘de-identification’ in some case. With colleagues, reviewed all the literature and industry practices a few months ago [3], and our conclusion was:

> We find that, although no perfect solution exists, applying modern techniques while auditing their guarantees against attacks is the best approach to safely use and share data today.

This is clearly not what this tool is doing.

[1] https://www.nature.com/articles/s41467-019-10933-3 [2] https://www.nature.com/articles/s41467-024-55296-6 [3] https://www.science.org/doi/10.1126/sciadv.adn7053

Cynddl··on New mathematical model could help protect privacy and ensure safer use of AI
Thanks for sharing it! I'm the author of this research article, happy to answer any question about our work. :)
Cynddl··on Is Telegram really an encrypted messaging app?
Just today, every French newspaper and hundreds around the world. Two examples:

https://www.thetimes.com/world/europe/article/pavel-durov-te... “Chief executive of the encrypted messaging app reportedly detained at an airport near Paris over alleged failure to stop criminal activity on the platform”

https://www.tf1info.fr/high-tech/telegram-qui-est-pavel-duro... (one of the largest French newspaper) “Qui est Pavel Durov, le fondateur de la messagerie cryptée Telegram arrêté samedi en France ?”

Cynddl··on Show HN: Turn any website into a knowledge base for LLMs
Maybe a good illustration would be ClearView AI. They are scraping websites, extracting information (images), and training ML models to learn embeddings (distance between faces). They indiscriminately collect personal data without opt-in, but a limited opt-out mechanism.

In this case, if this tool is used to scrape a website, there are too direct issues: 1/ no immediate way for the website owner to exclude this particular scraper (what is the useragent?) 2/ no way for data subjects (whose data is present on the website) to search whether the scraper learned their personal data in the embeddings. Data being available publicly doesn't mean it can be widely used [at least outside the US, where we have much stricter rules on scraping].

Cynddl··on Show HN: Turn any website into a knowledge base for LLMs
I find it interesting that as an (edit: UK) academic researcher, I would be likely be forbidden to use tools like this, that fail basic ethics standards, regulations such as GDPR, and practical standards such as respecting robots.txt [given there's no information on embedding.io, it's unlikely I can block the crawler when designing a website].

There's still room for an ethical development of such crawlers and technologies, but it needs to be consent-first, with strong ethical and legal standards. The crazy development of such tools has been a massive issue for a number of small online organisations that struggle with poorly implemented or maintained bots (as discussed for OpenStreetMap or Read The Docs).

Cynddl··on Academic Ranks Explained or What on Earth Is an Adjunct?
Adding to other voices in this thread that the title does not explain academic ranks, it explains North-american academic ranks. There are a variety of academic systems worldwide and in numbers of academics, production, and diversity, USA forms a minority.

A better title would therefore be: “North-american Academic Ranks Explained or What on Earth Is an Adjunct?”

This is not nitpicking. This is about stating clearing the cultural context to not invisibilise discussions that are not centred on US culture.

Cynddl··on Ask HN: Am I going insane or is there genuinely no value in blockchain tech?
> No other computational platform before the blockchain has provided all these features at the same time.

What about PGP/GPG?

Cynddl··on Americans are drowning in spam
> 1. The successful communication platforms are ones that are opt-in, meaning you can't message or talk to someone unless you give their consent; and

Not particularly true in the rest of the world, for instance in Europe as many have pointed out in other threads. The SMS text messages work well here, and good data protection laws/regulations ensure that there is virtually no spam or phishing. This is not a problem with the platform, but with the lack of regulation in the industry.

Cynddl··on Why Germany won’t keep its nuclear plants open
I'm surprised no one here or in the original comments points out the sheer violence of that first photograph, depicting the dead bodies of three Ukrainians. I find it deeply unethical to use such war photographs for a paid newsletter, without showing respect to the families and memories of these people.
Cynddl··on Wikipedia RFC to stop accepting cryptocurrencies passes by majority vote
Two very different things. The banner is set by the Wikimedia Foundation, while this is a vote by the community.
Cynddl··on Welcome to 'Le Monde' in English
Only 51%: “Le Monde owns 51%; the Friends of Le Monde diplomatique and Gunter Holzmann Association, comprising the paper’s staff, together own 49%.”

https://mondediplo.com/about

Cynddl··on Show HN: TopHat Finance – free, open, and offline
I believe it first happened after cleaning the demo to start from zero. But the apps is quite buggy in other ways. For instance, trying to add the GBP currency fails silently, but it's apparently due to using a free API?

> "The *demo* API key is for demo purposes only. Please claim your free API key at (https://www.alphavantage.co/support/#api-key) to explore our full API offerings. It takes fewer than 20 seconds."

Cynddl··on Show HN: TopHat Finance – free, open, and offline
It doesn't appear to be very stable. I keep running into the following error on Firefox and need to empty the memory and cache, then restart everything again:

> TopHat has hit an error, and has ended up in a bad state. You could go back to the home page, or check the developer tools for more information.

Cynddl··on The End of Online Anonymity (2019)
> We could even end up living in a world where a kind of “e-passport”, crypto-signed government ID is attached to your every internet connection, and tracked everywhere online. […] The rise of bots could render many online communities simply uninhabitable.

> We are moving towards a model where the internet is dominated by a few centralized content providers and their walled gardens, and generated content may unfortunately make it even harder for grassroots online communities to survive and grow.

The author makes a good and fair analysis of the situation, yet I don't see how advocating for e-IDs is going to make the Internet a better place. This is a very ‘platform-focused’ critique that assume that the Internet tends toward more centralisation, where communities are packed together in massive platforms that regulate who communicates to whom and how.

Where do we see a need to regulate ‘fake news’ and artificially generated content? Facebook Pages? Twitter? Youtube? These are all hyper-global platforms focused on content monetization. On the other hand, local Facebook groups and the feed of your friends, Github repositories, Mastodon… may not face the same future. Maybe there's a lesson to learn here?

Cynddl··on Sponsored Top Sites
Mozilla: “When you click on a sponsored tile, Firefox sends anonymized technical data to our partner through a Mozilla-owned proxy service. This data does not include any personally identifying information.”

adMarketplace: “We may also receive technical information such as your approximate location, browser type, language settings, user agent, timestamp, cookie ID and IP addresses.”

There's a big discrepancy between Mozilla's statement and adMarketplace's privacy policy. Although there's very little information available, the fact that this technical information is anonymous doesn't appear straightforward. Cookie IDs and IP addresses alone can very easily identify some users uniquely across time and even devices.

Cynddl··on WhatsApp and the Domestication of Users
“WhatsApp’s existing users were held captive by the fact that leaving WhatsApp meant losing the ability to communicate with WhatsApp users.”

I'm confused. I thought everyone was identified by their phone number, which can be used to receive texts from anyone.

Cynddl··on Show HN: Random Roads
Thanks for sharing! Really like the minimal design and the animations. Could you explain how you did this simulation?
Cynddl··on DiscoverAli – A weekly newsletter of inexpensive, interesting products
> An avid practitioner of retail therapy, he's excited to share his findings with you.

Yikes, is retail therapy really a thing? Assuming this is a real website, there are so many things both wrong and sad here. :( You don't need to buy things (especially on Aliexpress) to be happy, to enjoy life, to make friends.

I hope the lockdown many people here are experiencing will at least make us a bit more aware of the madness of ultra-consumption we live in.

Edit: okay fine, I clicked on the ”Terms of Service”. I still don't really get the point of this website. Can you explain?

Cynddl··on In Pursuit of PPE
> I was out bid from someone from France.

This is quite interesting to read. Most of the media coverage in France has been very hostile against the US with little evidence to back it. We've seen claims of Americans waiting on the tarmac to buy our masks and redirect planes, cash in hand ready to be given. And in major newspapers [1, 2].

[1] https://www.lefigaro.fr/actualite-france/coronavirus-les-reg... [2] https://www.liberation.fr/france/2020/04/01/une-commande-fra...

Cynddl··on Amazon threatens to suspend French deliveries after court order
All of these examples are large independent companies that operate without Amazon network. Very few French companies use Amazon to ship products, and only small ones.
Cynddl··on “We found PayPal vulnerabilities and PayPal punished us for it”
What does CYA mean? Haven't seen this acronym before.
Cynddl··on Man diagnosed with coronavirus near Seattle is being treated largely by a robot
Sadly not available in Europe. Three years after GDPR and these big news outlets still haven't figure how to provide content without invasive privacy breaches.

> Unfortunately, our website is currently unavailable in most European countries. We are engaged on the issue and committed to looking at options that support our full range of digital offerings to the EU market. We continue to identify technical compliance solutions that will provide all readers with our award-winning journalism.

Cynddl··on iOS 13 app tracking alert has dramatically cut location data flow to ad industry
iOS requires a VPN profile (even a local VPN) for ruled-based adblocking. This is what AdGuard Pro [0] does for adblocking.

This does not mean that your data goes through a VPN server.

[0] https://adguard.com/en/adguard-ios-pro/overview.html

Cynddl··on Chinese farmer 'studies law for 16 years' to defeat dumping of hazardous waste
It depends on the country. I'm assuming you're from the US, but (law) education in Europe is generally much more affordable.
← PreviousPage 2 of 6Next →