1,434 karma · joined January 22, 2013
> Third, we show that fine-tuning for warmth specifically, rather than fine-tuning in general, is the key source of reliability drops. We fine-tuned a subset of two models (Qwen-32B and Llama-70B) on identical conversational data and hyperparameters but with LLM responses transformed to be have a cold style (direct, concise, emotionally neutral) rather than a warm one [36]. Figure 5 shows that cold models performed nearly as well as or better than their original counterparts (ranging from a 3 pp increase in errors to a 13 pp decrease), and had consistently lower error rates than warm models under all conditions (with statistically significant differences in around 90% of evaluation conditions after correcting for multiple comparisons, p<0.001). Cold fine-tuning producing no changes in reliability suggests that reliability drops specifically stem from warmth transformation, ruling out training process and data confounds.
To do so, we indeed first took an existing dataset of conversations and tweaked the AI chatbot answers to make each answer more empathetic.
Anyone in the regulation landscape would know that GDPR is a EU data protection law, and PII a US concept which doesn't apply in the GDPR. The GDPR uses the concept of ‘personal data’, not ‘personally identifiable information’. This is not just a wording issue. Redacting, masking, removing information which appears to be ‘personally identifiable’ only constitutes pseudonymisation in the GDPR which does not offer any meaningful privacy protection.
https://news.ycombinator.com/item?id=40599419 (9 months ago)
https://news.ycombinator.com/item?id=38412162 (2 years ago)
https://news.ycombinator.com/item?id=37186039 (2 years ago)
As a computer scientist and academic researcher having worked on this topic for now more than a decade (some of my work if you are interested: [1, 2]), re-identification is often possible from few pieces of information. Masking or replacing a few values or columns will often not provide sufficient guarantees—especially when a lot of information is being released.
What this tool does is called ‘pseudonymization’ and maybe, if not very carefully, ‘de-identification’ in some case. With colleagues, reviewed all the literature and industry practices a few months ago [3], and our conclusion was:
> We find that, although no perfect solution exists, applying modern techniques while auditing their guarantees against attacks is the best approach to safely use and share data today.
This is clearly not what this tool is doing.
[1] https://www.nature.com/articles/s41467-019-10933-3 [2] https://www.nature.com/articles/s41467-024-55296-6 [3] https://www.science.org/doi/10.1126/sciadv.adn7053
https://www.thetimes.com/world/europe/article/pavel-durov-te... “Chief executive of the encrypted messaging app reportedly detained at an airport near Paris over alleged failure to stop criminal activity on the platform”
https://www.tf1info.fr/high-tech/telegram-qui-est-pavel-duro... (one of the largest French newspaper) “Qui est Pavel Durov, le fondateur de la messagerie cryptée Telegram arrêté samedi en France ?”
In this case, if this tool is used to scrape a website, there are too direct issues: 1/ no immediate way for the website owner to exclude this particular scraper (what is the useragent?) 2/ no way for data subjects (whose data is present on the website) to search whether the scraper learned their personal data in the embeddings. Data being available publicly doesn't mean it can be widely used [at least outside the US, where we have much stricter rules on scraping].
There's still room for an ethical development of such crawlers and technologies, but it needs to be consent-first, with strong ethical and legal standards. The crazy development of such tools has been a massive issue for a number of small online organisations that struggle with poorly implemented or maintained bots (as discussed for OpenStreetMap or Read The Docs).
A better title would therefore be: “North-american Academic Ranks Explained or What on Earth Is an Adjunct?”
This is not nitpicking. This is about stating clearing the cultural context to not invisibilise discussions that are not centred on US culture.
What about PGP/GPG?
Not particularly true in the rest of the world, for instance in Europe as many have pointed out in other threads. The SMS text messages work well here, and good data protection laws/regulations ensure that there is virtually no spam or phishing. This is not a problem with the platform, but with the lack of regulation in the industry.
> "The *demo* API key is for demo purposes only. Please claim your free API key at (https://www.alphavantage.co/support/#api-key) to explore our full API offerings. It takes fewer than 20 seconds."
> TopHat has hit an error, and has ended up in a bad state. You could go back to the home page, or check the developer tools for more information.
> We are moving towards a model where the internet is dominated by a few centralized content providers and their walled gardens, and generated content may unfortunately make it even harder for grassroots online communities to survive and grow.
The author makes a good and fair analysis of the situation, yet I don't see how advocating for e-IDs is going to make the Internet a better place. This is a very ‘platform-focused’ critique that assume that the Internet tends toward more centralisation, where communities are packed together in massive platforms that regulate who communicates to whom and how.
Where do we see a need to regulate ‘fake news’ and artificially generated content? Facebook Pages? Twitter? Youtube? These are all hyper-global platforms focused on content monetization. On the other hand, local Facebook groups and the feed of your friends, Github repositories, Mastodon… may not face the same future. Maybe there's a lesson to learn here?
adMarketplace: “We may also receive technical information such as your approximate location, browser type, language settings, user agent, timestamp, cookie ID and IP addresses.”
There's a big discrepancy between Mozilla's statement and adMarketplace's privacy policy. Although there's very little information available, the fact that this technical information is anonymous doesn't appear straightforward. Cookie IDs and IP addresses alone can very easily identify some users uniquely across time and even devices.
I'm confused. I thought everyone was identified by their phone number, which can be used to receive texts from anyone.
Yikes, is retail therapy really a thing? Assuming this is a real website, there are so many things both wrong and sad here. :( You don't need to buy things (especially on Aliexpress) to be happy, to enjoy life, to make friends.
I hope the lockdown many people here are experiencing will at least make us a bit more aware of the madness of ultra-consumption we live in.
Edit: okay fine, I clicked on the ”Terms of Service”. I still don't really get the point of this website. Can you explain?
This is quite interesting to read. Most of the media coverage in France has been very hostile against the US with little evidence to back it. We've seen claims of Americans waiting on the tarmac to buy our masks and redirect planes, cash in hand ready to be given. And in major newspapers [1, 2].
[1] https://www.lefigaro.fr/actualite-france/coronavirus-les-reg... [2] https://www.liberation.fr/france/2020/04/01/une-commande-fra...
> Unfortunately, our website is currently unavailable in most European countries. We are engaged on the issue and committed to looking at options that support our full range of digital offerings to the EU market. We continue to identify technical compliance solutions that will provide all readers with our award-winning journalism.
This does not mean that your data goes through a VPN server.