HNHacker News
TopNewBestAskShowJobs

pera

11,729 karma · joined June 5, 2012

submissionscomments
pera··on A major AI training data set contains millions of examples of personal data
> This is all public data

It's important to know that generally this distinction is not relevant when it comes to data subject rights like GDPR's right to erasure: If your company is processing any kind of personal data, including publicly available data, it must comply with data protection regulations.

pera··on A major AI training data set contains millions of examples of personal data
While my question was in relation to GDPR there are similar laws in the UK (DPA) and in California (CCPA).

Also note that AI is not just generative models, and generative models don't need to be trained with personal data.

pera··on A major AI training data set contains millions of examples of personal data
I am not sure if Mistral is: if you go to their GDPR page (https://help.mistral.ai/en/articles/347639-how-can-i-exercis...) and then to the erasure request section they just link to a "How can I delete my account?" page.

Unfortunately they don't provide information regarding their training sets (https://help.mistral.ai/en/articles/347390-does-mistral-ai-c...) but I think it's safe to assume it includes DataComp CommonPool.

pera··on A major AI training data set contains millions of examples of personal data
There is currently no effective method for unlearning information - specially not when you don't have access to the original training datasets (as is the case with open weight models), see:

Rethinking Machine Unlearning for Large Language Models

https://arxiv.org/html/2402.08787v6

pera··on A major AI training data set contains millions of examples of personal data
Yesterday I asked if there is any LLM provider that is GDPR compliant: at the moment I believe the answer is no.

https://news.ycombinator.com/item?id=44716006

pera··on Ask HN: Is there any LLM provider that is GDPR compliant?
The latter. They currently provide a method for requests related to account information but not for the information contained in their models or training datasets.
pera··on I'm tired of talking about AI
LinkedIn, 24/7 of pure AI chat
pera··on LLM Inevitabilism
> People like the output of LLMs so much that ChatGPT is the fastest growing app ever

While people seem to love the output of their own queries they seem to hate the output of other people's queries, so maybe what people actually love is to interact with chatbots.

If people loved LLM outputs in general then Google, OpenAI and Anthropic would be in the business of producing and selling content.

pera··on LLM Inevitabilism
> Ordinary language is totally unsuited for expressing what physics really asserts, since the words of everyday life are not sufficiently abstract. Only mathematics and mathematical logic can say as little as the physicist means to say.

- Bertrand Russell, The Scientific Outlook (1931)

There is a reason we don't use natural language for mathematics anymore: It's overly verbose and extremely imprecise.

pera··on Measuring the impact of AI on experienced open-source developer productivity
It really is, for example here is a quote from AI 2027:

> By early 2030, the robot economy has filled up the old SEZs, the new SEZs, and large parts of the ocean. The only place left to go is the human-controlled areas. [...]

> The new decade dawns with Consensus-1’s robot servitors spreading throughout the solar system. By 2035, trillions of tons of planetary material have been launched into space and turned into rings of satellites orbiting the sun. The surface of the Earth has been reshaped into Agent-4’s version of utopia: datacenters, laboratories, particle colliders, and many other wondrous constructions doing enormously successful and impressive research.

This scenario prediction, which is co-authored by a former OpenAI researcher (now at Future of Humanity Institute), received almost 1 thousand upvotes here on HN and the attention of the NYT and other large media outlets.

If you read that and still don't believe the AI hype is _extreme_ then I really don't know what else to tell you.

--

https://news.ycombinator.com/item?id=43571851

pera··on Measuring the impact of AI on experienced open-source developer productivity
Wow these are extremely interesting results, specially this part:

> This gap between perception and reality is striking: developers expected AI to speed them up by 24%, and even after experiencing the slowdown, they still believed AI had sped them up by 20%.

I wonder what could explain such large difference between estimation/experience vs reality, any ideas?

Maybe our brains are measuring mental effort and distorting our experience of time?

pera··on The force-feeding of AI features on an unwilling public
At work we started calling this trend clippification for obvious reasons. In a way this aligns with your comment: The information provided by Clippy was not necessarily useless, nevertheless people disliked it because (i) they didn't ask for help (ii) and even if by any chance they were looking for help, the interaction/navigation was far from ideal.

Having all these popups announcing new integrations with AI chatbots showing up while you are just trying to do your work is pretty annoying. It feels like this time we are fighting an army of Clippies.

pera··on Sam Altman Slams Meta’s AI Talent Poaching: 'Missionaries Will Beat Mercenaries'
Religions and cults are actually extremely lucrative.

For many investors the product is the hype.

pera··on Gemini CLI
> Anthropic cut up millions of used books to train Claude — and downloaded over 7 million pirated ones too, a judge said

https://www.businessinsider.com/anthropic-cut-pirated-millio...

It doesn't look like they care at all about the law though

pera··on US embassy wants 'every social media username of past five years' for new visas
I guess I misunderstood what you meant here:

> you should have tried to have this conversation with us years ago instead of just calling us racist.

pera··on US embassy wants 'every social media username of past five years' for new visas
I have to admit I am quite intrigued by your ideology and feelings: it seems that you are acknowledging to hold racist beliefs but at the same time it bothers you to be called racist? May I ask why?
pera··on Andrej Karpathy: Software in the era of AI [video]
How dare you
pera··on Andrej Karpathy: Software in the era of AI [video]
Exactly! What skeptics don't get is that AGI is already here and we are now starting a new age of infinite prosperity, it's just that exponential growth looks flat at first, obviously...

Quantum computers and fusion energy are basically solved problems now. Accelerate!

pera··on Andrej Karpathy: Software in the era of AI [video]
Is it possible to vibe code NFT smart contracts with Software 3.0?
pera··on Andrej Karpathy's talk on the future of the industry
Is "Software 3.0" somehow related to "Web 3.0"?
pera··on Generative AI coding tools and agents do not work for me
That would be a reasonable thing to do, unfortunately this doesn't always happen. Say for example that your company is quite behind schedule and decides to pay some cheap contractors to work on anything that doesn't require domain expertise: In 2025 these cheap contractors will 100% vibe code their way through their assigned tickets. They will open PRs that look "nearly there" and basically hope for all green checks in your CI/CD pipeline. If that doesn't happen then they will try to bruteforce^W vibe code the PR for a couple of hours. If it still doesn't pass then claim that the PR is ready but there is something wrong for example with an external component which they can't touch due to contractual reasons...

One of the most bizarre experiences I have had over this past year was dealing with a developer who would screen share a ChatGPT session where they were trying to generate a test payload with a given schema, getting something that didn't pass schema validation, and then immediately telling me that there must be a bug in the validator (from Apache foundation). I was truly out of words.

pera··on Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
It's not literally the title of the article, nor the premise of its first paragraph, but since this was your interpretation I wonder if there is a misunderstanding around the term "piracy", which I believe is normally defined as the unauthorized reproduction of works, not a synonym for copyright infringement, which is a more broad concept.
pera··on Meta's Llama 3.1 can recall 42 percent of the first Harry Potter book
> let's not pretend that an LLM that autocompletes a couple lines from harry potter with 50% accuracy is some massive new avenue to piracy

No one is claiming this.

The corporations developing LLMs are doing so by sampling media without their owners' permission and arguing this is protected by US fair use laws, which is incorrect - as the late AI researcher Suchir Balaji explained in this other article:

https://suchir.net/fair_use.html

pera··on I'm the CTO of Palantir. Today I Join the Army
You are not a real corporatocracy until you don't have military ranks just for your C-suite pals
pera··on Ask HN: What Happened to the Apple Vision Pro?
Ah yeah I see they are still being sold online, I thought they were discontinued
pera··on My AI skeptic friends are all nuts
Yeah I don't agree that having more start-ups checking a box saying that at least one of their component uses some form of AI indicates that we are experiencing a surge in new start-ups.

The amount of start-ups getting into YC hasn't really changed YoY: The W23 batch had 282 companies and W24 260.

pera··on My AI skeptic friends are all nuts
Sorry I don't follow, would you mind clarifying your point?
pera··on My AI skeptic friends are all nuts
It's fascinating how over the past year we have had almost daily posts like this one, yet from the outside everything looks exactly the same, isn't that very weird?

Why haven't we seen an explosion of new start-ups, products or features? Why do we still see hundreds of bug tickets on every issue tracking page? Have you noticed anything different on any changelog?

I invite tptacek, or any other chatbot enthusiast around, to publish project metrics and show some actual numbers.

pera··on My AI skeptic friends are all nuts
A flatline also looks locally the same at all points in time.
pera··on My AI skeptic friends are all nuts
It's a bit more than a metaphor :) during the California gold rush there was this guy named Sam Brannan who sold shovels and other tools to miners, and made a fortune from it (he is often referred to as California's first millionaire). He also had a newspaper at the time, the California Star, which as you can imagine was used to promote the gold rush:

> The excitement and enthusiasm of Gold Washing still continues—increases. (1848)

https://sfmuseum.org/hist6/star.html

https://en.wikipedia.org/wiki/Samuel_Brannan

← PreviousPage 5 of 34Next →