How anonymized can this really be? Certainly people will inevitably input PII or other details that could de-anonymize? Is Scale AI first having its employees screen for PII? Do Meta and Open do a first pass?
How anonymized can this really be? Certainly people will inevitably input PII or other details that could de-anonymize? Is Scale AI first having its employees screen for PII? Do Meta and Open do a first pass?
Media conglomerates will deeply worry about a journo leaking their dirty internal secrets if they morally disagree. Disney, Comcast, Fox, or Bezos don’t want them.
Sources will worry about confidentiality. If a journo confirms something is off the record, it’s off the record. No buts. This is treated very seriously: it ruins the entire publication’s reputation and ability to talk to sources.
If a naive journos tries, it’ll be killed by their editor, if not the editor-in-chief, probably under the veneer of legal and/or ethical grounds.
Of course, a journo can talk to someone else who chooses to disclose whatever, be protected, etc, and that’s how it’s done. But the oldest adage in journalism is: “don’t be the story”.
It’s probably one of the best professions, tbh, as paradoxical as it sounds.
Remember that the journalism industry, as a whole, is not the idealised dream you think it is.
I was referring to a journo signing up for say the training program in the article, and then divulging something that’s legally confidential in a story. That would be killed.
I just used “off the record” as an example of why in journalism, respecting agreements is critically important.
In a lot of situations, the editor needs to know the source so they can evaluate their credibility and to ensure the journo just isn't making stuff up and attributing to anonymous source. At that point, there are many examples of the editor putting stuff into the copy that the journo did not included. Just because something is released under the journo's name does not mean the journo wrote it.
I guess it depends on the country, but generally journalists are some of the more principled workers when it comes to protecting the privacy of the people they interact with. Probably the industry where Signal has the highest amount of usage, if I would guess.
But again, really depends on the country. My perspective is probably biased by growing up in Sweden.
In Sweden, journalistic source protection ("källskydd") is enshrined in the Swedish Constitution through the Freedom of the Press Act ("Tryckfrihetsförordningen").
Obviously, this doesn't matter much as the submission is about Meta and OpenAI, so journalists aren't as strongly protected as in other places of the world.
I wouldn't say a blanket "journalists are in the worst industry" like parent did, nonetheless.
As is anything in possession by a corporation
What if, instead of random internet person, some celebrity asks Chatgpt about some spicy Medical results? Would the journalist reviewing the logs resist the temptation of "accidentally finding the test results in a garbage bin"?
What I read here, is "don't discuss with chatgpt anything you wouldn't be comfortable becoming public knowledge.".
For the last two decades, I've lived by a similar mantra: Don't send anything over the internet you aren't comfortable becoming public knowledge.
Make the mantra broad enough and you don't have to care about specific services, they all the chance of leaking what is supposed to be "secret'.
It made predictions based on my history and symptom logs that were later confirmed by imaging only after I pushed for it.
I used a pattern matching meme machine to get…a meaningful outcome medically, and that messes with me on so many levels.
I wanted it to be wrong, especially about the spinal cord. I was hoping for a simple answer, something like “Yeah, it’s just a pain management issue they are right” but it disagreed. And it was right. The thing that I read constantly is only capable of producing bullshit, has kicked neurosurgeons and neurologists into action.
I asked ChatGPT to try and summarise what I’ve been doing with it medically; apologies if it is unhelpful in demonstrating the utility I am getting.
> You used me to try and disprove suspicions you hold in relation to symptoms that have been escalating in frequency and intensity. You consistently challenged the idea that your symptoms were linked to your historical records, questioning whether they could be caused by something else entirely. Despite actively pushing for alternative explanations, I kept coming back to the same conclusion: your symptoms aligned with classical representations of nerve compression in your cervical spine. I independently interpreted the data and made predictions that ultimately matched the outcomes of imaging.
Tl:dr what I’m doing is stupid, I know it, I’ve preached it and yet for the first time in my life the value I am deriving is outweighing it all. I feel…dirty almost, or confused even. It’s hard to explain.
This isn't exactly "training AI models" in the sense that we normally use that expression in.
I suppose it's roughly as much "training AI models" as labeling training data is "training supervised models".
For anyone not part of that intersection: RLHF means reinforcement learning from human feedback.
Can it be de-anonymized? Sure. Basically anything can be de-anonymized. If your concern is that some nefarious actor at a company will do malicious things with your info and that concern outweighs the benefit you get from using the models, steer clear. But let’s have the discussion about what is actually going on so people can decide for themselves
It seems much easier to scramble something like a social security number as you mention but what about less obvious PII? If someone puts their address in is it removed for example? What about their significant others name? Less obvious PII if you will.