Anthropic's Claude is said to improve on ChatGPT, but still has limitations
techcrunch.com
techcrunch.com
Based on the ability to get around ChatGPT's self-censoring functionality, it's unlikely the model itself was being retrained to avoid certain outputs but there was some other model used to identify inappropriate prompts. The approach from Anthropic instead seems to change the model's latent space to prefer outputs aligned with its constitution.
It's likely that OpenAI will incorporate this approach of using human oversight to train a constitutional supervised model to then train the chat model with ‘RL from AI Feedback’ (RLAIF). OpenAI's current RLHF approach seems related to improving quality of outputs while this is more related to its self-censorship (which didn't work well). I suppose this might be why they haven't released the constitution since it might be serving as some kind of moat (or the principles may be differentiable?). It's still not clear to me how a constitution can impact hallucinations.
This is giving me very strong Asimov's "three laws of robotics" vibes:
First Law
A robot may not injure a human being or, through inaction, allow a human being to come to harm.
Second Law
A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.
Third Law
A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.
This has always felt like a gaping hole to me. It seems like to work it would have to a) always make perfect predictions of the future, and b) agree with relevant humans what "harm" is.
I wouldn't say there's no AI risk, or risk from having robots and powerful intelligences, but he generally seemed to convey that we should be capable of making quite safe and quite friendly robots, given that we can make them at all. Which makes sense to me. The engineering of ethics in robots and AI doesn't seem too much more alien than the task of engineering AI itself (specially goal-oriented AI). It requires effort, attention to detail, and a progress in understanding morality and ethics that is going to be difficult for us, but it shouldn't be that fundamentally scary, as long as we collectively have the will to make them safe (in the spirit of the three laws).
I think Hollywood (in 'I, Robot' the movie) exaggerated his impression of AI danger (perhaps because of the 'Frankenstein complex'[2] Asimov coined), and the interest in catastrophe. I think 'The bicentennial man' is a complementary movie more in the hopeful spirit of Asimov.
[1] https://en.wikipedia.org/wiki/Robbie_(short_story)
Quote: " The story centers on the technophobia that surrounds robots, and how it is misplaced. Almost all previously published science fiction stories featuring robots followed the theme 'robot turns against creator'; Asimov has consistently held the belief that the Frankenstein complex was a misplaced fear, and the majority of his works attempted to provide examples of the help that robots could provide humanity. "
> A robot may not injure a human being or, through inaction, allow a human being to come to harm.
Hmmm I wonder about trolley problems
So it's always hilarious when people use them unironically/uncritically in other contexts.
Hard-coded behaviors cannot express the abstract ideas in those laws, while the ML part cannot be relied upon to accurately behave.
I believe the ML part can be relied upon to accurately behave as long as it can exhibit logical thinking at all (which, at least to a significant degree, it seems to be able to) -- in principle it would seem like its ethics potential should be proportional to its general (linguistic) reasoning potential.
I think this is exactly the benefit we have: AIs are fuzzy (like humans are), so then can understand fuzzy laws (which seemed to be a huge problem in the early days). The problem is how to get them to incorporate those laws that they can surely understand into their motivation. This doesn't seem insurmountable to me: a parallel critic prompt "Does this question and answer follow the following ethical guidelines: ... ?" (or something more complex but largely equivalent), or maybe some other form of engineering the AI thinking and motivation to include abstract goal evaluation at some stage.
And I'm not talking about logic or fuzziness here. I'm talking about something much more basic. For instance, do you have confidence that our AI can identify humans correctly? A robot can fail to follow rule 1, when it misidentifies a human as a walking cabbage. Do you have confidence that AI can identify "harm"?
I'm not saying that I don't like AI, or that they don't have benefits. I love robots and artificial intelligence, and I see plenty of practical benefits.
Asimov's theorized robots ("positronic brains") used potential-based computing (interestingly, AFAIK they predate digital computers; there's actually a few scenes in some stories where characters start using computers as New Shiny Thing) - so I'd actually argue that the original Three Laws are specifically for heuristic-based computing.
The three laws aren't really "strict", nor are they what we'd commonly consider "laws". An Asimovian robot (approximately, and IMO) doesn't think of things to do, and then discard the ones that don't match the laws; the three laws are the direct creative impetus for generating possible actions, each action "coming with" some % value in each law, and if the sum passes some threshold, the robot "decides" to do the action.
Basically every story in the "I, Robot" anthology is a story of debugging this system... which is essentially a heuristics system. One of the clearer one is the rover orbiting a hazard that it's supposed to investigate at a radius where the balance of the 2nd and 3rd laws even out.
Armchair AIist that I am... if I were to try to implement the Three Laws with current AI systems, I'd probably stick one system on at the "front" to add/modify/create/interpret prompts in a way that adds the three laws, and then stick one at the end to measure (and then filter on) how well the output adheres to them.
AI might fail to follow rule 1, simply by misidentifying a human as something else, or misidentifying what is harmful. Like how a Tesla can crash into an obstacle, because it failed to recognize that there is an obstacle.
This isn't a criticism of Asimov, of course. I think we simply haven't solved many basic problems that he probably regarded as "surely by the time robots need rules to follow, these aren't going to be issues".
> I'd probably stick one system on at the "front" to add/modify/create/interpret prompts in a way that adds the three laws, and then stick one at the end to measure (and then filter on) how well the output adheres to them.
I bet that we are going to get some sort of end-to-end simulated robot QA system. Imagine pushing your code to Github, and in a few minutes, Travis CI emails you, "28/1000 tests failed. 32 cases of bodily harm against human were reported." What a world to live in!
That'd be pretty cool, actually! A little... maybe ironic? How many stories are there about "they didn't stop to ask if they should, only if they could", and then... outsourcing that check to an AI system. Like I love it, but also, sus.
> bugs
Yep! I think positronic brains would notionally use a heuristics system ("human detector says 89%"), and AFAIK so do our AI systems. That said, your larger point still stands: what happens when such a system either fails, or is mis-calibrated, or is calibrated in a sus way?
(such as facial recognition, at first, not working on BIPOC faces... because the devs used themselves as the test subjects and weren't BIPOC)
I'm almost certain there's at least one story by Asimov on this issue; I know I've read other SciFi on this issue, although I think it's much more common to have the AIs act "better than the humans" rather than vice versa. Something like: "The rules you programmed in say Group X are humans even tho you don't treat them that way". Usually (I think?) when it's "Group X aren't humans by the programmed rules" it's apocalyptical because no-one / almost no-one fits.
I think you could probably write a cute short story about robots anthropomorphizing a lot of things in order to catch all the odd human edge-cases ("not all humans have a face").
- You give us a dollar, we'll give you four quarters!
- People ask us how we make money. The answer is simple: _volume_.
I suspect a big part of why stable diffusion managed to consume so much mindshare is that it can run on ordinary consumer hardware. On that point, I would be excited about an open-source RETRO (https://arxiv.org/pdf/2112.04426.pdf) model with comparable performance to GPT-3 that could run on consumer hardware with an NVMe SSD.
The biggest bloom I personally have run on cpu only in this fashion is 7B. It requires 4x7B of RAM plus some. On my hardware it tends to use all 32GB RAM and about ~4GB of storage during inference. At the moment I believe there is still a limitation of the smallest layer fitting in memory at once. This is why I haven't tried bigger bloom, but I believe there are ways to overcome it. Once this problem is resolved one should be able to use the same tech to use GPUs with less vram (like my 2070 with 8GB) for parts of larger models.
"The [$580M] Series B follows the company raising $124 million in a Series A round in 2021. The Series B round was led by Sam Bankman-Fried, CEO of FTX. The round also included participation from Caroline Ellison, Jim McClave, Nishad Singh, Jaan Tallinn, and the Center for Emerging Risk Research (CERR)."
Edit: to avoid ambiguity, this response is written re morality, not legality.
If the investment was made with stolen funds they can be clawed back.
If he paid any taxes, can those be clawed back from the IRS too?
OpenAI would be positively thrilled to give SBF back his money and cancel his shares. I’m not sure his creditors would similarly salivate at that deal. (EDIT: Nvm.)
Legally, its complicated, and the reason for that is because it is viewed as both morally complicated (as well simplistic approaches being viewed as creating undesired social incentives.)
Not at tolerable length for HN and do the subject justice. In very brief summary, the problem of how to balance the resolution of the interests of the two victims in this type of case, is a rather old one that has been recognized in the common law for quite a long time, leading to nuanced handling of different circumstances, which have themselves evolved over time, taking into account factors like the kind of property (real property having one general set of rules, personal property having another set of rules, but money and negotiable instrumenets having at times a different set of rules than personal property generally, etc.), the relations between the three parties, etc.
There’s been a lot of ink over the years written on the issue, though, “bona fide purchaser for value”,”innocent purchaser for value, “good-faith purchaser” are all terms for the issue. (Often, though, it will be written of in the context of a particular domain, e.g., real property transactions, commercial transactions under the UCC as compared to some particular preexisting law, etc., but it is all the same broad issue.)