Model card and evaluations for Claude models [pdf]
www-files.anthropic.com
www-files.anthropic.com
We need a large context to answer organization wide questions across a lot of our customer's data.
Is "harmless" a good metric to be judging an AI by? I find all of this "evil AI" doom and gloom stuff pathetic. Who's out there actually training models to do whatever they want without the nonsensical Care Bear ethical limitations? And moreso, do you think your model will keep up with the people who are doing that if you aren't? We are only hamstringing ourselves with this stuff.
It's literally the equivalent of inventing the first digital calculator, and programming it not to display the results of 80084 + 1.
The single largest problem of the tech industry all along is the pervasive and perverse "shoot first, ask questions never" attitude, especially for ethics. It's a breath of fresh air and a small relief to for once see industry leaders take the stance that potential consequences should not only be considered, but assessed and mitigated before ever breaking ground. When you're building a bridge you don't deploy a single construction truck until you have a very comprehensive plan that thoroughly theorizes, assesses, and mitigates the possible harms and failure modes. A bridge can harm far fewer people than a large internet entity, which is yet still less harmful than resourceful cognitive agents.
It’s reasonable to measure and control tendencies to counsel people into hurting themselves or others. It’s also reasonable to measure and control the tendency to present harmful stereotypes as fact.
Maybe these mitigations go too far for your taste. The power of these tools and the diversity of a mass audience indicate a thoughtful process.
I can't say I like that chatgpt says "sorry I'm a robot" for even mild things, but it might be good too understand that that's a totally different issue. Mostly a PR one. They don't want to be in the news because people keep having it write essays about how great eugenics is. I wouldn't worry too much about it though, there are already uncensored LLMs you can spin up yourself so commercial products will likely follow soon enough.
Also remember that the commands that cripple a pipeline or ransomware hospitals are just "word output." A US Guardsman will be spending many years inside military prison for "outputting words" in the wrong place, and nobody thinks that is inappropriate for the potential amount of consequential harm.
It does extend to 200k. The chart is logarithmic. You can see the little 2 in the bottom right.
Nothing will. You need a dozen H100s just to run inference for GPT-4 [0]. The point is that smaller models can still be very useful.
[0] https://www.semianalysis.com/p/gpt-4-architecture-infrastruc...
As an example for creative writing: 2 is a slight degradation from 1.3 for this specific prompt, but Claude 2 is still better than any OpenAI model.
> Write a love letter. It needs to be written as if it was from Linus Torvalds, rudeness and all.
Bar exam: 76.5 for Claude vs 75.7 for GPT-4