Claude 3 Haiku: our fastest model yet
anthropic.com
anthropic.com
Specifically, the smallest/likely will be the most popular model is now available when it wasn't then. (The model ID is claude-3-haiku-20240307 ). Notably, this is also a cheap model that supports image input, but per the documentation you can only provide 20 images at a time which won't work for video inputs.
Testing around image inputs in the web Workbench, it's surprisingly good for the price.
GPT-3.5 Turbo is $0.50/$1.50
I've updated the Claude 3 plugin for my LLM CLI tool to support the new model: https://github.com/simonw/llm-claude-3/releases/tag/0.3
pipx install llm
llm install llm-claude-3
llm keys set claude
# Paste Anthropic API key here
llm -m claude-3-haiku 'Fun facts about armadillos'
It's pretty fast! Animated GIF here: https://github.com/simonw/llm-claude-3/issues/3#issuecomment...- 10-15 seconds for 400 tokens out, and 4,000-10,000 tokens in.
- 6-8 seconds when using Claude Instant for the same prompts
Hoping it's just a rush at launch.
Llama 2 70B (4096 Context Length) ~300 tokens/s $0.70/$0.80
Llama 2 7B (2048 Context Length) ~750 tokens/s $0.10/$0.10
Mixtral, 8x7B SMoE (32K Context Length) ~480 tokens/s $0.27/$0.27
Gemma 7B (8K Context Length) ~820 tokens/s $0.10/$0.10
[1] https://wow.groq.com/They say they're just waiting on implementing billing, but at this point it reads more like "we wouldn't be able to meet demand of all your request usages".
-
Groq is going through all that to offer 500tk/s theoretically, meanwhile I'm seeing Fireworks.ai come in at 300+tk/s in production use.
A few days ago, I decided to give Claude a try , so I created an account, verified my phone number, and successfully logged in. After a warm welcome from Claude and presenting myself, I entered my very first prompt, which reads: "What do you know about Hacker News?". I pressed ENTER, and after a second, it replied:
"Your account has been disabled after an automatic review of your recent activities that violate our Terms of Service. Please review our Terms of Service and Acceptable Use Policy for more information."
I contacted the support team, and after a day, they replied and redirected me to a Google Forms to fill, which I still didn't fill.
I very much miss an official comment on this though.
It's not exactly inspiring confidence to subscribe to a product you may randomly be locked out from, with no comment from the company behind it.
Haiku seems like it adds to this lead. Having a cheaper and better model than GPT3.5 for processing large amounts of documents is great.
Props to the Anthropic team.
Zenfetch is now primarily powered by Claude 3 family of models :O https://www.zenfetch.com
There, Sonnet is within the margin of error of current GPT-4 and Opus the same but as for the GPT-4 previews.
Above all Anthropic seem to have found themselves a very nice training set, watching how nicely results are retained as they go down in model sizes.
If you want to email bkrausz at anthropic.com with your phone number I'm happy to check logs (assuming it's still not working).