OpenAI slashes the cost of using its AI with a "mini" model
wired.com
wired.com
llama3-groq-70b-8192-tool-use-preview llama3-groq-8b-8192-tool-use-preview
Sonnet3.5 also has tool usage, and is slightly cheaper than 4o ($3 input vs $5).
If you _only_ need tool usage, then 4o-mini or others will be cheaper...
mistal's Mixtral 8x22B has function calling for $2/$6, which is a bit cheaper than 4o.
a few prompts recently, it just answered for me instead of refusing, to my surprise, so if feels like it's been getting better, but unfortunately there's no hard data I have on that, it's more of a feeling, so it's hard to prove that's actually true to someone else.
then again, ChatGPT-4o was recently found to be unable to do 9.11 - 9.9 correctly unless pressed, so there's still a long ways to go.
https://chatgpt.com/share/47d260ae-0c62-48ab-8d3b-84767acd7f...
Why do you think these tools understand how they work enough to explain it? If this were the case they wouldn't be a blackbox to begin with.
Q: Write parameterized Python that gives the answer to 9.11 - 9.9 and 9.12 - 9.9. Run the Python to show the answer to each pair of parameters.
A: The results for the parameter pairs are as follows:
- 9.11 - 9.9 = -0.79
- 9.12 - 9.9 = -0.78 [>_]
The [>_] is a link. Clicking it opens a pop-up window, showing: def calculate_difference(a, b):
return a - b
# Define the pairs of parameters
pairs = [(9.11, 9.9), (9.12, 9.9)]
# Calculate and print the results for each pair
results = {f"{a} - {b}": calculate_difference(a, b) for a, b in pairs}
results
Result:
{'9.11 - 9.9': -0.7900000000000009, '9.12 - 9.9': -0.7800000000000011}// Note, if I ask it simply "compute the answer to 9.11 - 9.9", I get nonsense such as a negative number, and no Python. It doesn't "automagically" write code, it has to be nudged. If you see "Analyzing", it's fired up a sandboxed Python container.
It blows my mind that there are people on HN that don't know this.
LLMs are basically Markov Chains on a massive dose of steroids. It looks at the last $context_window tokens and decides what the next token should be. They just work differently in that a neural network creates more "fuzzy" matching than a Markov Chain, which is purely statistical.
How? They have a hard cut off context, and randomly generates next state from that. That is the definition of a markov chain.
It is a markov chain with a pretty large state at every point, but still a markov chain.
Adding to this, reducing cost means they've reduced compute and improved quality per unit of computation.
That matters! It matters that these systems currently require an embarassing amount of energy to run otherwise. It matters that these models could be portable to the point where they can be run locally on portable consumer electronics rather than being locked into remote compute and HTTP calls.
Oh and after pulling the pledged money, that other party could go buy a social media site and use it to whine about how they aren't a non-profit anymore... while knowing they're the cause of it.
I'm sure there's a method to permit it, but it doesn't seem like anyone's worked it out yet.
Couple of issues I can see: (1) most devices out there would probably be mobile, so no NVIDIA/CUDA for you; (2) even with binding to, say, Apple Silicon, you might still be memory limited, i.e. can you fit the entire net on a single mobile GPU; (3) network latency?
The greatest limiter for training isn't raw computing power, but storage. If your model is 400B parameters, you need 800GB to store it, assuming fp16. Then you need another 800 GB for calculating gradients. Sharding all this out means transferring a lot of data. If your device only stores 1B parameters out of the 400B, that means having to download 2 GB of data, doing your share of the work, then uploading 2 GB of results. Even with gigabit internet, you'll spend an order of magnitude more time transferring data than actually processing it.
At that point, it'd be faster to train on standard-specced PC that had to constantly page out most of the model.
Looking at https://openai.com/index/gpt-4o-mini-advancing-cost-efficien... , gpt-4o-mini is better than gpt-3.5 but worse than gpt-4o, as was expected. gpt-4o-mini is cheaper than both, however. Independent third-party performance benchmarks will help.
Good to know, although this is targeting cheaper API use for specific applications in which a second-tier model is sufficient. Note however that according to the LMSYS Leaderboard, GPT-4o rates slightly higher than Claude 3.5 Sonnet.
> mondain stuffs
Mundane, not mondain.
No idea what the leaderboard says, it's just been night and day for me since Sonnet 3.5 got out. Maybe my use case is just what sonnet does best.
Since the launch of Claude 3 Opus, and then Claude 3.5 Sonnet, they have been significantly behind Anthropic in terms of the general intelligence of their models. And instead of deploying something on par or better, they are making demos of video generation (Sora) or audio-to-audio models, not releasing anything.
GPT-4o is quite bad at coding, often getting stuck in a loop, and “fixing” buggy code by rewriting it without any changes.
GPT-4o is speculated to be a distillation of a larger model, and now GPT-4o-mini is an even dumber smaller model. But what’s the point?
Who is actually using small/fast/cheap/dumb models in production apps? Most real apps require higher reliability than even the biggest/slowest/priciest/smartest models can provide today. For the use case of transformers that has taken off, aiding students and knowledge workers in one-off tasks like writing code and prose, most users want smarter, more reliable outputs, even at the expense of speed and cost.
GPT-4o-mini seems like a move to increase margins, not make customers happier. That, like demoing products without launching them, is what big old slow corporations do, not how world-leading startups operate.
Extremely concise, formal. As short as possible. Assume I am an industry expert in any topic we discuss. Answer assuming I have the highest level of intellect possible, and do not require explication regardless of the sophistication of the topic. In cases where one approach among many is superior, offer an opinionated argument in favor of that approach.
Edit: I’m amazed by how offended some people are by such a simple question.
But let me know if you find something. I just don't think something tiny like phi-3 which could run on a VPS, although great for it's size, is at all comparable to this stuff in terms of ability.
My AI server takes about 60W idle and 300-350W while running a query in llama3. At a kWh price of 0.15€ that ends up at about 7-10€ a month if it's not loaded too heavily. Not bad IMO.
The server could be more energy optimized though. But that would cost me also.
You probably want to limit yourself if you do have a 16gb card because you still need to fit the context window in memory too.
I hope the new mistral comes out soon for ollama.
True about the context window but with llama3 this was not a problem as it has such a small context window anyway.
> Priced at 15 cents per million input tokens and 60 cents per million output tokens, the GPT-4o mini is more than 60% cheaper than GPT-3.5 Turbo, OpenAI said. It currently outperforms the GPT-4 model on chat preferences and scored 82% on Massive Multitask Language Understanding (MMLU), OpenAI said.
...
> The GPT-4o mini model's score compared with 77.9% for Google's Gemini Flash and 73.8% for Anthropic's Claude Haiku, according to OpenAI.
For some more context: We don't know the size of 4o-mini but Mistral's just released NeMo 12B scores 68% on the MMLU. [2]
[1]: https://www.reuters.com/technology/artificial-intelligence/o...
Gemma 2 27B scored: 75.2 in MMLU
LLama 3 70B scored: 79.5 in MMLU
Haiku scored: 75.2 in MMLU
GPT 3.5 scored: 70.0 in MMLU
Based on pricing I see in openrouter.ai across different providers this seems like the cheapest model for this kind of performance.
ref: [0] https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bb...
[1] https://blog.google/technology/developers/google-gemma-2/
Quality to Price graph suggests gpt3.5 was the worst now 4o-mini inched out all others of that lower league. It supposedly gets you flash/llama3-70B tier quality at around llama3-8B price.