Adding to this, reducing cost means they've reduced compute and improved quality per unit of computation.
That matters! It matters that these systems currently require an embarassing amount of energy to run otherwise. It matters that these models could be portable to the point where they can be run locally on portable consumer electronics rather than being locked into remote compute and HTTP calls.
a few prompts recently, it just answered for me instead of refusing, to my surprise, so if feels like it's been getting better, but unfortunately there's no hard data I have on that, it's more of a feeling, so it's hard to prove that's actually true to someone else.
then again, ChatGPT-4o was recently found to be unable to do 9.11 - 9.9 correctly unless pressed, so there's still a long ways to go.
https://chatgpt.com/share/47d260ae-0c62-48ab-8d3b-84767acd7f...
Why do you think these tools understand how they work enough to explain it? If this were the case they wouldn't be a blackbox to begin with.
It blows my mind that there are people on HN that don't know this.
LLMs are basically Markov Chains on a massive dose of steroids. It looks at the last $context_window tokens and decides what the next token should be. They just work differently in that a neural network creates more "fuzzy" matching than a Markov Chain, which is purely statistical.
How? They have a hard cut off context, and randomly generates next state from that. That is the definition of a markov chain.
It is a markov chain with a pretty large state at every point, but still a markov chain.
Q: Write parameterized Python that gives the answer to 9.11 - 9.9 and 9.12 - 9.9. Run the Python to show the answer to each pair of parameters.
A: The results for the parameter pairs are as follows:
- 9.11 - 9.9 = -0.79
- 9.12 - 9.9 = -0.78 [>_]
The [>_] is a link. Clicking it opens a pop-up window, showing: def calculate_difference(a, b):
return a - b
# Define the pairs of parameters
pairs = [(9.11, 9.9), (9.12, 9.9)]
# Calculate and print the results for each pair
results = {f"{a} - {b}": calculate_difference(a, b) for a, b in pairs}
results
Result:
{'9.11 - 9.9': -0.7900000000000009, '9.12 - 9.9': -0.7800000000000011}// Note, if I ask it simply "compute the answer to 9.11 - 9.9", I get nonsense such as a negative number, and no Python. It doesn't "automagically" write code, it has to be nudged. If you see "Analyzing", it's fired up a sandboxed Python container.
llama3-groq-70b-8192-tool-use-preview llama3-groq-8b-8192-tool-use-preview
Sonnet3.5 also has tool usage, and is slightly cheaper than 4o ($3 input vs $5).
If you _only_ need tool usage, then 4o-mini or others will be cheaper...
mistal's Mixtral 8x22B has function calling for $2/$6, which is a bit cheaper than 4o.