And that’s a win?
And that’s a win?
Sure, it absolutely might be a win. It depends on just how much accuracy they needed in the checking system in question.
It's also worth noting that one could utilize both. The assumed fast, low cost 50 lines of code on your server that takes care of the easy 97%. And then throw GPT4 at the stray hard cases. It requires being able to correctly identify when your code isn't up to the task of course.
Would be very interested in the longevity of this solution. It works today, but will it work in a month/year? A library file on the computer running the rest of the code isn't going to change.
So for example: gpt-4-0314, or gpt-3.5-turbo-0613, etc.
The latency issue is definitely true. Ideally the cost could be limited to a very small percentage of hard cases (which you first have to identify).
[0] https://matt-rickard.com/foundational-models-are-not-enough
[1] https://arxiv.org/pdf/2308.02828.pdf
[2] https://www.sitation.com/non-determinism-in-ai-llm-output/
[3] https://towardsdatascience.com/the-magic-of-llms-prompt-engi...
You can to an extent dictate GPT's determinism with settings you can pass along in the API, combined with the parent already proclaiming they saw a 100% success rate.
So how do you know it wouldn't be enough? The parent is already saying their test suite indicates it is enough. What tests have you run counter to their claim to show it fails? And how do you know the parent can't increase the determinism even further beyond what they were already using in their testing (and decreasing the risk of negative outcomes by doing so)?
I think this is a cute use case. I've recently outsourced categorizing the titles of user created tutorials into groups by relative similarity, to great effect. Took a few minutes.
It's definitely a win in my book.