This is the problem I have with the GPT models. I don't think I can trust them for anything actually important.
This is the problem I have with the GPT models. I don't think I can trust them for anything actually important.
[1] https://www.hacker-jobs.com [2] https://marcotm.com/articles/information-extraction-with-lar...
Not exactly true; https://platform.openai.com/playground
You absolutely should think about different kinds of models, especially for tasks that don't truly require generative output.
If all you are doing is classification, I'd grab some ML toolkit that has a time-limited model search and just take whatever it selects for you.
Binary classifiers are the epitome of inspectable. You can follow things all the way through the pipeline and figure out exactly where we went off the rails.
You can have your cake & eat it too. Perhaps you have a classification front-end that uses more deterministic techniques that then feeds into a generative back-end.
You can't, not absolutely. You can have some level of confidence, like 99.99%, which is probably good enough tbh (and I'm a sceptic of these tools) and honestly, it is probably better than a human, on average, at this!
But if that is a deal-killer (and it sometimes is!) then yeah, sorry - there aren't workarounds here.
I've noticed this discussion tends to get too theoretical too quickly. I'm uninterested in perfection, 99.99% would be good enough. 70% wouldn't. The actual number is something specific, knowable, and hopefully improving.
You can get to 99.9%+ with good data and well designed prompts. I'm sure it would be above 90% even with almost intentionally bad prompts, tbh.
This afternoon I tried to use Codium to autocomplete some capnproto Rust code. Everything it generated was totally wrong. For example, it used member functions on non-existent structs rather than the correct free functions.
But I'll give it some credit: that's an obscure library in a less popular language.
This isn't what I said at all. I said with summarizing data.
I also would not trust it with anything important, but there can be good applications for something that works 9/10 times.
Useful answer - fine tune on large training set, set temperature to 0, monitor token probability and highlight risk when probability < some threshold.