Wow, interesting. Do you have any example for this?
I've realized that LLMs are fairly good at string processing tasks that a really complex regex might also do, so I can see the point in those.
Wow, interesting. Do you have any example for this?
I've realized that LLMs are fairly good at string processing tasks that a really complex regex might also do, so I can see the point in those.
(separately, synthesizing a trace from this kind of data is impossible to get 100% correct for other reasons, but hey, it's a fun thing to try)
https://www.artisana.ai/articles/gpt-4-outperforms-elite-cro...
You can also probably distill a large model to a smaller one while maintaining a lot of performance. DistillBert is almost as good as Bert at a fraction of the inference cost.
GPT-3.5 and 4 also currently aren’t deterministic even with temperature zero, which is a nightmare for debugging.
What's definitely true is that getting decent data often takes some care, especially in how you define the task. And mechanical turk is often especially tricky to use well.