This is unnecessary and rude, you should hang your head in shame for that. I wish some people in this community weren't so reactionary and would engage with empathy instead of trying to personally roast people as soon as they don't agree with something.
Someone can tell a story on the internet, it doesn't have to be some rigorous experiment or proof.
Computer Science has a massive ethics crisis. Uncritical adulation and a total lack of accountability or consequence is part of that.
There is a massive misallocation of capital which is burning opportunity for our society. Users are getting terrible experiences and systems because people read this sort of thing and believe it. Trust in our technology is eroded and this has consequences for actual people. We have abandoned the standards that protected people and you have the view that these standards, or a shadow of them even, are unnecessary?
Someone has taught a generation that this is all ok, it isn't.
You want to have a conversation about quality and ethics in computing and how this post can be pushing a narrative that is not in line with your views on this, I think that is worthwhile to have. But personal denigration of someone else isn't necessary in doing that.
It's difficult to set up evals, especially with production code situations. Any tips?
The principal you need to work to is that you need to create evidence that other people will find compelling and then show that you have interrogated your results to show that you have checked that it's really working better than chance and not the result of some fluke or other. Finally you need to find a way to explain what's happening - like an actual mechanism.
1. Find or make a data set - I've been using code_search_net to try and study the ability of LLM's to document code, specifically the impact optimising n-shot learning on them*, this may not be close enough to your application, but you need many examples to draw conclusions. It's likely that you will have to do some statistics to demonstrate the effects of your innovations so you probably need around 100 examples.
2. Results from one model may not be informative enough, it might be useful/necessary to compare several different models to see if the effect you are finding is consistent or whether some special feature of a model is what is required. For example, does this effect work with only the largest and most sophisticated modern models, or is this something that can be seen to a greater or lesser effect with a variety of models?
3) You need to ablate - what is it in the setup that is most impactful? What happens if we change a word or add a word to the prompt? Does this work on long code snippits? Does it work on code with many functions? Does it work on code from particular languages?
4) You need a quantitative measure of performance. I am a liberal sort , but I will not be convinced by an assertion that "it worked better than before" or "this review is like an senior, I think". There needs to be a number that someone like me can't argue with - or at least, can argue with but can't dismiss.
*I couldn't make it work, I think because the search space for finding good prompting shots (sample function)is vast and the output space (possible documents) is vast. Many bothans died in order to bring you this very very very (in hindsight with about $200 of OpenAI spending) obvious result. Having said that I am not confident that it couldn't be made to work at this point so I haven't written it up and won't make any sort of definitive claim. Mainly I wonder if there is a heuristic that I could use to choose examples a-priori instead of trying them at random. I did try shorter examples and I did try more typical (size) examples. The other issue is that I am using sentence similarity as a measure of quality, but that isn't something I am confident of.
People can and are free to tell stories if they want. It's not some failing. You don't have to engage with it anymore than that.