'Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?'
'Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?'
https://huggingface.co/datasets/mbpp?row=98
At the end of the day LLMs in their current iteration aren't intended to do even moderately difficult tasks on their own but it's fun to query them to see progress when new claims are made.
"My favorite thing to ask the models designed for programming is ....... None of them ever get it right"
I read "benchmark".
I'm glad they can't quite manage this yet. Means I still have a job.
The whole point is to prompt less?
it is not. But the artifacts generated through the steps will be code. The last prompt will have most of the code supplied to it as the context.
What has happened to HN discourse recently?