Trying to get it to generate some pretty simple verilog code with extensive prompting.
It seems really bad?
Like specify what the module interface should be in the prompt and it ignores it and makes something up bad. Utterly rubbish code beyond that. Specify a calculation to be performed yet it calculates something very different.
What am I missing? Why is everyone so excited? Seems significantly worse to me than llama. Both o1-mini and claude haiku imperfect, sure, but way ahead. Both follow the same prompt and get the interface and calculation as specified. Am I doing it all wrong somehow (more than likely)?
After fixing up my open-webui install I tried "testing 1 2 3, testing. Respond with ok if you see this." Deepseek-r1:8b started trying to prove a number theory result.
Is there a chance this thing is heavily optimised for benchmarking not actual use?