My initial goal is to let users make webapp prototypes and iterate on them by writing tickets for the AI to complete.
I for some reason call it Code+=AI: https://codeplusequalsai.com
My initial goal is to let users make webapp prototypes and iterate on them by writing tickets for the AI to complete.
I for some reason call it Code+=AI: https://codeplusequalsai.com
For example they're oh-so keen on doing math themselves, even though they're shit at it and I instructed them not to do any math.
It's also hit and miss if they implement the right method or not, even with low temperature.
In my case I was experimenting with translating simple word problems into matlab scripts, so a resultcould be computed.
Do you find the AST approach helps? Or is it mostly just throwing compute at it, ie larger more better?
However, I have noticed that the AST code quality heavily depends on how common it is in the training set. I think I will have to add documentation to it through RAG or something - because OpenAI's models that I'm using seem to have limited experience writing esprima for JavaScript for example.
So it's hit or miss. In some cases I do feel like I'm throwing stupid compute at solving small problems and it's unnecessary - however, as I work on the project, it is getting better and better at successfully making the modifications. Some of that is me improving the prompts, some of it is OpenAI improving the models themselves, and some of it is the infrastructure I'm building for the project itself.
I did notice a huge improvement when o1-mini released. It is dramatically better at writing the AST code than GPT-4o or 4o-mini. I haven't tried Claude 3.5 yet but I've been hearing it does an exceptional job at code writing - not sure about my AST requirements though!