ChatGPT4 writes Stan code so I don’t have to
statmodeling.stat.columbia.edu
statmodeling.stat.columbia.edu
Using it to help trouble shoot errors in logs when you’re feeling frustrated, and get a basic framework up from scratch to start modifying toward the your goal are the best use cases I’ve seen. The ability to get yourself unstuck in minutes is life changing. The biggest constraint right now is short term memory loss. It can nearly always recognize on a second pass through what mistake it made.
Keep in mind 32kb token requests are already in beta… and 4kb requests are what you have access to now. So… the memory thing is potentially a dead point soon. Though at 6 cents per 1kb if it actually has to volley 32k over to you it’ll almost cost $2.00 based on current API pricing.
That 32k option would allow GPT-4 to hold every non-monolithic project I’ve ever worked on in working memory. The pertinent parts anyway, it wouldn’t need to read CSS files or imported modules unless it really gets forced down a specific rabbit hole.
This will be one more reason new projects won’t be monolithic architectures. Easier for the AI to help out.
> Working with ChatGPT4 for this task was like working with a good intern programmer who is rather forgetful and a bit scattered.
And the key take-away -- again, for me -- is this:
> But: This has already changed the way I code. I am not sure I am ever going to write code completely by hand again! And I expect these tools to get better fast.
https://learnprompting.org/docs/intermediate/chain_of_though...
Providing context, getting GPT to explain the steps before doing them, creating feedback loops, etc is super helpful when doing this stuff.
It's still useful to just go "write this function in x language" but anything beyond a Stack-Overflow type copy/paste script it helps to do a lot more hand-holding. The ROI is worth it.
Worst case scenario, throw the code back into GPT and ask it to make the changes you want to make.
This is also one of the reasons I transitioned to developing AI models myself. At least I'll be safe until they start rewriting themselves from the ground up. And at that point we'll all have bigger worries to deal with anyway.
The key is to a) give it context, and b) give it constraints, and to do it from the very first prompt itself.
Most chatGPT failures that I've seen add context and constraints after the initial prompt, which someone sends GPT into a loop.
If you have data in a particular format, mention that in the first prompt. Also helpful to tell it to optimize the code for scalability and to use any external libraries if necessary.
These tools are essentially software programs, and like any piece of software, you need to know the right commands to get maximum utility out of them.
"write a function in java to decode a rfc 3588 diameter header"
They both need a lot of hand-holding. Too much to make it worth using them.
Things they get wrong...
1) They are confident about the structure, but are 100% wrong.
2) They get Java types wrong, and try to put unsigned data into the same size signed type (8 bits into a byte, 24 bits into an int).
3) They marshal data into ints in a way that is correct for C code, but wrong for Java code (missing the 0xFF mask)
4) They forget previous instructions and keep re-introducing bugs.
5) ChatGPT kept thinking that version only had information in bits 6-8.
If you know _anything_ about what you're trying to do, it's a waste of time to use a LLM. If the code has anything interesting, a LLM will write broken code.
On the plus side, Bard was quoting sources, complete with license. That led to my first pull request to hopefully fix Bard's upstream source, and let me understand what (byte[0]&0x77) << 16 does, why it works and why byte[0]<<16 doesn't. But then, the LLMs didn't help with that, my need to fix it did.
http://blog.jason.pollock.ca/2023/04/trying-to-write-code-wi...
For example, I needed a GitHub action that would automatically check out build and deploy a node project to a VPS. I also needed it to pipe out the secrets into a local environment variable (.env) and copy that over to the VPS as well.
GPT4 got it in one try. Its this kind of DevOps boring mundane crap that GPT excels at eliminating. Even though I had very little familiarity with the syntax for GH action yaml, I was able to read it and reason out that it was correct.
Could I have sifted through the documentation, or waited through 50 stack overflow pages? Sure, but this was a hell of a lot more convenient.
I've been trying the same prompt since ChatGPT was released, and I keep revisiting it.
They are getting better at writing code, Bard's addition of sourcing information is very important to me.
I do this so I can see if it will help _me_ in what I'm trying to do, and where the failure modes are. Their failure modes are very sneaky and very dangerous for the kind of programming I do.
I would use them to maybe give me a first pass at some code in a language I have zero familiarity with, and only to solve general problems.
I'm thinking throw-away, one-off code where the operational range is known up front and fully middle of the road. I'm thinking reports.
I might try to get it to write something useful in R, I have no knowledge of that.
I can see them as really handy refactoring tools. They take instruction pretty well. They just need to be stable in the code they produce - they keep reintroducing bugs you've instructed them to fix. Changing to an edit stream against code instead of re-generating the code might be a fix?
Still has the "won't fit in an int" problem, but it did use byte & 0x77 to get them into ints.
Amazingly, it got the field correct and in the right order!
It did not reference a source, but that seems to be a Bard feature right now.
It needed:
1) write a function in java to decode a rfc 3588 diameter header
2) 32 bit values will not fit into a java int, please fix the code.
3) 24 bit values will not fit into a java int, please fix the code.
I think I'm going to have to come up with a more complicated problem!
Pretty impressively, this is what it produced:
public class DiameterHeader {
private int version;
private long messageLength;
private int commandFlags;
private long commandCode;
private long applicationId;
private long hopByHopId;
private long endToEndId;
public DiameterHeader(byte[] data) {
if (data == null || data.length != 20) {
throw new IllegalArgumentException("Invalid Diameter header data");
}
version = data[0] & 0xFF;
messageLength = ((long) data[1] & 0xFF) << 16 | ((long) data[2] & 0xFF) << 8 | ((long) data[3] & 0xFF);
commandFlags = data[4] & 0xFF;
commandCode = ((long) data[5] & 0xFF) << 16 | ((long) data[6] & 0xFF) << 8 | ((long) data[7] & 0xFF);
applicationId = (((long) data[8] & 0xFF) << 24) | (((long) data[9] & 0xFF) << 16) | (((long) data[10] & 0xFF) << 8) | ((long) data[11] & 0xFF);
hopByHopId = (((long) data[12] & 0xFF) << 24) | (((long) data[13] & 0xFF) << 16) | (((long) data[14] & 0xFF) << 8) | ((long) data[15] & 0xFF);
endToEndId = (((long) data[16] & 0xFF) << 24) | (((long) data[17] & 0xFF) << 16) | (((long) data[18] & 0xFF) << 8) | ((long) data[19] & 0xFF);
}
// Getters for the fields
}Bard links to a github repo with license information.
What phind does is pulls code from search results and modifies it to suit your needs.
Oddly it went off on a tangent about how the problem can't be solved in Python, which I never mentioned.
But when I explicitly told it I wanted c++ it gave a decent answer.
It did, however, get some nuances seriously wrong.
> Precision loss for large integers: When encoding to JSON, UInt64 and Int64 are encoded as strings of base-10 digits in ASCII, since JavaScript's numeric type only affords 53 bits of precision. This means that large integers might lose precision when converted to JSON.
https://www.phind.com/search?cache=319c7cf5-65e4-4704-a8f3-0...
Is it any better? (Your prompt is related to a completely different domain than the one I work in.)
The code has the standard "you can't fit 24 bit or 32 bit unsigned values in an int" bug. I get the feeling this is a common bug. Probably insufficient testing in the hobby implementations used as training data.
My reading of "int length = first4Bytes & 0xFFFFFF" is that it will truncate values incorrectly. I'm pretty sure that "-1 & 0xFFFFFFF != 0xFFFFFFF". It goes (-1 & 0xFFFFFFF) is a long, which is then downcast back to int, truncating it?
commandCode is wrong, GPT-4 is using 28 bits instead of 24.
I _think_ the command code flag decoding is wrong, but it's been a while since I've done bit fiddling. The spec says Request is bit 0, but the code has it in 0x80 - bit 8?
That's much improved from ChatGPT and Bard. The use of the ByteBuffer avoids many of the problems. I've seen that in some GitHub Diameter decoders.
The RFC is here: https://www.ietf.org/rfc/rfc3588.txt - Page 31 for the header.
>All told, this took about 2.5 hours. >And all I did was end up with a model very similar to one I wrote by hand several months ago. >But that one took me something like six hours! So 2.5 hours is a huge speedup.
It's likely that the attempt being the second time explains 80% of the time reduction. Design and implementation is a learning process, and you being it through once will speed it up massively.
Maybe, who knows? The author cannot go back and do the task for a second time (only for a third time and beyond). So it's difficult to say how long it would have taken them. Either way, gaining half an hour to an hour of productivity, while great, is certainly well outside the hype that we hear about GPT. Indeed, it's not even consistent with the headline, "GPT4 writes Stan code so I don't have to".
I'm pushing against what I see as unwarranted hype. We're hearing breathless reports by both clueless journalists and by seasoned professionals who are using these tools from the perspective of someone with years of experience, who are then extrapolating that experience out across the spectrum of all possible users. My point is that, if the Stan code example required so much hand-holding, then it's not really helping anyone who hasn't already done it before.
This is not an anomaly, indeed, when you ask GPT4 to produce code for non-trivial examples, you need to already have the expertise in that domain to evaluate whether that output is correct, efficient, and error-free.
He gained more than that. what he was saying was that no experience might have made the ordeal 3 or 3 and a half hours instead, still a good deal less than the time he originally spent.
>I'm pushing against what I see as unwarranted hype.
You are hardly equipped to say what is and isn't unwarranted hype for anybody but yourself. It's rather obnoxious to try and impose whatever you think on people giving their personal experiences. No one asked you to pushback on anything so don't. If it's not as useful as you say, then it'll become readily apparent.
>We're hearing breathless reports by both clueless journalists and by seasoned professionals who are using these tools from the perspective of someone with years of experience
There are many giving similar reports without any/little experience at all.
>you ask GPT4 to produce code for non-trivial examples, you need to already have the expertise in that domain to evaluate whether that output is correct, efficient, and error-free.
No you don't. Code is code. It runs as intended or it doesn't and it's extremely easy to see if it runs as intended.
It's more useful if you have experience certainly but that's it.
>I have had to be very attentive because it sometimes produces code that runs but does not do what I asked
It's like someone putting sand on your toothbrush and in your shoes. The subtle way it messes up are in many ways worse and require more attention than dumber tools. It's even worse than humans because the ways in which it fabricates nonsense aren't the kind of mistakes a human makes, it's not like you're coding with an intern but like you're coding with ALF.
The constant context switching between "it's nice" and "it's sabotaging random things" is so taxing if you don't want to mess up. It's like self-driving that way. An automated system that works 95% is worse than one that works 20% of the time.
I also have used ChatGPT for coding and it hallucinated APIs, twice now, that sent me researching down a rabbit hole.
When I have asked it to do things I know how to do and how to evaluate, it has sometimes given very good answers. Other times it misunderstood the problem in a subtle way— and if I hadn’t known enough I would have accepted a program that performed the wrong analysis.
If you have to accept the quality of the solution on faith you are playing russian roullette with your reputation.
I asked a student to analyze the randomness of set of data. He got a program from ChatGPT without knowing anything about how statistical testing for randomness works. He thus accepted a program from the bot that was hilariously unreliable.
in many cases DSLs do the reasoning that LLMs cannot
(though in this case the author works at the place where stan was invented and is probably doing really good prompting)
(also to the extent LLMs let us solve everything with boilerplate generation, that's horrible and I weep for the youth)
At the end of the day you need to know coding to use ChatGPT to help you code, it’s a loop
What kind of data structure permits it to keep that code in particular to be reviewed? Is that part of the model? Or is this phrase a hallucination?
This is why you see it start getting wierd after a bit, it loses the history and you've basically started a new different conversation.
However
Is anyone working on something like this? I would be surprised if Wolfram wasn't working on it.
[1] I don't know what is used behind the curtains [2] worse because I much prefer to read a text than to watch a video