Until LLMS start to get there, we still need to save the source code they produce, and review and verify that it does what it says on the label, and not in a totally stupid way. I think we have a long way to go!
Until LLMS start to get there, we still need to save the source code they produce, and review and verify that it does what it says on the label, and not in a totally stupid way. I think we have a long way to go!
There’s a related issue that gives me deep concern: if LLMs are the new programming languages we don’t even own the compilers. They can be taken from us at any time.
New models come out constantly and over time companies will phase out older ones. These newer models will be better, sure, but their outputs will be different. And who knows what edge cases we’ll run into when being forced to upgrade models?
(and that’s putting aside what an enormous step back it would be to rent a compiler rather than own one for free)
IIUC, same model with same seed and other parameters is not guaranteed to produce the same output.
If anyone is imagining a future where your "source" git repo is just a bunch of highly detailed prompt files and "compilation" just needs an extra LLM code generator, they are signing up for disappointment.
Models are so large that random bit flips make such guarantees impossible with current computing technology:
If the answer is no, then we cannot be sure to use it as a high-level language. The whole purpose of a language is providing useful, concise constructs to avoid something not being specified (undefined behavior).
If we can't guarantee that the behavior of the language is going to be the same, it is no better than prompting someone some requirements and not checking what they are doing until the date of delivery.
[1] https://thinkingmachines.ai/blog/defeating-nondeterminism-in...
Genuine question, but why not set the temperature to 0? I do this for non-code related inference when I want the same response to a prompt each time.
[1] https://thinkingmachines.ai/blog/defeating-nondeterminism-in...
Which is exactly one of the AOT only, no GC, crowds use as example why theirs is better.
They are, actually. A "fresh chat" with an LLM is non-deterministic but also stateless. Of course agentic workflows add memory, possibly RAG etc. but that memory is stored somewhere in plain English; you can just go and look at it. It may not be stateless but the state is fully known.
Perhaps more importantly, how would I quantify such “memory”? In other words, how could I verify that two memory inputs are the same, and how could I formalize the entirety of such inputs with the same outputs?
Without taking anything else into account that the JIT uses on its decision tree?
If the compiler had an issue like LLMs do, the half my builds would be broken, running the same source.
But that’s not the point I’m trying to make here. JIT compilers are vastly more predictable than LLMs. I can take any two JVMs from any two vendors, and over several versions and years, I’m confident that they will produce the same outputs given the same inputs, to a certain degree, where the input is not only code but GC, libraries, etc.
I cannot do the same with two versions of the same LLM offering from a single vendor, that had been released one year apart.
Using low code platforms with AI based automations, like most iPaaS are now doing.
If the agent is able to retrieve the required data from a JSON file, fill an email with the proper subject and body, sending it to another SaaS application, it is one less integration middleware that was required to be written.
For all practical business point of view it is an application.
Give a spec to a designer or developer. Do you get the same result every time?
I’m going to guess no. The results can vary wildly depending on the person.
The code generated by LLMs will still be deterministic. What is different is the product team tools to create that product.
At a high level, does using LLMs to do all or most of the coding ultimately help the business?
I think there are important distinctions there, predictably one of them.