From the paper:
> The COBOL source is passed through an internal deterministic Migrator to produce a generated Java target.
Also, humans are not deterministic either. Give the same COBOL -> Java translation to multiple developers and each will come up with a different solution. Heck, even the same developer will produce a different output for the same task, depending on the day of the week.
Indeed, and that's why so few dare to migrate them, and so many who do fail or blow the budget many times over.
I think that was the implicit point of the comment: don't expect that with AI, suddenly we can convert all those COBOL apps with a single prompt.
And that's why it's usually a stupid idea to nilly-willy migrate large code bases to different languages (also I'm getting really tired of the "but humans aren't either" trope).
Just set the sampling temperature to zero and remove any unintended non-determinism during the parallel computation of the token probability distribution. The problem is solved? Of course, not. Non-determinism has little to do with LLMs' mistakes.
Most compilers are also not deterministic, at least by default, but they don't usually make mistakes. Determinism isn't an important quality here. And if it were, AI can be deterministic, it just isn't normally for much the same reason compilers typically aren't (hint: performance).
I find it bizarre that I keep reading this here. Just one of those things that keeps getting blindly repeated without receiving any thought?
can you get some arbitrary number of 9s for consistency? Yea. You can’t get 100% deterministic though.
Computers are deterministic. When we talk about non-determinism we're just talking about where there are hidden inputs, but the fundamentals of computing means that all hidden inputs can become visible if you want them to be (although possibly at the cost of things like performance), so, yes, you can reach 100% determinism just fine if you wish to. It's just not worth the tradeoffs in most cases. But if a deterministic solution solved a problem here it would be worth it.
However, a deterministic LLM doesn't help here. LLMs don't "make mistakes" because they are typically non-deterministic. They would equally "make mistakes" when deterministic.
My understanding is thats not possible (different from being practical), wonder if you have any literature, research to back up that claim?
1. Local hardware, no networking or HTTP requests 2. No other parallel requests 3. 1000 runs bounded by 1000 tokens
I have worked a little bit in academia and I am not a big fan of the way favorable samples for the hypothesis are kept and unfavorable ones are thrown away. I might be wrong here, however its highly unlikely that the team would have just worked with one query. Chances are that a lot of different prompts with varying number of runs would have been tried to see what works and supports the hypothesis.
Correction: AI is not deterministic, the only realistic low-error solution is not a more complex use of non-deterministic AI, but deterministic transpilation.
The problem is that this results in COBOL-in-Java which runs correctly but it is a nightmare to maintain.
As a side note, I can't fathom why they chose Java. It's easier to teach COBOL to a competent engineer than to rewrite everything in a language that encourages onion architectures. Why not Go, for example?
The tests need to be green, the math needs to math, but the code can look a little bit different here or there.
I don't think its easy at all to educate someone on COBOL. Its carrier suicide if you want to ever do anything else again after that gig. I was writing java for years, switched for 2 years to php for another company and had a few companies complaining about these 2 years...
A good Java architecture is possible, though, but it requires planning, discipline, and restraint; and yet the ever-changing business needs can mar your beautiful castle given enough time. COBOL is simpler and more straightforward, you have to go out of your way to mess it as badly as Java allows.
At least if tescoverage is good, but well... That's something llms can also be used for