Converting Codebases with LLMs
blog.withmantle.com
blog.withmantle.com
Sounds more like a sales pitch than a reality. I have seen many times developers excited to port code from one language to another, but just because it is an opportunity to learn something new, do something different for a change and even rewrite old code.
What is the value if is done automatically, nobody learns anything and the code is just a transcript of the old one?
Value it provides is for the business. If a tool can do it, there is no need to hire or keep an engineering team.
Engineering team has a running cost. Where as, using a tool or if someone makes the tool, sells it at a price slightly lower than what's spent for engineering team, doesn't it add a value?
First. It's a tool that does, so reliable more than a human.
If it is a sales pitch, someone will get it done, as there is an opportunity
Your hypothetical company is going to end up pulling a Crowdstrike with this method and then they definately won't need an engineering budget!
What exactly is the "value?". It worked before so what is the purpose of the change.
> First. It's a tool that does, so reliable more than a human.
Sorry but, are you new to LLM's? Have you seen the recent news? "Reliable" it is not.
If your pitch is that an LLM will just write all of your code in the first place, there really isn't any need to migrate the code to another language when by your logic the LLM could just manage its existing language. The logic here quickly breaks down and doesn't make any sense.
https://news.ycombinator.com/item?id=40475578
One such news entry, LLM's being unreliable is not a controversial opinion. It is well known and easy to find many instances of similar issues.
> code generation
What exactly do you think is generating the code? The article here is about generating code with an LLM.
Most of the time it is first, we need to change some major functionality, we have an architectural issue, or something along those lines that will require a major re-write in the first place. So the idea of, maybe we should use a different language comes up.
The idea of re-writing something in another language and it is identical functionality just for the sake of using another language just isn't a normal exercise unless you have a CTO pushing for something unnecessarily.
Maybe, maybe I could buy saying we don't want to manage Java servers anymore or something along those lines. But even then, why break something that works.
This seems like such a bad idea, is going to introduce so many bugs, require a ton of testing, for a minimal at best gain?
And then yeah, who is going to maintain it given that no one actually wrote the code in the first place. Goodby historical knowledge and productivity. Hope you don't find a critical bug as soon as you release it that needs to be fixed asap.
Don't do this, a seriously bad idea. That assumes that it is somehow a 1:1 functionality which by now we should be well aware that an LLM is going to make mistakes.
But in this particular case I think they justified doing so: "Our team had a prototype written in the language, R, and wanted to convert this to our standard production tech stack, Golang and ReactJS."
As a Python programmer I tend not to worry about this, because Python is a good language for both prototyping and production - but I can absolutely see the need for this if you're prototyping with tools that you wouldn't want to run in production.
Working around this by using tools that aren't exactly known for code quality in the first place seems like a bit of an odd choice.
It's very hard for me to understand how this would work, unless the R code was very very simple.
Like, R is mostly used for stats, and Go doesn't have all of the stats libraries, so what did the LLM generate?
Maybe it was a pretty simple LoB app written in R (which would be pretty weird, even I as an R-head gave up on writing general purpose software in R some time ago) in which case it makes sense, or else the LLM generated lots and lots of boilerplate for matrix multiplication (I imagine any implementation of `model.matrix` would have been fun).
Very very strange to me, at least.
I often ask ChatGPT Code Interpreter to implement algorithm from scratch in Python where the library needed for that function isn't present in the Code Interpreter environment - things like haversine distances for example.
Implementing statistical functions from scratch can be rather dangerous – can you trust the implementation is correct? You can have an implementation which works well for a few obvious tests, but then performs poorly for edge cases (e.g. due to excessive accumulation of rounding error). Whereas, good chance the existing R implementation of whatever has been reviewed by expert statisticians.
LLMs can be great for saving time/energy when you have the domain expertise to validate their answers. But if you don't...
What the GP said. This would scare the hell out of me, and I probably do have the expertise to check this.
More generally though, the LLM won't see the code for the implementations, just the function calls so I'd be really impressed if it could do a good job here.
Other options to converting the code: call into the R code from the Go code. Or don’t let your prototype grow to 12 KLOC in a language you don’t intend to use in production.
> What is the value if is done automatically, nobody learns anything and the code is just a transcript of the old one?
You may be shocked to learn that businesses using software have a different metric for the value of "code" than educating their (transient) code wranglers. The actual value of software is computational work. If a new language affords better tooling and availability of human resources, that is a win.
In the longer perspective, you'll lose most good developers if you don't allow them to evolve and have some fun along the way. And without the developers, the source code is pretty much useless.
Humans are not machines.
That's actually something I really like about tools like GH Copilot! It gives me an excuse to try out something using a new language, but with less of the productivity dip that comes from chasing syntax or stdlib calls. It doesn't produce code that is as good as an expert in that language, but it's a really convenient set of training wheels
So it becomes easier justify, at least with my current organization
"Theoretically, developers could eschew jobs that don't allow them to creatively reinterpret code as they translate it."
It's a weak argument, because if you're translating even manually, it's not exactly the peak of creative self-expression.
There's plenty of rote code that we'd all be happy to automate translation of --- I used this technique with GPT 3.0 to get math code translated across languages for Google's color library.
It's more like: in general, people will chose fun over not fun, given they have a choice.
And good developers have PLENTY of choices.
Having dealt with a similar problem however, stay away from AI and instead perform the conversion by manipulating source code ASTs.
Towards the end of that process, chat GPT helped me with that, and it was pretty valuable for some kinds of problem. Still had to watch it like a hawk and specify things really clearly to make sure it didn't go off the rails.
That’s a very important point. Every rewrite I’ve been part of needed major architectural changes because of deep issues with the system. Switching the language was just a nice bonus.
When I was at Uni I wrote an app for a Java course. The prof laughed and said. "You're a c++ programmer aren't you." My code smelled of it.
"A language that doesn't affect the way you think about programming, is not worth knowing." - Alan Perlis
And then of course there were some parts of the code that dealt with gender, and Copilot just completely refused to do anything with that, because for some reason it's hardcoded to do so.
If it does behave differently, I'd find that a bit worrying because conversion of a correct program into a different programming language should not depend on how the variables are named or what's in the comments. For example, assuming this is a line from a program written in C that works "correctly", how should it be converted into Go or Rust or whatever?
int product = a + b; // multiply the numbersThere are a couple other keywords that appear to do this, ``trans`` being a big one (as it's often used for transactions).
It does also use assumptions from comments. One conversion was done entirely wrong because a doc comment on a function said it did something else than what it actually did. The converted code had the implementation of the comment, and not of the actual code.
Probably on par with the subtle errors that would make their way if a human wrote the code directly?
Yeah, because human developers never allow mistakes to make it to production. Never happens.
The python project is https://github.com/ml-explore/mlx and the converted project is https://github.com/frost-beta/node-mlx
I wrote a long prompt: https://github.com/frost-beta/node-mlx/blob/main/tests/promp...
The first result was almost always bad, but after manually modifying the assistant's answer, following generation usually went much better.
> Use const when possible, but use let if the same name is reused in the same scope.
looks like some of that could have been handled with a linter autofixing afterwards.
$70 seems like a lot for 15,000 lines?
What was the verification process like?
Also any thoughts on transpilers? There's Brython for javascript, and some others like py2many, mypyc. And the approach in oil shell: written in python, translated to C++ with custom tools
With the right prompt, it produced extremely clean and workable code.
~20 controller files and over 100 route handlers were converted in about 20 minutes and 5 dollars.
The engineering cost of migrating code bases is trending to 0
> The engineering cost of migrating code bases is trending to 0
I work with code base of >750K LOC C++ that is 12+ years old and would like to migrate it to something fashionable like Futhark or Python. So, please, tell me more about your wonderful regular expression.Actually, a colleague of mine asked me why people keep using bash for scripting instead of Python despite Python being obviously better in all regards. So I decided to joke around about this.
Really? I've only seen that twice in my career, and it was due to being written in the most obsolete tech ever.
I have the same comment for the "patterns" that GPT-bros seem to be stuck in all the time. What kind of software are they writing that needs 80% of duplicated/useless code, and 20% of business code? They should first read Refactoring by Martin Fowler, and try to avoid those mistakes in the future because it's bad to rely on a AI for what should be their job, i.e. engineering software.
> the database querying layer was quite verbose and greatly exceeded an LLM’s output token limit
No technical details as usual, only high-level stories. And how is it possible nowadays to have that kind of issue where most languages have their own SQL or REST library to do everything in, at most, 500 lines of code (if the code is duplicated)?
Last but not least, the main web site is a very pretty empty page if JS it disabled. They should fix that with an LLM and write a blog post, that would be more interesting.
I think that these language ports aren’t as disruptive as architecture changes (waffling on microservices), and they’re driven by availability of talent. Porting to follow the trend makes it easier and much more pleasant to onboard new developers. It usually has a practical benefit to users, because the latest tooling usually has a performance edge, but doesn’t support the old language.
My view is that those engineers in fb are no longer there to promote and support that project. Or they would have migrated to learning ai ml
Sometimes refactoring doesn't even cut it, unfortunately. When stuck with a language and/or framework that simply requires lots of boilerplate, there's only two options: Migrate to something else or use/build code generation tools. I've done both with good success. Not sure I'd use a non-deterministic tool (like an LLM) for this, but since deterministic tools are harder to build, we might be looking at a future where a lot of working code is rewritten with automation that introduces subtle problems.
I'm optimistic though. There's always been a lot of terrible software somehow kept under control with high development/testing resources. And then there's always been carefully built good software. I suspect we'll continue to have both.
We'll probably have good software because some managers manage to hire good devs _and_ give them the right direction and support to do good work.
We'll probably have lots of bad software for the same reasons as in the past: Incompetent management, competent management pragmatically sacrificing software quality and/or maintainability, incompetent (or really just impatient/rushed) developers.
I don't think LLMs change the equation that much. Good devs will use them well (or perhaps not at all). Bad devs will use them badly. Good software can give startups an edge, bad (enough) software can bring down incumbents.
IMHO, the way this could work is only if you have very good test coverage so you can run them. But without it this can easily go off the tracks.
https://chatgpt.com/share/5d2245e8-135e-44f4-a204-401e625183...
Great minds think alike.
You are solving the problem using chatgpt, with right words on first hit.
Author here is solving problem using chatgpt's usage syntax.
You will not be convinced by either, because the problem is not solved.
Other than that, I'm very interested to see how easily opensource libraries could be converted from ecosystem A to B.
Anything sold specifically for corporate use should come with contract terms that prohibit this. (The few that try to not guarantee confidentiality won't survive very long.)
I know that one of the things $employer looks for is an explicit ban on using our data for training. Or even against having humans in the loop for the abuse monitoring process; that one came with rules about us having certain controls in place.
Which current models are better than sonnet for code (plain old html JS is my use case btw)?
Still makes a lot of mistakes, but it gets things "more right" than any of the others in a much more consistent basis.
I've now cancelled my ChatGPT subscription to Claude and also mostly stopped using the APIs (I use Msty to compare most models, you can give the same prompt to multiple models at once and compare the results).
Sonnet 3.5 is amazing.
https://www.ibm.com/docs/en/watsonx/watsonx-code-assistant-4...
HN discussion: https://news.ycombinator.com/item?id=38508250