Maccarone: AI-managed code blocks in Python
github.com
github.com
https://en.wikipedia.org/wiki/Macaronic_language
I finally have the right term to describe the warning signs from 1960's-era mainframes that coined "blinkenlichten".
https://en.wikipedia.org/wiki/Blinkenlights
There should be a German term for this, but "gefälschter deutscher" doesn't quite capture it.
>The strength of your faith in GPT-4.
I got a chuckle out of that
Conceptually, you tell the computer how to validate that the conditions you state are correct with your proof, and it checks that each step follows.
One just needs a language to write those conditions in.
One the one hand you have declarative languages, and on the other imperative languages.
Not sure where I'm going with this, but I think there's something there.
https://arxiv.org/abs/2303.11366
Essentially, AI’s output is fed into a checker, whose output is fed back into the AI for “reflexion”. Then the AI often corrects (leading to noticeable improvement in GPT-4 perf).
FYI, I think my open source tool aider would work out of the box to serve this use case. You would just run:
aider file.py —msg “implement the comments”
Of course aider works with any popular language, not just python. And it can do a lot of other coding tasks. It's like pair programming with an AI.Question: can aider work with Ooba/llama.CPP/meta code llama on local llm? If not yet, are you planning on it?
So many users just can’t use GPT4/Copilot because of corporate policy. But they have Macs with M2.
1. GPT-3.5 is just barely capable of editing code to provide aider's interactive "pair programming" style workflow. None of the other models seem to be as capable as GPT-3.5 yet.
2. Just "hooking up" aider to a new model by connecting to its API is almost certainly not enough to get it working in a useful way. Getting aider working well with GPT-3.5 and GPT-4 was a significant undertaking, involving specific code editing prompts and backends for each model and extensive benchmarking [0]. Officially supporting each new LLM will probably require a similar effort to tailor the prompts and editing backends.
Numerous users have done experiments with numerous models. None of these experiments have yet identified other models that look like they are capable of working well with aider. Claude has been the most promising so far, and the new Code Llama looks very interesting on first glance.
Once we see signs that a particular model is capable of code editing, it would be reasonable for aider to attempt to officially support such a model. Until then, aider will simply maintain experimental support for using alternative models.
There is more information on connecting aider to other models, local models and Azure models in the FAQ [1]. There are also ongoing discussions about LLM integrations in the aider discord [2]:
[0] https://aider.chat/docs/benchmarks.html
[1] https://aider.chat/docs/faq.html#can-i-use-aider-with-other-...
[2] https://discord.com/channels/1131200896827654144/11330607806...
One logical concept that's also been noodling in my brain was to construct a DFA(Deterministic finite automata) from the code seen in all the files and then offer the n-1 tokens to the language model and constrain the nth token's selection from the valid ones. I recall someone did this for things that produce DFAs that are fairly small in size(like JSON) and that essentially produced 100% valid JSON without hallucinations(It could be garbage JSON).
So for example if I had a `class ABC` then typing `abc.` could produce: 1. all the methods on it that were valid and 2. had arguments from the surrounding code informed by the LLM.
I'd like to make something more constrained. Instead of a fully-general programming language, let the LLM configure data-flows between pre-defined modules, field mappings, or presentations.
Then, hopefully, we could let the end-user more directly edit the prompt.
I always get squeamish when I see magic comments
The decorator invokes AI completion only the first time the function is run.
edit: I lost interest before I was able to get arguments to work ¯\_(ツ)_/¯
The potential issue I see here is that comments are valid anywhere while decorators may not be, but the parser is hopefully resilient to that. You could see a multi-phase LLM that uses the interpreter to ensure the code runs / works as expected
It’s completely different from a decorator that generates code at runtime? As in, when the code runs?
what if the output has comments?
in my question, there is no need for the decorator to be handled at runtime, the tool that does the LLM stuff has to parse the Python code to know what to generate. It can just key off of decorators rather than comments, or at least this is my question & hypothesis.
you can alternatively split generated code from human written code with files, keep the mapping in something more structured like a config file
I just normally see a better way to do the same thing a magic comment does, generally speaking. There is typically a better language construct if you limit yourself to that language (most common), and config files offer much more structure with existing tooling (mostly decode in your preferred language)
You should be squeamish about running the code without reading it first, given that you're pair-programming with a bot.
It's funny, the first version of this project[0] let you do exactly that, e.g.,
def main(path: str):
#<<filenames = a list of filenames under path>>
for fn in filenames:
#<<size = size of fn in bytes>>
print(fn, size)
#<<use argparse and call main>>
and then run that program like any other Python script (using import magic): $ python -m examples.file_sizes /etc
…
/etc/wgetrc 4942
/etc/nsswitch.conf 542
/etc/adduser.conf 3028
/etc/ethertypes 1816
But yeah, it never felt exactly practical. :)It could replace template rendering in the long run.
I'd much rather focus on teaching people how to think about tests and what they want to do and, when they're stuck on syntax or patterns, make sure they're thinking enough about the problem so that they know why they take the decisions they do. They can then use a search engine or AI to cover a specific technical gap.
Implenting or writing code is rarely the bottleneck for software development after enough experience so something like this doesn't seem useful for me but it's definitely cool to see people trying to integrate new technologies in different ways.
I prefer to use Claude for code generation if using a newer framework or language (the 2021 cutoff with gpt-4 is unfortunate)
It'll be hell to debug an ever shifting codebase though
Practically, how often does this lead to new errors from the AI managed codeblocks when you update code elsewhere?
What prevents my program from behaving differently after each preprocessing run?
- The strength of your faith in GPT-4.Sigh...
And that thorough code review prevents bugs is, at best, a debatable assertion. See e.g. https://www.microsoft.com/en-us/research/publication/code-re...
It finds _some_ bugs. CI/CD, and a massive investment in automated testing has probably had the largest impact in moving software quality forward. (See e.g. "Accelerate", Forsgren, Humble & Kim)
Code review is an excellent tool to socialize knowledge and train up more junior engineers, but in terms of preventing bugs, it's low-value.
Before we had the ability to just add a patch and let the user download it, the end result needed to be very solid, because once that disk was purchased and taken home, it was static.
Now less attention is paid to these things, because it's just assumed to be tomorrow's problem.
But no real argument with the concern. An LLM will generate bugs, and that may be a reason this kind of thing never makes sense in practice (isn't that an argument against copilot, too, though?).
But yeah, I'm personally not comfortable with Copilot or its ilk, either.