Diff Models – A New Way to Edit Code
carper.ai
carper.ai
My idea of enjoyable high-quality programming isn’t to dip a spoon into an ocean of soup made of other people’s random design decisions and bugs accumulated over fifteen years, hoping to get a spoonful without hidden crunchy insect bits.
I know the soup is nutritious and healthy 98% of the time, and eating it saves so much time compared to preparing a filet mignon myself. But it’s still brown sludge.
By averaging, a lot of imperfections get diluted away.
Like in Anna Karenina "happy families are all alike, unhappy ones are each in its own way". The defects are idiosyncratic, the commonalities are good.
[1] https://fstoppers.com/portraits/average-faces-women-around-w...
Multi-sourced accumulated unmaintained amateur software without clear provenance or ownership is more like creating a feature I'll call "insta-legacy": now you're responsible for a bunch of code you didn't write that by definition nobody you have access to understands.
This is absurd.
It's not going to stop people from doing it. The industry is clinically insane.
It allows people who do bad work to do more of it quickly. Before they had to manually shovel garbage into projects but now they have a dumptruck.
You know what? It might be fine. Maybe we're going to have a world of fast food programming where minimum wage coders pump out trash and there's going to be Michelin star programmers where you go to for the real stuff.
If that's the case, we'll have to somehow educate the public on the difference so they don't think it's the same thing. McDonald's and The French Laundry are both successful restaurants. That world is possible in programming as well.
It might already be like that. The cheap rates for shady contracting firms that do trash work are probably already using these things
https://arxiv.org/abs/2206.08896
The idea is: LLMs know how to modify code in semantically useful ways. Evolutionary algorithms are great at search, but don't learn new mutations by themselves. So combine them together to generate new data and retrain the models.
So the old system is: scrape human code and train on it. the new system is: generate code, keep the good parts and retrain. It only costs electricity and is open-ended.
> Maybe we're going to have a world of fast food programming where minimum wage coders pump out trash and there's going to be Michelin star programmers where you go to for the real stuff.
It sounds like your concern isn't that it's going to to a poor job, it's that it's actually going to do a good job and you will no longer be able to differentiate your work.
If what you're delivering is so much more valuable then there is no threat and no concern to be had. I believe the concern is that this actually will solve peoples problems or much cheaper to almost free and as a result people will use it. And use it tons. And that's a legit fear to have, but I don't think it should be wrapped up in calling it's output McDonalds's of code.
The better analogy in my mind is a collection of wine connoisseurs seeing the rise of something like 2 buck chuck and trashing it for not being snobby. Deriding it for not having "grapefruit mouthfeel" or something. When in fact most people just want a easy to drink wine that goes with what they're having or dinner - and none of the the extra.
If this code doesn't help people out then people won't use it, if it does, and the "French Laundry" of code is important only to the chef's working there and not anyone else, then we'll find that out pretty soon.
The real issue is potentially getting "French Laundry" quality (or a step just shy) for an every day meal price .
No it's not about me. I won the startup lottery. This was called "RAD" in the 90s, OOP in the 80s and was the promise of "structured programming" in the 70s. The effort to deprofessionalize software development goes back decades.
I care about the craft and the well-being of my fellow engineers. Requiring less knowledge is a mixed bag. Sometimes it's fine, such as compilers handling your C code, and other times it's a problem, such as Word handling your HTML code. It's best when some tooling sophistication is still exposed.
> If this code doesn't help people out then people won't use it
Incorrect! Human behavior and planning is aspirational and emotional, not rational. Choices are made based on narrative appeal and mistakes take years to unravel.
Think of all the once hot frameworks that you'd be simply crazy not to love that are now unmitigated disasters to maintain and lead to mass abandonment and rewrites. People do this stuff, it's how they decide things. Not everybody, but enough to fuck things up for the rest of us.
Also known as:
• Getting hired at a software company that already exists
• Having co-workers leave the project
• Importing frameworks or libraries you didn't write
(1) and (3) are different. In (3) you have access to the people, documentation, things are versioned and bugs are fixed and there's forums of people using identical software.
In (1) the people are still there, you can open tickets against them, go and talk to them, etc
(2) is correct and that's not a good thing. It's snowflake code where you can't do the other things.
The point is you're producing this worst case scenario that good companies try to avoid at great cost, instantly.
Software on the other hand is a logical environment with clear, logical requirements. It ought to work, not just fall apart randomly.
There is no guarantee that sticking the average of one software into a completely different one will satisfy the logical requirements by any means whatsoever.
I can see us switching from programing languages to some new type of logical construct that AI excels at.
They're called humans.
We had built computers also to overcome the human limitations of fuzziness and imperfection.
Now apparently we think putting these limitations back into computers is a good thing.
The goal of all of this is to get humans out of every single process except at the very highest level [1].
You don't farm and hunt your food. You don't make and hand wash your clothes. Why the heck would a business want a person to turn their requirements into repeatable execution units? That person won't be required for much longer.
This entire career only exists as a stepping stone. It fills a business need that can't currently be done better and cheaper. What we see today is not how things will always be.
[1] (At some point that too might disappear.)
The directionality of this is not set my individuals, but by all of us participating in the market economy. It's happening, and all we can do is prepare and find out how we fit into the new paradigm. The Luddites had to, and so shall we.
Maybe it's time we think if that "AI panacea" that many envision -- namely have all the benefits of thinking humans with none of the drawbacks -- is even possible. I get increasingly skeptical with time.
Mind you, I am one of the people who absolutely would create Skynet if he had the time and resources... but I am just not sure it's even possible for us in this day and age.
Part of the reason i don't do these things is because i cannot make consistent clothing, consistent meals, etc. Which isn't to say that all items/food i buy _is_ consistent, but consistency is a valuable metric behind a ton of things we buy and do.
Consistency also seems to be a thing humans are pretty bad at. At least in the capitalist model where we produce millions of Units of any one thing.
I'm curious what this Diff Models paper does with tensorflow's source code. Can it already suggest improvements?
> Various criteria for intelligence have been proposed (most famously the Turing test) but to date, there is no definition that satisfies everyone.
https://en.wikipedia.org/wiki/Artificial_general_intelligenc...
Also, when you average you don't really kill "defects", but rather outliers. An "outlier" statement within a program is very likely to do something important, e.g. taking care of a corner case, otherwise it wouldn't be there.
(My experience with Copilot and ML-assisted programming has been extremely positive, I would not choose to go without it at this point)
Also, unrelated: Leo Tolstoy had his number of problems with his wife and wrote them into Anna Karenina. In fact, it is quite contrary: most disfunctional families fall into several textbook scenarios, while happy families has their own distinct inner dynamics, just looking the same on the outside.
This "study" and your argument are meaningless.
> Futurists in 1950: Automation will free mankind from meaningless tedium to focus on creative pursuits only humans can master.
> Techbros in 2023: We coded AI to write all your books, music, and TV so you can focus on the meaningless tedium of your cubicle farm.
From this popular tweet by @stealthygeek https://twitter.com/stealthygeek/status/1618997354199400449Are they going to follow the luddites and call for people to storm data centers and smash the GPUs?
Most creative work companies need can likely be substantially automated either now or soon.
Most creative work individuals want will take much longer to automate. Things like movies, where even humans can't robustly figure out which ideas will work, will take more human discretion.
The problem of course is that for the most part, the current creative community (whether paid or unpaid) is much more about execution and craftsmanship than creativity. Ideas are a some a dozen, and having exceptionally great ideas mostly matter for the top 2%. The rest is mostly shining in execution, which AI is rapidly attacking right now
Personally I like the way it has removed much of the need for a professional artist, I much prefer having decorations made (or generated) by myself friends and family; it now feels like buying art was something done to surpass the quality possible from a hobbyist, rather than an actual desire to support artistic professions.
Contrast, automation of driving, where there is a hard cap on skill. After a certain number of hours of driving you are not going to get any better at driving.
The pivot of AI form automating tedious work to automating creative work is simply tragic. I consider it one of the major forks in the road between a future utopia and a future dystopia.
I do not think the tech will get any further than it did for self driving cars. It will still be 80% there with the quintessential last 20% out of reach. But it has the potential to do lasting damage.
Automating tedious jobs runs the risk of sudden large scale unemployment if done abruptly but this can be solved by slowly and deliberately visibly phasing in the tech over a few decades.
Automating creative jobs runs the risk of creating barriers to entry and destroying the pipeline to mastery. The jobs eliminated will be the junior levels everywhere. And with no more juniors coming in eventually you will have no more seniors in any of the fields. And their job will likely not be automated.
Think of the demographic crisis China is in, but this time just in terms of skilled workers.
Also, all of those juniors are paying taxes part of which go to pay for pensions. Will the AIs pay taxes?
My point was that the skill pipeline will be nuked. Amateurs and hobbyists will not have sufficient time and resources to reach the same skill level and eventually there will not be enough amateurs and hobbyists to teach others in a sustainable way and keep the craft going forward. The existence of formal training is important because it provides structure, continuity and certain knowledge is only highlighted in a formal setting. Hobbyists too benefit from networking with professionals.
To move away from art to a different creative field, imagine how the software ecosystem would look if there were only hobbyists and AIs(in the service of corporations creating all of the commercial software). It might be a hobbyist FLOSS utopia but certain knowledge would simply be inaccessible. Say you are a hobbyist and have a tricky question solved by some obscure but commonly thought algorithm, who are you going to ask? The AIs have no reason to spend time on StackOverflow. If an AI can write all software, I see no particular societal need for people to be formally trained in software engineering, yet I think humanity would be worse off by having lost this knowledge.
AIs will be appliances not tools. Like appliances, they will serve their purpose but unlike tools they will not elevate the user in any way. Any skill, knowledge or capability an AI has will be sealed within the black box of AI. This goes against one of the defining features of the human species, the ability to transmit knowledge by encoding it.
Hollywood studios are already using game engines like Unreal to create virtual sets in real-time. I could see some sort of hobbyist pipeline being created that will get a decent starting point made.
Workflow would be like this:
>use GPT to generate scenes and ideas
>GPT then used to create multiple different scripts, human chooses the best for each scene
>AI voice synthesizer used to do the dialogue
>Stable diffusion or equivalent used to create multiple 3D models for characters, finishing touches by human
>models then used by game engine to act out the scene, not sure how much of this can be done via AI
>live action stuff can use deep fake technology with any random person being deep faked to look like the AI generated unique character
I can see some crazy animated/CGI movies being produced much cheaper than traditional Hollywood style used now. We could see indie projects with the look of much bigger budget projects thanks to automation. It will level the playing field somewhat and allow people with better ideas to flourish, rather than just people with connections to get funding from studios.
Or we could see creative work drown in massive numbers of half automated generic garbage. But to be honest, most of the movies today seem to be generic and it is very hard already, to find the gold nuggets.
So yes, there is also great potential, but I am less confident that it will level the field and rather make true artists stay niche.
Moreover, math is NOT hard for AI. Anytime the LLM detects a numerical math problem, it should be smart enough to go to a calculator and enter the numbers and give you the answer. This is not hard to implement and I'm sure someone has already done this.
a big part of me not liking windows and preferring the open source world was also not wanting to use the lowest common denominator quality operating system and applications.
handmade, handcrafted stuff will be always head and shoulders above the rest, software included.
In both cases the problem is the same, and hearkens back to Reflections on Trusting Trust. The total amount of code necessary to implement a useful system is far too large for anyone to fully understand and audit.
But I think it should be clear that "well we had a black box AI make them" is not going to be a satisfying answer for militaries trying to remove hostile powers from their electronics supply chains. No different with software.
A question of volume, surely? It might be OK for snippets, but once you've added a million lines of AI code to your codebase, when are you going to get around to reading it?
Like libraries, you hit a button and add thousands to millions of LOC. If it appears to work, are you going to read it?
What I see is a person who copy and pastes crap around until it works and calls it a day. I think code assistants can and will compete with them.
I love to solve _problems_ and to help people with it, but sometimes I just hate to write code to solve them. I wish my computer could have a clear picture of the solution that is in my mind so I didnt have to write a single line of code, so I could focus on the creative part of the problem solving
But I totally agree with you that it would be a positive outcome if we spent less time writing lines of code and more time using better tools to direct computers in solving problems and (I think just as critically) understanding the dynamics of those solutions. A major facet of my skepticism is that I think progress on that second part seems to be lagging way behind...
I foresee a lot of "we had a team, who have all now left, that used AI to write this system and it's mostly working right except in all these ways, and you need to fix it, good luck!" in all of our futures.
It is going to be incredibly easy for people to be the 100x programmer who always delivers on time and promptly leaves for a higher paying job, leaving the debris in the hands of some poor sod who knows nothing about the code or the decisions that led to it.
What AI will do in this scenario is make the "creative" part much easier and jolly and the maintenance part much more painful and frustrating.
So I'm not sure what exactly AI would worse here?
What I hope is that we'll also figure out ways to get AIs to help us just as much with the debugging and verification part as well. But I think it is current a bit ominous that I don't see a breathless article per week on the improving-software side, like I do on the writing-code side.
But yeah, if AI can make us 1000x faster at writing code and 1000x better at fixing it when it isn't working right, then that will be awesome! I'm just a bit skeptical that's where we're headed at the moment.
Today a tech job in whatever tech, will at least use the common tools of that tech.
AIs allow you to eliminate as much of the supply chain as possible and do it in house. And eliminating as much of the supply chain as possible will be done for very good reasons: efficiency, flexibility, supply chain attacks, etc.
Imagine a world in which every client has their own different tech stack. A world in which there are no longer Java, Python, .net, JS, etc. jobs. Instead there are only company specific DSL jobs.
[0]: https://www.semafor.com/article/01/27/2023/openai-has-hired-...
The six figure salaries won't last another decade. For some of us, maybe, but certainly not most of us.
Learn AI now.
Good luck, everyone.
It really doesn't take that much time to teach the CS fundamentals needed to help you figure out where to put the generated code. Or even to know how to prompt it.
Our careers as high earners are collectively doomed. I really didn't expect LLMs to get this good until 2025 minimum. Pretty sad that "learn to code" is gonna be dead. It was the only place where the American dream was still alive.
Right now we don't punch cards or write asm; we write in higher-level languages with lots of existing libraries and autocomplete suggestions. These current AIs are just moving our work up another level. Instead of writing the function with the for loop directly, which turns into the appropriate machine code, we write a natural-language-ish instruction that turns into the function with the for loop.
As the coding help becomes more sophisticated, we'll just do more design and architect-ing and less typing individual lines of C# or whatever. I suspect there will be fewer "programming" jobs available eventually, but they will be just as important to business, if not more so.
------
*If the AGI can even be convinced to spend its time making chat apps for dimwitted meatbags...
But I don't buy the "joy of writing code" argument. Coding is all about making a computer work for you, and I think that taming AIs to be more efficient without letting it introduce random crap will become both important and enjoyable. I think the techniques we have now are too crude for that, but it will improve. Keep in mind that even if you are writing C, you are already at high level, using libraries and compilers other people wrote, bugs included.
Now there is a certain charm being close to "hands on" programming, but if that's the case, go get an Amiga and make a few demos. It won't pay the bills, but it can be fun.
That is an absolutely valid point of view. But it doesn't apply to everyone. Programming is something that can take me into the zone like nothing else. And it has the added side effect of making me think more precisely about higher level problems as well. It's one of those excercises that help me stop fooling myself (in the Feynman sense).
Traditionally you'd put a list of recipe ratios into a spreadsheet and then calculate what you need. But there's a mod called Helm that can handle all those calculations for you, so you just need to specify what logistics components you'll be using.
His immediate comment was "that takes all the fun out of it", to which I responded "it just moves the fun elsewhere."
In this case, the programmer still provides intent and strategy for the bot. We know roughly how efficient the final algorithm should be, and roughly what the data model might be-- being able to get 90% of the code written in 30s-1min should free up more time to think about the system as a whole, I think. (Though this point has been beaten nearly to death, now that I ruminate on it.)
At least in my own experience with Copilot it's been very convenient not having to worry about the finer details. More "should I model the problem this way?" and less "and now I bring in the inner for loop, then I overwrite the first part of the buffer up to index j...".
As a human programmer, is this not what your own brain looks like? What are you doing to the information you take in that allows you to avoid regurgitating the "crunchy insect bits" of your own training corpus?
Maybe something like GPT, style transfer, and OpenAPI combined.
That's not to say plumbing doesn't take skill, it certainly does, but the point of it is that nobody except the next plumber cares how the pipes are laid out, so long as it works and works well. It's when one blows and you have to fix it, or install a second bathroom, that shit really tends to come out. If I'm the one that has to do the work, I can only hope the previous plumber had some idea of what they were doing, and didn't just leave it entirely to automation.
But if they did I'd hope they trust, but verify.
depends what you are working on, I suppose the average web developer could probably be considered this. But there are people working on problems that require major CS knowledge, domain expertise, etc. Those types would definitely be closer to engineers building the tools that the plumber uses
Programmers when the plumbing works: But I don’t want to just stitch components together! This is boring.
There's a reason they call Europeans "europoors". There are advantages to identifying yourself by your work.
What I meant by that is that there are no source images that have previously existed, and especially not in these animated latent space forms. And no, I’m not the only person to be using the tools like this and I never claimed to be.
https://williamcotton.com/articles/the-making-of-distant-des...
What I think I’m doing with this is revisiting some live audio visual stuff I was working with over a decade ago, but instead of having to hand draw as I did in this video:
Around 2:20 has some landscape animations. There’s some really low quality temp imagery in that video as well because it take a lot of time to draw stuff! I basically stopped doing this kind of stuff because it was already too much work to do on top of writing and performing music let alone all the nonsense related to getting shows and managing a band…
I don’t have time to make these base visuals, animate them, program, etc.
Also, I want to try to map audio inputs like amplitude to vectors in the latent space… snare hit increases (lightning:1.0) vector, kick the (ocean:1.0) vector, etc.
Or maybe I don’t do anything other than explore and get inspired.
Neil Young’s song Unknown Legend:
Somewhere on a desert highway, she rides a Harley-Davidson Her long blonde hair flying in the wind She's been running half her life, the chrome and steel she rides Colliding with the very air she breathes The air she breathes
Pedro the Lion’s Leaving the Valley
Long desert highways Where the wheel stops, no one knows My sister breathing This song playing on the radio
It’s a pretty common theme in American music… Wild West, cowboys, Texas, truckers, Harleys etc.
But really this is annoying as fuck for one simple reason: You're talking to me like I claimed I'm some avant-garde artistic genius when all I was saying was that whatever the fuck I'm doing with Stable Diffusion has nothing to do with how any of the plaintiff's have ever drawn any images. It is original work. Just like that song that references desert highways. It might suck, it might be cliched, whatever, but it's still original work.
I write songs that suck all the time. I'd say somewhere less than 1% are any good!
Prepare for the same thing with electronics which you didn't consider as containing much software before - central heating units, AC units, fridges, stoves, light switches, LED light bulbs, vacuum cleaners, electric shavers, electric toothbrushes, kids toys, microwave ovens, really anything which consumes electricity.
Prepare for the support of the vendors of those appliances not taking phone calls anymore, only text communication.
Prepare for the support not understanding the random problems you encounter.
Prepare for the answers you get from support being similarly random.
And maybe, with an unknown probability, prepare for your house burning down and nobody can tell you why.
Just learn how to be a plumber.
Car manufacturing has been automated very much and there was still a need for welders and other skilled workers in different fields. If phased in slowly enough, automation of repetitive work does not have such bad repercussions and has happened all throughout history.
But we've had all of history to regulate quality control in many of these fields. All of this regulation worked to slow down adoption of automation. And this is a good thing. Without regulation roads would be full of alpha quality self driving cars (Tesla manages to ignore this). And even when the tech is ready, switching too quickly is bad.
Creative fields are far less regulated and require far longer training and education. The transition to alpha quality 80% good enough AI has the potential to be far more abrupt and to never actually eliminate higher skilled work but to instead destroy the pipeline towards that higher level of skill.
On the other hand, an utility (truck, taxi, etc.) driver, for example, after a certain number of hours of driving will no longer get any better at driving. Repetitive tasks have an upper limit of skill. Contrast for example a lawyer, since we recently had that AI startup, there is no upper boundary for skill because at a high enough level the comparison is fuzzy. And lower stakes cases serve as training for higher stakes cases. Also contrast how road regulation slowed start-ups like Waymo and Cruise (but not Tesla) vs the reason DoNotPay is facing setbacks: not because there is regulation specifying a minimum level of quality of lawyer work but due to receiving threats from State Bar prosecutors.
Think of other examples of jobs we have automated away: textile making, blueprint drawing, etc. After a number of years working the loom or drawing blueprints a worker would no longer get any better at it. Overall humanity is better off having automated those tasks and the transition has been gradual.
Not sure whose skillset is being threatened, 5-year-olds?
https://github.com/giuven95/chatgpt-failures has more failures, some were fixed, laughed a bit at:
me: "write a sentence ending with the letter s"
ChatGPT: "The cat's fur was as soft as a feather."2) These cherry-picked gotchas are the exact responses I'm referring to. Even in its current form, chatGPT is an incredibly useful resource, and if your reaction to it is to smugly point out its flaws, that speaks more to your own mental rigidity than to the limitations of the model. At the very least, "centaur" workflows will replace raw coding, and in the process devalue much of developers' expertise at the margin. That's already underway.
The gotchas point that this tool unable to understand the letter s and more is just that, a tool, a fancy hammer, in no way it is an arm, and even less it is a brain-mind-agent knowing which nail to hammer and how that nail will fit in the larger picture. And as any tool, it comes with its own downsides. Sure, some sweatshops will be replaced by some even more middle managers managing themselves and the increase of the shareholder profit will continue. The completely messed up state of the world is not a technological issue and will not be solved by technology.
It's referenced in the article.
You'd like the other Inverse Scaling Prize winners too.
[1] https://www.lesswrong.com/posts/DARiTSTx5xDLQGrrz/inverse-sc...
"No GMO"!
"Chicken eggs and cow eggs are produced by different animals and have some notable differences. Chicken eggs are much smaller than cow eggs and have a smooth, hard shell. Cow eggs, on the other hand, are much larger and have a thicker, bumpy shell. Additionally, chicken eggs are typically used for human consumption, while cow eggs are not."
then I followed up with how to make an omelette from cow eggs, they answered:
"To make an omelette from cow eggs, you will need the following ingredients:
2-3 cow eggs
Salt and pepper, to taste
1 tablespoon of butter or oil
Instructions: 1. Crack the eggs into a bowl and beat them together with a fork or whisk.
... (general omelette steps)
7. Serve immediately and enjoy!
Note: cow eggs may be larger than chicken eggs, so adjust the amount of eggs you use accordingly."Technically correct, hard to argue. Perhaps the AI label should specify "Programmed by Pattern Matching: No general understanding or common reason involved".
[1] https://www.thecattlesite.com/articles/1031/anatomy-of-the-c...
So, technically, the largest egg, as in laid egg, is the ostrich's (around 6 inches), although, the largest egg in relation to body size belongs to the kiwi (25%) [2]. For more eggs facts press 1.
[1] https://www.nhm.ac.uk/discover/do-sharks-lay-eggs.html
[2] https://savethekiwi.nz/about-kiwi/kiwi-facts/enormous-egg
Yes, we can automate systems, in certain aspects we can even externalize some decision-making: is this a good apple or a bad apple, should the car break or take a left, but when the chips are down, we are rather far from any externalization of reasoning, meta-reasoning, higher-order thinking, and so forth.
[1] https://en.wikipedia.org/wiki/Mars_Climate_Orbiter#Cause_of_...
[2] "One way to see the MCAS problem is that the system took too much control from the pilots, exacerbated by Boeing’s lack of communication about its behavior. But another way, McClellan suggests, is to say that the software relied too much on pilot action, and in that case, the problem is that the MCAS was not designed for triply redundant automatic operation." https://www.theatlantic.com/technology/archive/2019/03/boein...
[3] https://en.wikipedia.org/wiki/Margaret_Hamilton_(software_en...
Even worse, prepare for them to enthusiastically take calls.
Don't worry, someone will plug ChatGPT into a text-to-speech model soon enough, and market it as a way to put the personal touch back into customer support. Maybe they'll even give it a folksy accent.
This is not a tool for generating applications using statistical methods (we have a lot of tools which do that already), but a tool for assisting human persons by taking boring/repetitive tasks from them and letting us focus on the meaning, the goal
I think this will lead to extreme cost cutting measures in choice of the developers which are used.
People who would have previously been totally ineligible to develop software will happily be chosen.
And they won't care about the garbage code they produce as long as it somehow seems to work from the outside.
They'll care about feeding their families in the dire situation they are in, not more.
It's adherence to safety regulations that's stopping your house burning down at the moment and this responsibility will be there regardless of how the code is written.
On HN, many seem to have interesting ideas about what goes on in the world of programming because they read HN articles and posts and thing everyone is adhering to the high standards advocated on here. It's not only enterprises though; plenty of startups (or small companies that are no longer strictly startups but not enterprise either) who are still running the code from the founders from the day 1 MVP. Held together back hacks and misery, deployed from version25_12_22_xmas_bugfix.zip.
The majority of electronics gets produced in foreign nations far far away.
Do you really think they obey the regulations?
If we're talking about Kitchen Aid / Whirlpool / Samsung / LG / etc, they're going to design for certification and have them produced in the foreign nation to those specifications.
If you're getting random things on Amazon or Alibaba, they definitely may not be produced to those regulations, and you may be risking your insurance coverage if one of those is found to be the source of a fire, as I understand it.
That said, if some vendors are illegally selling products that _don't_ meet safety standards, I'd be doubtful of the GP's claim - that the reason they aren't burning down your house is because of the calibre of software devs working on the product.
and most of all, there's reputation. it still works
Like said, this is already the case since somewhere beginning '00 when the outsourcing boom started taking off.
Tools like this probably simply will lower the bar to $2-3/hr 'data entry' 'specialists', who were ignored before for programming work.
I already see people directly around me who normally couldn't really write much of anything (be it natural language or code) with ease or at all who suddenly (since chatgpt saw the light) produce both with success. They could already do that with gpt3 or copilot, but that takes prompting; chatgpt lowers the barrier to entry significantly.
> And they won't care about the garbage code they produce as long as it somehow seems to work from the outside.
It would be a black box for sure; json in, json out. When something is broken, that 'nano service' is just replaced by a new black box nano service that does the same thing but without the reported bug(s).
And you would logically be without a job hence your fear of these tools?
Maybe it’s much more likely that these tools entrench current software developers who did in fact learn the craft before these tools and can successfully use them to make themselves much more productive?
Does the recent memory of bootcampers getting paid as much as industry vets after a year or two have an impact on this psychology of feeling replaceable?
The currently semi-bad low-grade coders will get pushed out and replaced with even WORSE ones.
The worse ones will be the ones responsible for choosing and approving the code which was generated by AI.
So let's say there is a "Q(c)" which measures worst possible quality of code c.
If the person who monitors this things has a bad worst possible quality Q(c), the code will also have a bad worst possible quality Q(c).
If current code has Q=0.1, "AI" has Q=0.3 and "person who monitors" has Q=0.02 end result might still be better. It's not a simple multiplication of coefficients there. Better baseline would pull result higher
I think my intuition is that the average quality of software may well improve (good!) but that when issues arise they will be more obscure and harder to debug and fix, because nobody will know what the system is actually doing.
this is a fun pattern I've seen play in other industries
The 50% success rate is also best out of 3200 completions. For best out of 1 completion, the success rate is in low single digits.
I think the lesson here is that these models bring a lot more value when: 1. you have unit tests, 2. can afford compute/time to let the model try many solutions, 3. have enough isolation to run unverified code.
But yes, the choice of scales for the graph was rather peculiar.
One of the first steps of a misaligned/unhelpful/virus type of a system, attempting to secure its presence would likely be inference/GPU/TPU compute access. And code injection is a vector. There are multiple other vectors.
When designing such systems, please do keep that in mind. Make sure code changes are properly signed and the originating models are traceable.
Same applies to datasets generated by models.
I on the other hand will survive: what sense is an AI to make of such classic messages as David Bowie's excellent "ch-ch-changes!", the five "fix CI maybe???"s in a row, or the eternal "fuck this shit"?
Something I haven't seen explored too much: navigation help. One of the things that takes me the most time when coding is remembering what was the next file / module / function I need to edit and jumping to it.
An autocomplete engine that would suggest jump locations instead of token could help me stay in the flow much longer, with fewer worries about whether I'm introducing subtle bugs because I'm relying on the AI too much.
1. Humans create programming languages which machines can understand. OK.
2. Humans build tools (LSP, treesitter, tags, type checkers and others) to help humans understand code better. OK.
3. Humans build (AI) programs which run on machines so that the computer can understand... computer programs???
Aren't computers supposed to be able to understand code already? Wasn't the concept of "computer code" created so as to have something which the computer could understand? Isn't making a (AI) program to help the computer understand computer programs re-inventing the wheel?
(Of course, I get that I use the terms "understand" and "computer programs" very loosely here!)
Arguably, we would benefit from even higher level abstractions so the LLM can fit more logic in a single prompt/output.
Maybe a future AI could generate machine code that could be "disasssembled" into higher level languages.
Not sure if that would be better.
My concern with AI across all fields are that people won’t gain the fundamental skills necessary for moving the bounds of what’s possible. Certainly, tools like this AI could produce good results. However, the underlying human is still providing the training data. More importantly, humans are producing the trajectory of development.
If humans are no longer capable of pushing the AI systems. Then the AI systems will either cease to improve, or the AI systems will learn to play off each other. In highly complex systems like many programs, I suspect they’ll play off each other and achieve local minimum/maximum locations. Ie because the “game” (program development) can be iterative they’ll constantly improve code. However, because the AI systems don’t interact with all data (particularly real-world data) when a customer shows a sad face at some UI/UX, it won’t completely develop a new feature that matches the desires of the customer.
Where I fear this will leave us is a class of less-skilled engineers and overly optimized AI. Basically, stuck in development.
So you train the LM to:
Input: code+commit output: diff
OpenAI has an 'edit' endpoint but it's 'in beta' and limited to 10-20 requests per minute. They do not acknowledge support requests about this. Azure OpenAI also has this endpoint I think but they ignore me as well.
So for my edits just like everything else I have been relying on text-davinci-003 since it has much more feasible rate limits. I have just been having it output the full new file but maybe this Unified Diff thing is possible to leverage.
Does anyone know, what would be the easiest way to try to run their 6B Diff Models thing against my own prompts for my service? Maybe Hugging Face?
Negative results are interesting in their own right. I’d rather read about why this isn’t better at the 6B parameter level than e see a hand wave that, well, the samples are more diverse and look the 350M model is better.
<NME> diff_model.py
<BEF> import argparse
import torch
import transformers
def main():
argparser = argparse.ArgumentParser()
argparser.add_argument('--checkpoint', default='CarperAI/diff-codegen-2b-v2', choices=['CarperAI/diff-codegen-6b-v2', 'CarperAI/diff-codegen-2b-v2', 'CarperAI/diff-codegen-350m-v2'], help='Model to use')
args = argparser.parse_args()
model = transformers.AutoModelForCausalLM.from_pretrained(args.checkpoint)
tokenizer = transformers.AutoTokenizer.from_pretrained(args.checkpoint)
# Use CUDA.
model = model.cuda()
text = '<NME> diff_model.py\n<BEF> '
text += open("diff_model.py").read()
text += '<MSG> Print a message after loading the model\n<DFF>\n'
completion = model.generate(tokenizer.encode(text, return_tensors="pt").cuda(), max_new_tokens=400)
completion = tokenizer.decode(completion[0], skip_special_tokens=True)
print(completion)
if __name__ == '__main__':
main()
<MSG> Print a message after loading the model
<DFF>
<DFF> @@ -1,4 +1,5 @@
'''
+import sys
import argparse
import torch
@@ -8,6 +9,7 @@ def main():
tokenizer = transformers.AutoTokenizer.from_pretrained(args.checkpoint)
# Use CUDA.
+ print("CUDA:", torch.cuda.is_available())
model = model.cuda()
text = '<NME> diff_model.py\n<BEF> '
@@ -10,3 +12,4 @@ def main():
completion = tokenizer.decode(completion[0], skip_special_tokens=True)
print(completion)
+ sys.exit(0)
\ No newline at end of file
[\ No newline at end of file repeated many times]
It takes about 150 seconds to run on a 3090 Ti when the model is already on disk.I also wonder if it could be useful in creating Coq proofs!
2023: engineers with their own AI model, typing “#fixed bugs” and spending the rest of the day by the pool.
It makes me wonder if it's related to recent protests in other creative fields in response to AI models, or just a weird dislike of openly released model weights?
And both are trained on large corpus of github sources
Is there a way to test it somehow? Public API maybe?
Do people actually use Copilot for that? I just let it work its magic uninstructed. I guess it sometimes uses comments and function/variable names for its suggestions but that's about it. 99% of the time it just looks at my code, the context and neighboring files to predict what I'm trying to do.
//insert a unicode dot between each character in the string, and convert the numbers to subscript
saved me a lot of copy-pasting.
- ok
- fix
- done
- test
- nice
You should fix that.
Or more likely they are working alone.
It's like good comments in software. Half the time, you're doing it for your future self.
Sounds kinda cool, even if trusting it would be a terrible idea.