AI in software engineering at Google: Progress and the path ahead
research.google
research.google
1) Making non-controversial fixes that save time and take cognitive load off the developer. The best example is when code completion is working well.
2) It’s making you smarter and more knowledgeable by virtue of the suggestions it makes. You may discard them but you still learn something new, and having an assistant brainstorm for you enables a different mode of thinking - idea triage - that can be fun, productive, useful and relaxing. People sometimes want completion to do this also, but it’s not well suited to it beyond teaching new language features by example.
The article makes an interesting assertion that AI tools “fail to scale” when the user has to remember to trigger the feature.
So how can AI usefully suggest design-level and conceptual ideas in a way that doesn’t require a user “trigger”? Within the IDE, I’m not sure. The example given of automated comment resolution is interesting and “automatic”, but not likely to be particularly high level in nature. And it also occurs in the “outer flow” of code review. It’s the “inner flow” that’s the most interesting to me because it’s when the real creativity is happening.
I'm just guessing here. But maybe make it part of some other (already natural and learned) trigger made by the user.
I'm thinking part of refactoring. Were your AI is not only looking at the code, but the LSP, recent git changes (both commit and branch name), which code files you've been browsing.
And if you want to make it even more powerful. I guess also part of your browser history will be relevant (even if there is privacy concerns)
https://upload.wikimedia.org/wikipedia/en/d/db/Clippy-letter...
An actual assistant that can preempt what you need and create it before you get there with a 95% success rate will not feel like Clippy.
Lastly, people are more or less aware of these other dynamics at work. Yet, people are people and respond to other social cues. The sheer popularity of something being novel, cute or especially "cool" does move the needle about adoption and implementation. Clippy was a direct marketing response to Apple getting "cool" ratings with innovative GUI elements.
Furthermore, saying "techies really lose perspective" comes across as dismissive and judgmental. It's important to remember that people are drawn to technology for various reasons, and many are deeply concerned about the ethical implications of what they build. Painting an entire group as lacking awareness isn't helpful or accurate.
If we want to have a productive conversation about the impact of technology, we need to avoid generalizations and engage with each other respectfully, even when we disagree.
Honestly think Clippy predated that, came out in 1996 for Office 97, Macs were only on System 7.5 back then while I think you're thinking of the early MacOS X era which was was 6-7 years later.
Most people are using systems to facilitate code review these days (we aren't sitting around boardrooms with printouts these days) so I wonder if there is a way to use the code review data streams combined with diffs to train AI to "do code reviews"?
At this point, asking an LLM to "implement feature X please" is not going to give you great results. however, unless you can type at 600 wpm, an LLM doing extremely trivial boilerplate code completion is a godesend.
They're really good for type / name finding and boiler plate generation. For larger suggestions as your pointed out they're too wrong to be used as is. They can give good ideas, especially if guided with comments, but usually at this stage I just use Phind.
In the past year since GPT-4 came out, I've also found this to be the case. I'm an ML/backend engineer with little experience in frontend development. Yet, I've been able to generate React UIs and Python UIs with GPT-4 in a matter of minutes and simply review the code to understand how it works. I find this to be very useful!
The person Im responding to was gatekeeping. I responded by sarcastically doing the same to an extreme degree. A lot of people Will have agreed with the person i'm responding to. "Oh yeah of course You should understand these things, the things that I already understand", genuinely not realizing that there's no basis for that. When they reae my response they realize what they were doing, and are less feeling embarrassed for their senseless (and pretentious!) gatekeeping.
In fact, in a high-trust system, e.g. a good engineering culture in a tech company, the reviewer will learn even less, because they won't be worried about the author making serious mistakes so they'll make less effort to understand.
I've experienced this from both sides of the transaction for code, scientific papers, and general discussion. There's no shortcut to the level of understanding given by synthesizing the ideas yourself.
Often the major trigger for a rewrite is that the knowledge has mostly left the building.
But then there's the cognitive dissonance; because the we like pretending that the system is the knowledge and thus has economic value in itself, and that people are interchangeable. None of which is true.
Reviewing the solution is limited. What you don’t get are the myriads of other ways that didn’t work out.
Elegant solutions are the result of weeding out dozens of other messy ways.
So what gets perpetuated here then is the Dunnimg Kruger effect.
While it might be speed things up in many normal circumstances, it devalues hard work in the long run. Not good.
One need only look at open source projects to see the massive variance in code quality that is out there.
This is not true. With any complex framework, the author first learns how to use it, then when they build the model they are drawing on their learned knowledge. And when they are experienced, they don't see all the details, they just write them out without thinking about it (chunking). This is essentially what an LLM does, it short-circuits the learning process so you can write more "natural", short thoughts and have the LLM translate them into working code without learning and chunking the irrelevant details associated with the framework.
I would say that whether it is good or not depends on how clunky the framework is. If it is a clunky framework, then using an LLM is very reasonable, like how using IDEs with templating etc. for Java is almost a necessity. If it is a "fluent", natural framework, then maybe an LLM is not necessary, but I would argue no framework is at this level currently and using an LLM is still warranted. Probably the only way to achieve true fluency is to integrate the LLM - there have been various experiments posted here on HN. But absent this level of natural-language-style programming, there will be a mismatch between thoughts and code, and an LLM reduces this mismatch.
Maybe that's not the case in all fields, but it is in my experience at least, including in software. Code I've written I know on a personal level, while I can much more easily forget code I've only reviewed.
I'm starting to feel like the programming community is just mad things are easier to learn now.
"Easier to get code submitted to a code base" is only marginally correlated with learning anything.
The danger seems to be code that is syntactically correct and compiles without errors, but is logically incorrect.
As a noob I copied code from Railscasts or Stack Overflow or docs or IRC without understanding it just to get things working. And then at some point I was doing less and less of it, and then rarely at all.
But what if the code I copied isn't correct?! Didn't the sky fall down? Well, things would break and I would have to figure out why or steal a better solution, and then I could observe the delta between what didn't work vs what worked. And boom, learning happened.
LLMs just speed that cycle up tremendously. The concern trolling over LLMs basically imagines a hypothetical person who can't learn anything and doesn't care. More power to them imo if they can build what they want without understanding it. That's a cracked lazy person we all should fear.
Are you sure that would work for others? And that other approaches might not be more effective?
I’ve learned lots of things from SO. The top voted answers usually provide quite a bit of “why” content which have general utility or pointers to more general content.
Yes, there are insufferable people on there, but there are gatekeepers and self-centered people everywhere.
Because they’re consistent and they follow (the good ones) a clear path to learn what you want to learn. The explanations may not be obvious at first glance and that’s when you may need somone to present to you in another perspective (your teacher) or provide the required foundational knowledge that you may lack. You pair it with some practice or do cross-reference with other books and you can get very far. Also they’re can be pretty dense in terms of information.
> which books, or what YT channels were useful.
I mostly read manuals nowadays, instead of tutorials. But I remember starting with the “Site du Zero” books (a French platform) for C and Python. As tech is moving rapidly, for tutorial like books, it’s important to get the latest one and know which software versions it’s referring to.
Now I keep books like “Programming Clojure”, “The GO Programming Language”, “Write Great Code” series, “Mastering Emacs”, “The TCP/IP Guide”, “Absolute FreeBSD”, “Using SQlite”, etc. They’re mostly for references purposes and deep dive in one subject.
The videos I’m talking about were Courses from Coursera and MIT. Algorithms, Android Programming, Theory of Computation. There are great videos on Youtube, but they’re hidden under all the worthless one.
In our startup we are short on frontend software engineers.
Our project manager started helping with the UI using an IDE (cursor a VS-code fork) with native ChatGPT integration. In the span of six months, they have become very proficient at React.
They had wanted to learn basic frontend coding for multiple years but never managed to pass the initial hurdles.
Initially, they were only accepting suggestions made by ChatGPT and making frequent errors, but over time, they started understanding the code better and actively telling the LLM how to improve/fix it.
Now, I believe they would have the knowledge to build simple functional React frontends without assistance, but the question is why? As a team with an LLM-augmented workflow, we are very productive.
I’m not saying that React is hard to learn. But I believe buying a good book would have them get there quicker.
I'm a senior React Dev and the code is totally fine. No more anti-patterns than juniors who I've worked with who learned without AI. The contrary actually.
Not gonna fault people for learning, think the FUD is more so in the vein of being ignorant while working.
Yea, you dont really need to know how transistors work to code, but you didnt need that for 2 generations. Personally think (and hope), LLM code tools replace google and SO, more so than writing SW itself.
I got my start on a no-code visual editor. Hated it because the 5% of issues it couldn't handle took 80% of my time (with no way to actually do it in many cases). See LLM auto generation as the same, the problems that the tool dosent just solve will be your jobs, and you still need to know things for that.
But I do think that we will have less depth of knowledge of the underlying processes. That's the point of having a machine do it. I expect this, however, to be a good trend: the systems will need to be up to a task before it makes sense to rely on them.
LLMs _can_ be a good tool under the right hands. They certainly have some ways to become a reliable assistant. I suppose in the way of LLMs, they need better training before they can get there.
The difference is that:
1) All of those things are deterministic [1]
2) In all of those cases I can debug at the level of the abstraction.
[1] Meaning: Do exactly what I say. Don't make it up.
In a certain sense I'd say optimizing compilers aren't deterministic:
The same source code can produce different object code, depending on data and algorithms into which a typical programmer has little insight.
As the parent comment suggested, UI elements are a great candidate for this. Often very similar (how many apps have a menu bar, side bar, etc) and full of boilerplate. And at the rate things change on the front-end, it's often a candidate for frequent re-writes, so code quality and health don't need to be as strict.
It'd be nice if every piece of software ever written was done so by wise experts with hand-crafted libraries, but sometimes it's just a job and just needs to be done.
Accessibility, cross browser+platform support, design systems, SEO, consistency and polish, you name it. You are most certainly not getting that from an LLM and most engineers don’t know how or don’t have a good eye for it to catch when the agent has gone astray or suggested a common mistake
But i started to read a lot more code than what i did 10 years ago
When someone is using an LLM they are still the author.
Think about it like someone who is searching through record crates for a good sample. They're going to "review" certain drum breaks and then decide if it should be included in an artwork.
The reviewing that you're alluding to is like a book reviewer who has nothing to do with the finished product.
people use tools to make things. Its okay. Some "hardcore folks" advance the "lower level" tooling, other creative folks build actually useful things for daily life, and mostly these two groups have very little overlap IMO.
I've been finding AI's suggestions -- even when rather wrong -- help me do that initial step faster. Which, I think, jives with their findings here.
Then we extended the energy in our calories reserves with crops/livestock.
Then we extended the length of our memories with writing.
Then we extended the breadth of our thinking with AI?
We extend our perception with remote sensors (video, sound).
We extend our muscles with machines.
...
I would suggest we have flexible RAM. Also, we have an awful lot of it. The analogy breaks down as soon as you look at it too seriously!
In IT we largely deal with compute, persistent storage and non-persistent storage. Roughly speaking: CPU, RAM, HDD. In humans we might be considered to have similar "abilities" but unlike IT there is a mostly a single thing that performs all of those functions - the brain. That organ is both compute and storage.
LLMs can be surprisingly useful but they are a tool. As with all tools they can be abused and no doubt you have spotted all those tech blogs that spout the same old thing and often with subtle failings (hallucinations).
Keep your tools sharp and know how to safely use sharp tools.
RAMs differentiating factor is increasingly just that it can handle a lot of read/write cycles, not it's speed. And that doesn't map to anything in biology
7 GB/s is the low end of the DDR3 performance range; DDR3 is 17 years old. Meanwhile, DDR5 performance ranges from about 33.5 GB/s to about 69 GB/s. RAM latency, even on DDR3, is measured in nanoseconds; NVMe latency is measured in microseconds, making it about three orders of magnitude higher.
I have no idea how I could even integrate AI into my workflow so that it's useful. It's even less reliable than search is for basic research and can't even cite its sources....
This argument held a lot more weight when it was a search engine playing the role of our memory.
Generative AI still has a ways to go for other forms of research, and it will never fully replace the utility of a search engine. They're two different tools for different but overlapping tasks.
Even AI code suggestions seem to be only a minor improvement over basic LSP integration. One major exception is tedious formatting of the text—say you want to copy over a table by hand to a domain value, copilot is really good at recognizing values and situating them appropriately in the parent l-value.
If chatbots could serve as my RAM, surely they'd be able to generate code relevant to the rest of the codebase or at the very least not require deep scrutiny to ensure their RAM matches mine (it most often does not).
What's your process? Can you give an example? So far for me, I found them to be most useful using LLMs as code copilot.
It's sort of like Cunningham's Law with a party of one. Giving me the wrong answer helps me clarify what the correct answer should look like.
Or, perhaps a better way to put it:
It's easier to criticize than to create. It gives me something to criticize and tinker with. Doing so helps me hone in on the solution I want. (Provided its suggestion was at least in the right universe, of course.)
A few weeks ago they enabled auto complete. I disabled it after a day. Reason: most of the suggestions weren't great and it drowned out the traditional auto complete, which I depend on. I just found the whole thing too distracting.
I also have the chat gpt desktop app installed since a few weeks. This adds a key binding (option + space) to be able to ask it questions. I've found myself using it more because of that. Copy some code, option+space, ask it a question and there it goes.
The main issue is that it is a bit ground hog day in the sense that I have to explain it in detail what I want every time. I can't just ask it to generate a unit test (which it does very well). Instead I have to specify that I want it to generate a unit test, use kotlin-test and kotest-assertions, and not use backticks for the function names (doesn't work with kotlin-js). Every time. If I don't it will just default to the wrong things because it doesn't remember your preferences for frameworks, style, habits, etc. and it doesn't look at the whole code base to infer any of that.
Mostly, progress here is going to come in the form of better IDE support and UX. The right key bindings would help. A bigger context so it can just grab your whole git project and be able to suggest appropriate code using all the right idioms, frameworks, etc. would then be possible. Additionally, it would be able to generate files and directories as needed and fill them with all the right stuff. I don't think that's going to take very long. Gpt-4o already has a quite large context and with the progress in OSS models, we might be doing some of this stuff locally pretty soon.
Have you experimented with GPTs much? I'd solve this problem by creating my own private "write a unit test" GPT that has my preferences configured in the system prompt.
In an IDE context a lot of this stuff should be solved in the UI, not by the end user.
https://www.jetbrains.com/help/idea/full-line-code-completio...
I don't know. There's an idealized vision that people have in their heads of what LLMs should be in which they're "undeniably useful". It's unclear if that already exists and is just constrained by the right UX concept, or if it doesn't exist and never will.
Am I missing something? I tell GPT to remember my preferences and it does.
> Continued increase of the fraction of code created with AI assistance via code completion, defined as the number of accepted characters from AI-based suggestions divided by the sum of manually typed characters and accepted characters from AI-based suggestions. Notably, characters from copy-pastes are not included in the denominator.
On a related note, maybe they should measure number of code characters that can be REMOVED by AI rather than inserted!
Boilerplate is often tedious to write and just as often easy to read. Abstraction puts more cognitive load on the developer and sometimes this is not worth the impact on legibility.
The internal dev tooling at Google is quite far ahead of what's available on the market rn.
The pressure, such that it is, is killing funding for the custom extension for IntelliJ that made it possible to use it with the internal repo.
Cider doesn't have the code manipulation featureset that IntelliJ has, but it's making up for that with deeper AI integration.
So this particular thing was a very well established program, path, and methodology by the time AI hype came.
Whether that is good or bad I won't express an opinion, but it might mean you get a different answer to your question for this particular thing.
Yes, there are oodles of people complaining about AI overuse and there is a massive diversity of opinion about these tools being used for coding, testing, LSCs, etc. I've seen every opinion from "this is absolute garbage" to "this is utter magic" and everything in between. I personally think that the AI suggestions in code review are pretty uniformly awful and a lot of people disable that feature. The team that owns the feature tracks metrics on disabling rates. I also have found the AI code completion while actually writing code to be pretty good.
I also think that the "% of characters written by AI" is a pretty bad metric to chase (and I'm stunned it is so high). Plenty of people, including fairly senior people, have expressed concern with this metric. I also know that relevant teams are tracking other stuff like rollback rates to establish metrics around quality.
There is definitely pressure to use AI as much as reasonably possible and I think that at the VP and SVP level it is getting unreasonable, but at the director and below level I've found that people are largely reasonable about where to deploy AI, where to experiment with AI, and where to AI to fuck off.
However, once I do something, I guess the LLM gets the nudge/prompt in the right direction and almost always auto-completes the full thing correctly.
And now...
> Just five years later, in 2024, there is widespread enthusiasm among software engineers about how AI is helping write code.
Which software engineers?
It may or may not be great for job prospects, but as someone who isn't one of you fancy startup/FAANG super programmers it's great for me to be able to ask it design problems, what the 'best' way to do a certain thing is, or tell it "I need a function given x,y, and z that does this and returns this."
There are plenty of instances where it doesn't make sense to use it, but I always have a tab with ChatGPT open when coding now. Always.
> implement also for Days
This fails to recognize that this is a bad feature that the Abseil library would explicitly reject (hence the existence of absl::CivilDay) [0], and instead perpetuates the oversimplification that 1 day is exactly 24 hours (which breaks at least twice every year due to DST).
Which is to say: it'll tell you how to do the thing you ask it to do, but will not tell you that it's a bad idea.
And, of course, that assumes that it even makes the change correctly in the first place (which is nowhere near guaranteed, in my experience). I have seen (and bug-reported!) cases where it incorrectly inverts conditionals, introduces inefficient or outright unsafe code, causes unintended side effects, perpetuates legacy (discouraged) patterns, and more.
It turns out that ML-generated code is only as good as its training data, and a lot of google3 does not adhere to current best practices (in part due to new library developments and adoption of new language versions, but there are also many corners of the codebase with, um, looser standards for code quality).
[0] https://github.com/abseil/abseil-cpp/blob/bde089f/absl/time/...
Notice parent said "fully"
It's like asking "How long until a bulldozer [1] fully replaces a human construction worker?" A bulldozer is not a full replacement for a human construction worker. However, 1 worker with a bulldozer can do the work of 10 workers with shovels. And 10 workers with bulldozers can build things no amount of workers with shovels can.
[1] I'm using bulldozer as shorthand for all automated construction equipment - front end loaders, backhoes, cranes, etc.
It'll be one guy with a bulldozer doing the work of ten men.
New tech drives job changes. Always has, always will. No guarantee that the changes are good for every individual. And there’s no guarantee that they are good in the aggregate, either.
Either it's 1 Software developer full-time equivalent (already happened) or all SWEs (never going to happen.)
I'm guessing it's 30+k, depending on how liberal you are with the job (eg. data engineering, SREs, etc).
> the first engineer has already been replaced
Realistically, many engineers never got hired because of this already. Also they've had a few rounds of layoffs, so probably plenty are fired by now.
Google already was known for their use of automated code-gen long before they invented LLMs so I wouldn't be surprised if they're on the verge of major systems being gen'ed or templated with LLMs.
Just try it....
At some region on the technological "singularity" timeline, there will can and will be fully-autonomous corporations. The question is: With corporate personhood, can a corporation exist, pay (some) taxes, reinvest in itself, and legally function without any human owners? This may also be contingent on whether or not a corporation conduct legal and business activities, i.e., if performed by a human agent or through some automated means, even if they were directed by AI management. IANAL, but I guess an autonomous corporation could be sued, and perhaps even the creators of the software used to create the AI that run it could also be potentially at risk.
What product of Google’s has been improved by this feedback loop? The trajectory of Google search itself in the past year has been steadily downhill. And what other products of theirs I use don’t really change much, at least not in positive ways. Gmail is just gradually being crushed from every side by other app widgets scrounging for attention. Chrome has added… genAI features and more spyware? Great.
This exactly. There is so much boilerplate involved in writing anything inside Google.
AI was great to cut that down a bit. It's still nowhere near what it's like in the outside and/or non-Java world.
Which isn't to say that this isn't progress - just that that stat should be taken with context.
Code is read more than it's written. And it should be written to be read.
If you are barfing out a lot of auto-completed stuff, it's probably not very easy to read.
You have to read code to maintain it, modify it, analyze its performance, handle production incidents, etc.
From my experience using LLMs, I'd guess the opposite. LLMs aren't great at code-golf style, but they're great at the "statistically likely boilerplate". They max out at a few dozen lines at the extreme end, so you won't get much more than class structures or a method at a time, which is plenty for human-in-loop to guide it in the right direction.
I'm guessing the LLM code at Google is nearly indistinguishable from the rest of it for a verbose language with a strong style expectation like java. Google must have millions of lines of Java, and a formatter that already maintains standards. An LLM spitting out helper methods and basic ORM queries will look just like any other engineers code (after tweaking to ensure it compiles).
If you already apply a code-formatter or a style guide in your organization, I'm guessing you'd find that LLM code looks and reads a lot like the rest of your code.
But I am saying it's not going to ever make the code significantly better
In my experience, code naturally gets worse over time as you add features, and make the codebase bigger. You have to apply some effort and ingenuity to keep it manageable.
So if everyone is using LLMs to barf out status quo code, eventually you will end up with something not very readable. It might look OK locally, but the readability you care about is a global concern.
i am still struggling to find value where the main selling points of using llm for code lies. for boilerplate this is a very inefficient although interesting way, but it does not help me do the thinking. well-written, language-specific boilerplate generation in most ide brings out the current effects just fine.
I'm gonna need a _lot_ of popcorn.
Maybe the fact could be leaked? Unlikely though.
IP getting coughed out of an AI will keep some lawyers busy and result in fig leaf changes.
If the infringing code powered a widely used service, replacing the code with a clean room refactor could be arduous and time-consuming and might require a court-appointed monitor to sign off.
People who work at your bank right now are pasting your personal details into language models and asking if you deserve a loan. People will figure out how to get this data back out.
"But the model only uses old training data" there are myriad lawsuits in flight where these companies took information they shouldn't have, in all forms. Prompt engineers have already got the engines to spew things that are legally damning.
A real hack, which might archive user inputs as well as exfiltrate training data. We're only beginning to imagine the nightmare.
Can people use these language models to get private Google info? Only if Google was dumb enough to include that in the training. Hint: Yes, it's much, much worse than anyone imagines.
We nominally have controls against this sort of thing, but if it's only weakly-enforced policy, then unfortunately it's easy enough to bypass.
Apple and Microsoft are gonna smoke these guys.
Kind of reminds me of the mad cow disease problem.
Same thing they’re doing with the general public and search.
Anyone who worked at google during + can see this is the same half hearted type of force feeding that will fail, with the rationalizations to boot.
Disclaimer: I hold google stock but don’t work at google
It does look like there's an auto-installed cider extension, which is fine, the worst case for this stuff is "it's in my autocomplete list" -- that's fine!
Certain aspects can be disabled but you cannot entirely remove all AI affordances. It’s force feeding.
It will eventually all be opt-out. This is obvious to anyone who works at google. Again, same as force feeding of gen AI search
I don’t mind optionality, but it can’t be disabled so it’s annoying.
I don't feel it's obvious at all that this will become forced. What would that even mean – you don't get to edit code anymore but can only interact via chat with an AI which edits the code for you? I think it's obvious that's not going to happen.
Will there be a day when you can't turn off AI suggestions for something? Maybe, but suggestions aren't requirements, no one is forced to use them so at worst they're a minor UI annoyance.
I'm interested to hear why you have such a negative take on it? I don't particularly enjoy most AI implementations in products (not Google specific), but they're almost entirely optional everywhere, and ignoring them really isn't much of a problem.
Not all AI affordances in cider can be disabled. There are plenty of people complaining about this on yaqs and buganizer.
> no one is forced to use them so at worst they're a minor UI annoyance.
This attitude is exactly the problem with Google and is why Google’s AI rollout has been terrible thus far. Luckily the stock is still going up, so whatever I guess.
I agree this may not be how they are used, but honestly, code review is a skill and many people are bad at it. Blindly suggesting things because an AI (or a presubmit) suggests it is bad form regardless of the source. People slam that "please fix" button without looking at what it's asking for, or without reading my detailed explanations of why I'm ignoring a recommendation, etc. This was already a problem and AI isn't changing that.
We can make code review better, but that comes through training reviewers better, not by ruling out AI tools just because they're AI tools.
If a reviewer makes a comment, Critique will create an AI suggestion based on their comment, even if they didn’t explicitly do so (unless they turn off giving AI suggestions on their end, but there’s no way to stop from seeing it from the other).
> not by ruling out AI tools just because they're AI tools.
This was not my point at all.
The point is that Google force feeds AI tools, and makes it difficult to opt out. In many cases it is not possible as shown with GenAI search.
We all have unique circumstances, but I can almost guarantee you that you'll be absurdly happier elsewhere.
Life's short, and no matter how much you save, something can take it away.
Better to start pursuing it now, than after being the sacrificial MI, or after the call from HR asking you to chill because a VP got upset and had lackeys reverse engineer your identity.
Well, that's easy for you to say. People have family to feed, right?
Thanks for your insightful contribution - I read the post again and found this right after what you found, like, right after. I hope this helps clarify
Guarantee with what? Personal money? I know people who have more than 3+ years of experience having trouble with getting an offer after months of searching these days. What can you offer to guarantee the "happiness" "elsewhere"?
Such a ridiculous, out of touch comment.
Optional. I'll bet.
As far as being non-optional: a code owner could very much refuse to approve your change until it conforms to their standards for their code base.
ML code suggestions at Google do not inspire that high level of confidence in me today, but I have no reason for thinking that will always be the case.
Also I would say that automatic formatters weren't popular until the mid 2010s from what I've experienced, despite being technically possible since pretty much the advent of programming. I even remember having to push hard for adoption of them in ~2018. Even if the AI tools were at the level (and they're definitely not yet, any of them), it could easily take a decade for it to become the norm or for it to be mandated.
When Blake Lemoine was talking nonsense about lambda being conscious in June ‘22, these internal models were already killer at producing code. ChatGPT wouldn’t be released for another 5-6 months.
https://theconversation.com/is-googles-lamda-conscious-a-phi...
I actually don't really use copilot as I didn't find it that helpful, so I don't really have the problem anymore, but I could see it was a danger. Bit like driving when tired.
With AI, you delegate, and then you need to review in the next minute while you are still exhausted.
if you're just delegating work you don't have an intern, you have an undercompensated employee
Copilot doesn't give a shit about that.
I don't know how to articulate it, but wherever I hear the financial analysts talking about how much work AI is going to do for us, I just have this spidey-sense that they're severely underestimating the social aspect of why anyone tries to achieve a good outcome
they think they can just spend 100,000$ on GPUs and get 10x the output of someone buying a house and raising kids getting paid a 6 figure salary
Maybe you're the exception!
Libraries are biggest pain point. You don't know what a function really doing unless you have used it before yourself. Docs are not always helpful even when you read them.
Lot of assumptions that may not be totally wrong but not right also. In c++, using [] on a map to access an element is really really dangerous if you haven't read the docs carefully and assumes that it does what you believe it should do.
It's not that bad, as it just inserts a default-constructed element if it's not present. What would you expect it to do, such that it returns a reference of the appropriate type (such that you can write `map[key] = value;`)? Throw an exception? That's what .at() is for. I totally agree that the C++ standard library is full of weird unintuitive behavior, and it's hard to know which methods will throw exceptions vs have side effects vs result in undefined behavior without reading documentation, but map::operator[] is fairly tame.
Meanwhile, operator[] to access an element of a vector will result in a buffer overflow if not appropriately bounds-checked.
The committee decided that since the reference can't be null, therefore, we'll insert something into the map when someone "query" for a key. Perhaps two mistakes can make a right!! It is hard to make peace with it when you are used to a better alternative.
It's funny to see that, since I'm used to folks coming to C++ from C and bemoaning that C is a much more straightforward language which doesn't engage in such under-the-hood tricks.
> Something akin to `std::optional` would have been great.
So, map.find(), which returns an iterator, which you can use to check for "presence" by comparing against map.end()? (Or if you actually want to insert an element in your map, map.insert() or, as of C++17, map.try_emplace().)
Again, the standard library is weird and confusing unless you've spent years staring into the abyss, and there are many ways of doing subtly different things.
C++ does not have distinct operators for map[key] vs map[key] = val. Languages like JS and Python do, which is why they can offer different behavior in those scenarios (in the former case, JS returns a placeholder value; Python will raise an exception). But, that's only really relevant in the context of single-dimensional maps.
Autovivification is rather rare today, but back in the early 90s when the STL first came about (prior to C++'s design-by-committee process, fwiw), what language was around for inspiration? Perl. If you don't have autovivification (or something like it), then map[i][j][k] = val becomes much more painful to write. (Maybe this is intended.) Workarounds in, say, Python are either to use a defaultdict (which has this exact same behavior of inserting elements upon referencing a key that isn't present!) or to use a tuple as your key (which doesn't allow you to readily extract slices of your dict).
If I can't tell based on the code what it's supposed to do, then it's a shitty library or api.
I don't have a strong conscience - I just don't trust it to do anything right.
Just in time to make every software product eat up even more CPU and RAM to do simple things.
It's a short release and I read it twice, so if it was there I feel like I'd have noticed.
> users accept nearly 30% of code suggestions from GitHub Copilot
> Using 30% productivity enhancement, with a projected number of 45 million professional developers in 2030, generative AI developer tools could add productivity gains of an additional 15 million “effective developers” to worldwide capacity by 2030. This could boost global GDP by over $1.5 trillion
They were probably just being disingenuous to drum up hype but if not they'd have to believe that:
1) All lines of code take the same amount of time to produce 2) 100% of a developer's job is writing code
[1]: https://github.blog/2023-06-27-the-economic-impact-of-the-ai...
With that perspective, "characters added by AI" is an ok metric to track.
Per the metrics I added to the integration when I began to trial it, I accepted about 27% of suggestions.
I didn't track how many suggestions I accepted unmodified, because that would have been orders of magnitude more difficult; I would be fascinated to see Google's solution to the same problem documented, but doubt strongly that I will. I'm sure it's entirely sound, though, and like all behavioral science in no sense a matter of projection, conjecture, interpretation, or assumption.
I turned off the Copilot integration months ago, when I realized that the effort of understanding and/or dismissing its constant kibitzing was adding more friction on net to my process than its occasions of usefulness alleviated.
I do still use LLMs as part of my work, but in a "peer consultant" role, via a chat model and a separate terminal. In that role, I find it useful. In my actual editor, conversely, it was far more than anything else a constant, nagging annoyance; the things it suggested that I'd been about to write were trivial enough that the suggestion itself broke my flow, and the things it suggested that I hadn't been about to write were either flagrantly misguided for the context or - much worse! - subtly wrong, in a way that took more time to recognize than simply going ahead and writing it by hand in the first place would have.
I've been programming for 36 years, and it's been well more than two decades since I did any other kind of paying work. The idea that these tools are becoming commonplace, especially among more junior devs without the kind of confidence and discernment such tenure can confer, worries me - both on behalf of the field, and on theirs, because I believe this latest hype bubble ill serves them in a way that will make them much more vulnerable than they should need to be to other, later, attacks by capital on labor in this industry.