TLDR explains what a piece of code does
twitter.com
twitter.com
That is, it's like it generates these kinds of comments:
// initializes the variable x and sets it to 5
let x = 5;
// adds 2 to the variable x and sets that to a new variable y
let y = x + 2;
That is, IMO the whole purpose of comments should be to tell you things that aren't readily apparent just by looking at the code, e.g. "this looks wonky but we had to do it specifically to work around a bug in library X".Perhaps could be useful for people learning to program, but otherwise people should learn how to read code as code, not "translate" it to a verbose English sentence in their head.
To your point about the utility of code comments describing the behavior this way, I agree it’s probably much more valuable for beginners. In fact when I’ve mentored early programmers, I sometimes ask them to write out essentially prose like this in comments before writing a single line of executable code.
Now, I’m far from a beginner. I’ve been considered a senior engineer long enough that friends discourage me from disclosing the amount of time, for fear of age discrimination. I can absolutely see the potential of this tool as part of my IDE. I’m on vacation now, but when I return to work I plan to take it for a spin as an aid for refactoring areas of code which clearly work as intended (well, for the most part) but the actual behavior and intent is much less clear.
Here’s why I think it’ll be valuable for refactoring: it can help limit the amount of mental context switching necessary to build a mental model of what the code does. I often find myself trying to produce prose much like this for my own reference, but I end up losing context as fast as I acquire it as I follow references into their respective rabbit holes. Having the tool do that for me can help me stay in a single area of focus. It could also be a useful reference for adding and improving type definitions, maybe even regression tests.
The best part is that it doesn’t, from what I’ve seen, do anything besides populate ephemeral annotations. It doesn’t try to write code or automate anything other than producing a narrative. Like at least one other commenter, I’m skeptical about the reliability of that. But unlike that commenter, I’m willing to take the risk… probably because I’ve learned to be skeptical of my own reliability performing the same task. If my instinct is right that I can use this tool the way I hope, I’ll still scrutinize it for accuracy. But that’s potentially much better than having only one imperfect, meat-based computer doing the work.
Perhaps there are some other examples that explain more of a "structural understanding" of the code, but I'm skeptical without more evidence.
TBH, I would also expect this tool to be a lot more affordable than a senior developer for language X.
Plus, chances are good that your senior developer for language X will barely scrape by as an extremely junior developer for language Y that you also use in some capacity, and that there is a problem which involves stuff written in both X and Y. Also, many languages have horrid gotchas, where innocuous looking code does something quite unexpected. Like the C++ classic log->debug("Timeout: %d", config["timeout"]). If some tool where to actually tell you "Writes a debug log with the current value of timeout, if it exists, or otherwise sets it to 0", well that would be pretty useful for someone who has only superficial knowledge of C++.
That specific code? Sure. Real world code that’s gotten more convoluted over a decade of maintenance?
I have a dozen pages of notes I’ve taken on the apparent behavior of a single function, all of its call sites, all of the known possible kinds of states it might encounter, all of the categories of implications those states might carry and the kinds of downstream effects its return value might have. It’s not even a particularly large function as far as those go. The notes aren’t even complete, if I had to guess they’re 1/3 there. To be complete they’ll span not just several modules but cross package and language boundaries and even repo migrations. Not because the function itself is that complex (although it’s much more complex than I’d prefer), but because the universe of its inputs and usage is enormous and the history which produced it is long and just as convoluted. And, importantly, because the space of very similar behaviors and functions I’ve discovered has grown each time I peel a layer of the onion off.
Now when I go back to work, given I get to continue untangling this thing, having a tool which helps even explain this universe would be a godsend for actually acting on it in a reasonably safe way. Especially if it does so with consistent language. I could literally shove it into a database and query it, or do all kinds of other analysis. I can’t do that if I’m spending all of my energy trying to just describe the thing in my own words with incomplete understanding. I mean, I can. The documentation I’m describing literally began with me wrapping up the previous day’s notes with “humans built this, a human can understand it, model the damn thing”. But on what timeline? And is it a good use of my time to do it when a machine could probably do it more quickly and hopefully just as reliably at scale?
I’d be happy to share the evidence if my instinct proves out.
Also complex indexing or expressions.
This stuff very rapidly stops being expressible in narrative form because the details can't be "compressed".
The explanation starts to look like legalese and you realise the code itself is the more compact and legible expression.
Ideally an autocommenter will be able to take a hand matmult and write "compute the surface normal" (this is an example where apparently high local detail can map to good conceptual compression) but actually that is the best possible case.
When it comes to important but somewhat arbitrary stuff like shipping or tax rules, english will not speed comprehension and is less good than a tageted dsl or well thought out tables.
The core problem is that "sometimes there is no simple explanation".
I’m well acquainted with its complexity but I need to articulate it coherently to justify changing any of it, not just to myself but to my team and stake holders. I’m not looking for a simple explanation. I’m looking forward to having tools help me navigate a very complicated explanation.
I also don’t understand the motivation here or upthread to splain the problem I’m solving away as some nihilistic unsolvable thing when I’m saying I see real prospects of this helping me. I mean, you’re welcome to whatever nihilism you see fit in your course of action but I personally don’t benefit from being told things I find potentially useful are not useful for me, actually.
I don't think we're there yet (because complicated side effects etc. are a thing) but I would rather have such a solution than some text to read, after all I'd be using such a tool for code I cannot comprehend just by reading it.
i += 2 // Increments i by 1
Now you're screwed. Added bonus - the commit message introducing this change is "Yay! Fixed all the tests!!!!""When code and comments disagree, both are probably wrong", Norm Schryer.
Edit: just thinking, the canonical example would be something like the fast inverse square root from quake. Is it going to summarize or is it going to tell you
i = 0x5f3759df - (i >> 1);
// Shift i right by one and subtract it from 0x5f...
(It would be cool if it does work here, even if it does, when someone makes up a new thing like this, it couldn't possibly comprehend why it's being done)e.g. it would probably "understand" common devops/organization/OS/library stuff even if it's not part of the language it's presumably reading - just because other users probably left comments on those lines similarly - but it's not gonna necessarily understand your application-specific business logic beyond what's actually happening. Would need some very specific examples though (by definition lol). Even dumb stuff like interpreting "make the div spin around" from some CSS/JS would probably work, as someone somewhere probably coded that similarly.
AI generated papers that got published in prestigious journals situation all over again and again. At glance looks amazing but the machine definitely doesn't have any kind of intelligence and the output is actually worthless.
This AI stuff work really well when they do something very specific that is quickly inspectable by human, for example generating interpolated frames in videos or extending a pattern or detecting anomalies kind of stuff.
The moment it strays away from human control it fails amazingly well.
$data = file_get_contents(filename: 'php://input');
to "We're getting the raw request body"
I still argue that developers should be able to read the raw code like that faster than turning it into an English sentence. I'm not saying the text generation isn't extremely impressive, I just don't think it's that useful. $data = file_get_contents(filename: 'php://input');
I would definitely have to google what this means, like why is it opening a file and how is 'php://input' a file? "We're getting the raw request body"
Is actually a good explanation.So, someone new to a project, and maybe unfamiliar with the tech, might find this useful.
Even for experienced developers, it is useful when navigating a new codebase (where you don't know what each API call does) or even a new programming language (where a particular construct may be unfamiliar like Rust's if let). Well supposing the tool is accurate (incorrect results may make the tool useless or harmful)
But if you're discounting the usefulness for beginners I think this is a mistake. The experienced devs of tomorrow are the beginners of today.
I have been “learning to program” for 20+ years and would absolutely find this useful as a quick way to get basic information about a chunk of code I’m unfamiliar with.
Not that learning to read code isn’t important, just not always necessarily worth the time (:
1. Write comment first, code later.
2. Tell the intent of code, not the instructions to achieve it
I do wonder if ML could be applied if you actually also train it on your particular (large) application where if might be able to 'cross-reference' domain knowledge present in commented code to other similar code. (But then you might want to just remove the pseudo-duplication anyway.)
"But the code shouldn't need comments," he'd complain, "it should be obvious what it does otherwise it's just bad code!"
Yes, you dobber, it *is* obvious from the code, the *how* is obvious, but the why and the what might not be. The comments explain what it's doing stuff to and why, and in particular why you'd want that and why it's important. Disk space is free, so in a big long comment just write up what that particular bit of business logic is intended to achieve and what it expects as inputs and outputs.
or, TL;DR "DOCSTRINGS BITCH DO YOU SPEAK IT"
There were horror stories with that codebase, like index.php was set as the default error handler, started by immediately returning 200 OK before going anywhere near URL dispatching, and then displaying a page that said 404 despite returning 200 if it couldn't find the URL.
But otherwise, I also want the Why not the What!
It knows how to create sentences that sound like something a human would write, and it's even good at understanding context. But that's it, it has no actual intelligence, it doesn't actually understand the code, and most importantly, it's not able to say "Sorry chief, I don't actually know what this doing, look it up yourself."
I can easily imagine the reason you're looking at an unfamiliar function in an unfamiliar language (hence needing such a line-by-line translation) is that there's some sort of bug and that edge case is exactly why. The tool would mislead you into thinking it's one of the other lines, because of how simple its translation is.
This is also the story of the past decade of "progress" in machine translation between natural languages.
You trust a human translator because their ability has been partially verified by an authority (language certificate) and many people have employed their services with few complaints. Similarly, you trust a translation program because it was partially verified (hidden test cases) and many people use it with an acceptable amount of complaints (given its convenience).
This is a prime example of the moving goalpost of what intelligence "actually" is - in previous eras, we would undoubtedly consider understanding context, putting together syntactically correct sentences and extracting the essence from texts as "intelligent"
The goalpost was never in that spot.
It reminds me of people who use "big words" without actually understanding them. If they don't overdo it or really miss the meaning of a term, they can seem much more educated than they are.
You're asserting here that it understands context, but you haven't provided any argument in support of that assertion.
I think you'll also need to define what you mean by "understanding" (because that term is loaded with anthropocentric connotations) and clearly state what "context" you think the model has.
I don't want to sound negative of course, and I expect many of these apps coming up, until Codex stops being free (if they put it on the same pricing as text DaVinci model, which Codex is a fine-tuned version of, it will cost a ~cent per query). I'm just wondering how come the information about this type of app reaches most people way before the information about "the existence of Codex" reaches them.
For all the publicity around Codex recently (and especially on HN), it still seems like the general IT audience is completely unaware of the (IMHO) most important thing going on in the field.
And to anyone saying "all these examples are cherrypicked, Codex is stupid", I urge you to try Copilot and try to look at its output with the ~2019 perspective. I find it hard to beileve that anything but amazement is a proper reaction. And still, more people are aware of the recent BTC price, than this.
Source: have been playing with Codex API for better part of every day for the last few weeks. Built an app that generates SQL for a custom schema, and have been using it in my daily work to boost ma productivity as a data scientis/engineer/analyst a lot.
The lack of control over it just makes it annoying. In many ways it's faster to just type out the algorithm than it is to lay the algorithm out and spend the time trying to understand what's there so I can successfully convert the code to what I need.
Then there's the lack of stability. Yesterday it did something different from what it's doing today, so I can't even use muscle memory to interact with it anymore.
Intellisense has _always_ had that annoyance factor of getting in your way sometimes, forcing you to write code in a certain way to minimize that. All this just makes it more annoying and I don't believe anyone who claims it truly makes them more productive.
I think you should be careful to realize that though it may not fit for you intellisense is very helpful for a lot of people and that it may be your tastes as to what you find annoying that do not generalize. I for one don’t even notice the things you’re saying bother you because the mental overhead to me is very little. Just quick glance, tab to auto complete if it’s useful otherwise keep typing.
Think of it like parenting.
I can either convince the little tike to clean up his room or do it myself in half the time. Only with babysitting I may take the time to teach them something, with things like autopilot I have no such motivation.
I don't even like autocompleting brackets and the like. It flat isn't consistent enough for me. It's ok-ish when writing fresh code, but completely gets in the way when editing code to the point that it's easier and faster for me to just disable it.
Using this beast as intellisense is just one application (called "Copilot") and it has all these annoyance factors sometimes. But I am not talking about that.
To me, this is like we found a way to transform iron to gold with low energy usage, and people are complaining that gold is not that useful. And most chemists not even hearing about the news. I'm constantly amazed by this, every single day, as I read threads like this one.
I'm also wondering why the news coverage of this is so abysmal. In the Netherlands there is very little if any awareness of this, in my bubble at least. (We all seem to very aware of every goddamn bowel movement of every soccer player.)
Even my government seem to just very recently become aware that using data and generally working in a structured manner is preferable to just winging everything all the time. Some departments are even starting to use basic statistics which some even call AI. Nobody is quite sure what anything means and how to make sense of it and all these high-level decision makers studied history, administration or something legal. Absolutely nobody with a clue about anything beyond '80s tech - if even that. It's downright disturbing to see this immense gap and we are supposed be somewhat advanced. But I digress..
I'll admit I haven't played with Copilot yet (since I don't think my employer would be happy for me to send off proprietary code to third-party servers, so I've effectively self-banned myself from using it at work*), but I'd feel that for anything non-trivial like your example of complex SQL queries I'd be reluctant to use the generated output without extra scrutiny (essentially a very fine-toothed code review, which is exhausting).
My opinion will probably change as the tools become more mature, but for now I'm treating them as toys primarily which limits the excitement.
Something like TLDR is less risky as it's not producing code, just summarising it, but I'd still feel wary to trust it since it's such a new field. Maybe this speaks more to my own paranoia than anything else!
EDIT: *and on this topic while I'm here: I'm actually a bit confused (and honestly... jealous?) on the topic of privacy for these kinds of external models. Is everyone who's using Copilot and tools like this working at non-Bigcos? Or just ignoring that it's sending off your source code to a third party server? Or am I missing something here?
It'd be against the rules to use external pastebins or other online tools that send off private source code to a server, so I'm kind of shocked how many devs are talking about how they use AI tools like this at work... is this just a case of "ask for forgiveness, not permission"?
And I'm not saying this can replace developers, as it clearly isn't capable of building complete codebases and reasoning about the system as a whole. But writing self-contained code snippets seems like a solved problem to me, and I think that's the biggest thing that happened in our field since a long time ago.
Generating correct code would be useful.
I have the same issues with these tools, but the one situation I can imagine it being really useful is people who are good at reading and understanding code, but are slow typists. Or more particularly, people who have to think about typing, no matter what the speed is (though I think they're usually the slower ones). I believe it's only once you can type without thinking about typing, and have done it for a while, that these tools become an annoyance because you've gotten used to not interrupting your thoughts on the problem at hand.
It's something run repeatedly, so small chances will occur. Amoung it's failure states are being very, very wrong in ways that are hard for a skilled human to detect without more work that writing from scratch.
And no one in the space is discussing ways to eliminate categories of bugs, only ways to reduce the frequency. Most of those solutions have the side effect of making the less frequent bugs harder to detect. On balance, that's worse.
And, less importantly, it's only useful for writing boring code that should probably be generalized to an API. Sure, I write plenty of that, but it's not an exciting area to follow in my spare time.
Really? What exactly is the bar then? I'd say most professionals I know hover somewhere between 95% and 99%.
Except you're pretty used to the sorts of bugs you write, and the AI isn't you. So these bugs will be harder to find.
So why is this better than writing by hand? Most of the hard work of programming is figuring out specs and debugging, not banging out well understood and specced implementations.
Also makes for a great plot summary for the original Jurassic Park
I read hn pretty regularly but unless you’re really excited about the AI space a lot of this news washes over you and you mostly ignore it.
Q: Best selling artist per country
A: https://pastebin.com/qBVu2mvc
Needless to say, this query works and returns the data I wanted. Whether this is useful or not is up for discussion. But I cannot understand how it's not amazing.
If each invoice has 1 line, why does a line table exist?
Here are some more examples if someone's interested: https://pastebin.com/5Vr08N7z
Call me when I can download and finetune the weights, like I can with Stable Diffusion.
In one scenario, I took the slow running long MySQL query and rewrote that with Codex in 2 mins.
But I think people have started to realize the potential now.
Pitch: My app https://Elephas.app brings GPT-3 and Codex to all applications in MacOS. Many business professionals are using it.
And does having this reduce the amount of discomfort badly readable code creates, and thus make you less inclined to take care the code is and stays easily readable?
i = 0x5f3759df - ( i >> 1 );I'm assuming that if you throw enough training-data at it, it will have seen the same "equation" (or at least the constant) right after an explanatory code-comment.
Other similar situations would probably confused it, but 0x5f3759df is pretty famous at this point.
4. What the fuck?
...
And returns the approximate inverse square root.
/* we don't use the actual price but the discounted price, as per email from Manager Bob on 2022-09-16 */
subtotal += price * customer_discount_factor;
or
/* note there's a 2ms delay while relays settle; this is subtracted from sample time, so timeout is not what you might expect */
select(0,&readfd,NULL,NULL,&timeout);
edit: Comment modified since thought submitted link was about the tldr pages rather another tool. For something like tldr pages but legal-wise there's https://tldrlegal.com (https://news.ycombinator.com/item?id=7367027).
Actual human documentation would read something like:
> Return true if the X-HELPSCOUT-SIGNATURE request header matches the
> base-64 encoded SHA1 hash of the raw request data.You can think of “too long” referring to the time it might take someone to reason out a particularly terse, dense line of code verse the actual length of the code.
I can see something this functionality being useful to explain dense, ungooglable code, like regex, or maybe APL. That said, I couldn't really trust current-generation ML to actually produce a correct explanation instead of being confidently and wildly wrong.
Probably this still gets confused like humans do when variables and functions are named less clearly or even plainly wrong.
I wonder if reading explanation like this makes you more likely to believe code is correct, even if some details are wrong.
In this signature example, you can read the wrong header, calculate hash the wrong way, compare hashes wrong way, etc. there are some many tiny mistakes.
Wound up being useful for explanations of long code paths across file boundaries in large code bases.
If it could give you some context about the implications I could see it being handy for static analysis one day.
Think about it has a general picture description vs a description pixel by pixel. Both have value depending on what you need.
Isn't this a variation on "who needs comments anyway, just read the code?"
The function adds 5 to i, and if j is less than three, the function sets it to 2.
Does it describe this code?
function f() {
i = i + 5;
if (j < 3) {
j = 2; // <--
}
}
... or does it describe this code? function f() {
i = i + 5;
if (j < 3) {
i = 2; // <--
}
}
The answer is yes, and this is ultimately why we use programming languages rather than natural language to instruct computers.In the English you can recognize the ambiguous parse and reject it.
Natural language doesn't have such rules. My example above isn't an example of incorrect English, there isn't a correct interpretation. The grammar is correct and both interpretations are correct.
The comments then should then be reserved for the cases where it is not obvious why or how something is done, or for assumptions that can't be expressed through code (type system or similar).
So instead of having the following
// set i equal to five
i = 5;
either skip the comment, or, since 5 might seem arbitrary, comment why 5: // first four slots are reserved
i = 5;That's not what this "TLDR plugin" does, because there is no magical translation. You put whatever you want in there. I think this is obvious.
Ironically, you basically rephrased the title, while objecting to it's use?
> Comments explain why the code is there
> TLDR explains what a piece of code does
It's still not clear how this explanation is supposed to be sufficient to explain the conclusion.
If a comment is collapsed in an IDE and you spend 1 click/key combo to expand it, versus some key combination to pull up the TLDR, the difference is what?
Comments being embedded IN code could be a thing of the past. No PR necessary to maintain them. Sounds like an improvement to me.