Autocomplete from Stack Overflow
emilschutte.com
emilschutte.com
Sure, we have libraries, apis, package managers etc. but every time I read a code base, there is always a utility function reinventing the wheel.
Someone wrote it because it is still difficult to discover modular code and reuse it easily.
It's pretty nuts when you think about it. Imagine mechanical engineers having to recreate the same CAD file because they can't find a component with the same functionality (which does happen but for other reasons e.g intellectual property).
But we don't have the same intellectual property challenges and yet the best tools we have to discover and reuse code are 90s search engines that treat code as raw text with zero contextual understanding and often outdated code snippets.
PS: I love your project btw.
Don't get me wrong, libraries are wonderful and useful and you should be using them where appropriate, but like everything they have tradeoffs. I've seen way too many projects that try to just tie every single library they can find together without having to write any code. I know these projects because they're a nightmare to install and keep working and I end up spending way too much time trying to hack around all of the code rot to get them running again.
I'm tempted to say developers shouldn't have been able to target specific dll versions, but that would cause different problems.
In the grand scale of things that are wrong with programming that's somewhere between non-problem and minor irritation. Disk space is cheap.
I don't understand what part of this doesn't qualify as a "reusable module" to you.
- How do you search for a function other than by description or name. Assuming the developer classified properly and you even happen to use the same domain language as the writer.
- When multiple modules are returned, similar packages, how do you compare and select efficiently without wasting hours reading the code, evaluating it, checking if it is still maintained and hoping there are no hidden bugs etc.
It is easy to share code. Package manager even facilitate updating code. But it is still a nightmare to find/discover code and reuse it
It's like building a lego set. You know the building block you want but if you have to spend hours looking for it, you give up.
I believe competent developers write shitty code not because they are lazy or inexperienced but because the cost of writing great code is very high.
You were asking for modules and now you're down to the function level. That's a much easier problem to solve. Hoogle, Javadoc, documentation and good index engines solve this problem easily.
Note that type only is pretty ambiguous and most of the time useless without additional metadata. Types tell you what, they don't tell you why or how.
> - When multiple modules are returned, similar packages, how do you compare and select efficiently without wasting hours reading the code, evaluating it, checking if it is still maintained and hoping there are no hidden bugs etc.
At some point, the human brain has to step in and make a call. You can't expect an automated tool to dictate to you the code you need to write.
Let's say you want that `contains` function from the original post. You'd search for `(Eq a) => a -> [a] -> Bool`, which describes a function that takes two parameters and returns a boolean. The first result[2] is the `elem` function, which is exactly what we wanted!
This is a bit of a contrived example, but I have honestly been surprised by how effective searching by type signature is in Haskell. I wonder if it is possible for a language with a weaker type system, like JavaScript.
[1] https://www.haskell.org/hoogle/
[2] https://www.haskell.org/hoogle/?hoogle=%28Eq+a%29+%3D%3E+a+-...
map (elem x) [list1, list2, ..., list10]
or, for folds, the function is the one least likely to change, so you put it first, and the list is most likely to change, so you put it last: foldl (+) 0 [1,2,3]
map (foldl (-) 100) [[1,2],[2,3],[4,19]]
Of course, it's always debatable if you can find some situation where you wanted to curry in a different order, but generally you pick in order to reduce forced named lambda parameters. As far as I can tell, that's how it's done.Edit: if you saw the sneaky edit, you'll know that this particular convention isn't easy to follow!
If there are multiple close matches, ordering of arguments can change which is first, but the one you're looking for is almost always in the first few.
[1] https://www.haskell.org/hoogle/?hoogle=%28Eq+a%29+%3D%3E+%5B...
It's not necessarily a problem. It actually cuts both ways.
The code is not going to benefit from updates, but the code is not also going to be harmed by updates. Think API or behavior changes. Additionally, if the use case for a dependency changes, sometimes the "new" solution doesn't match well.
Ex:
I work with jscript in ASP sometimes. A javascript library changed such that instead of doing iterative processing of nested items (pushing an item on to an array, then looping, popping off the item and processing it, then repeating), they changed to use node's nexttick with some logic like 'we want to move to nested function calls for processing, and doing that iteratively would blow the stack, so we'll just use nexttick so it won't have an increasing stack'. Well, jscript in ASP doesn't have nexttick or any equivalent timer. So while the original code itself worked flawlessly, the use case of the authors moved and we pulled that code in as an external dependency.
That's obviously an extreme case, but I can't really count the number of times that an API change in an NPM package for node has meant modifying code, without a change in functionality.
So yes, you get updates, and sometimes those are going to be security updates and real bug fixes. Other times though that update is going be adding new functionality (and possibly new attack vectors), or dropping support for your use case, etc.
Like many things in our field, it's wisest to look at the risks in all cases, evaluate them for the specific situation at hand, and then choose the appropriate one, rather than cargo-culting one 'best practice'.
There probably isn't a nirvana just waiting for That One Great Tool. Or, alternatively, if there is, it probably involves having to switch to something like a dependently-typed system or the very bleeding edge of where the Haskell community is right now, which is probably still not really quite where it needs to be yet for this. And that would still involve a lot of cost for that switch.
Oh, and this is one of those places where I perhaps may be justified in pointing out again that for all the sound and fury in software development, in reality our field does not move that quickly. There's a ton of package managers out there, but broadly speaking, they're the same set of features shuffled into various combinations. Nothing wrong with that, per se. Just a lot less innovation than may initially meet the eye.
That's because those modules also have multiple (and competing) dependencies though.
If most modules in all languages depended on the same, lower level modules (with rigidly fixed APIs) that wouldn't the case.
Sometimes the wheel is invented because the existing wheel isn't built for the same purpose (ex: off-road).
ex:
I recently needed to check a buffer to see if it's contents were valid UTF-8 in nodejs. I wanted to do this without having to run it through something like StringDecoder, because it's an intermediate buffer, and I didn't want to have the overhead of actually converting it since it wasn't for my use. So I do 'npm install is-utf8'. Then:
let isUtf8 = require('is-utf8')
isUtf8(new Buffer('\u0000')) //false
isUtf8(new Buffer('\u000B')) //false
isUtf8(new Buffer('foobar')) //true
So the package is-utf8 is really more like "Is a printable UTF-8 string". Which is certainly a reasonable choice, and a possibly reasonable default, but was not what I needed.The problem is we don't have a safe, easy to build and package with everything else, and easy to consume (without performance penalty) language. The best we have is C, which is not enough.
That said, the real value of stackoverflow is not only in the code but the surrounding explanation and the following comments which discuss the merits and demerits of each solution. So you get to read multiple solutions to the same problem and realize why the top answers standout from the others. In a way, you learn to smell good code and bad :)
So searching on StackOverflow can be a learning experience!
But these kind of developments might be the next step and who are we say?
One problem I foresee is, the highly rated solution could be for a specific version of the stack/software ( the latest stable or the obselete, unmaintained one). In which case, you might end up with an error amplifying the mess.
int add(int a, int b) {
if (a == 1 && b == 1) return 2;
if (a == -1 && b == 0) return -1;
if (a == 13 && b == 7) return 20;
return 0;
}
But at least it passes my unit test suite!Which is how AIs avoid overfitting, in general, by penalizing complexity.
Then you set it loose on its own source - and you have an AI that takes over the world. (Or at least github, which is more or less the same thing.)
Come up with a general purpose algorithm for evaluating the quality of code and we can get somewhere. AFAIK that's the whole point of being a developer: you can compare the code intelligently.
I guess we as programmers should seriously think about the future of our craft.
If, for example, we feel we are basically using the same building blocks over and over again, we should seriously think about organizing libraries and code snippets and questions asking for such snippets in an organized way and provide ways to transpile 1 solution in different languages etc..
We should not accept the current state of our craft as final and rather think about how to improve in general.
If for instance something like StackOverflow has become the Wikipedia of Code then let's think hard about how to make into a full blown tool, with all the features and semantics we need. It was a nice project the way it started and grew but it doesn't have to stay like that forever!
So an accepted JS answer with 49 points would not be included, nor would a JS answer with 500 points that was not marked accepted.
Its amazing how many people cant.
Being able to search and separate the wheat from the chaff is most definitely a skill. Unless you always want to be reinventing wheels.
EDIT: Unless your job is to invent a better bread recipe.
If we encourage the DRY paradigm in the whole development cycle, copy/pasting a function is just a further reach of such paradigm. Now, if a programmer decides whether or not study and understand the pasted code has nothing to do with the quality of the final product.
Furthermore, IDEs by design encourage copy/pasting to speed development, most offer functionality to store code snippets. In that sense SO is like a web extension of IDEs.
I deeply disagree with this and for such obvious reasons.
It may be a hammer when you're looking for a screwdriver.
Even if you find yourself copy and pasting your -own- written code, you should question why you're having to do that.
Is what you're doing perhaps better abstracted to somewhere else? Be that in an abstract class, a helper class, or even in a separate library from which you can reuse it.
IDEs excel in understanding your code, which helps finding where you've stored something and allows auto-completing access to it, which is almost definitely their most important strength.
No, it's the opposite. The whole point of DRY is you don't copy/paste, you put the thing in a common place and have references to it.
Many SO answers are good enough to be libraries. But they should be libraries, not snippets that are copied and pasted. Fortunately we're moving in that direction - our library infrastructure is becoming good enough that it's worth releasing even small pieces of code as reusable libraries, and people are.
But even if there's copying going on at the implementation level, it's an important conceptual distinction to make.
I know you know, I know they know, and I know you know they know. I just wanted to say it, ok?
Siri, make me a facebook clone. ... but for cats! There, now you're a facebook clone.Really, Apple? Nobody thought of this use case? You shipped this thing with that little forethought?
Anyway, this gave me a little laughter today.
This is exactly the direction we should be going. We build the world's most sophisticated engines for predicting human social behaviour, why are we stuck in the 1990s when it comes to autocomplete for a tightly scoped domain like writing software?
Bravo!
https://github.com/byteface/chode
I considered hooking into a variety of other resources but haven't really bothered with it as have other things going on.
var contains = function (needle, haystack) { return haystack.indexOf(needle) !== -1 }
IE8 and older to be specific[0]. But, if you're supporting IE8 still you have bigger problems at hand.
Microsoft stopped supporting any version lower than IE11 as of January 12th of this year[1]. So unless all of those older IE machines are accessing an intranet (and only an intranet) their risk of being hit with malware is way higher than average.
For my clients I do a cursory IE9+ review, but if anything lower is a requirement then it's an additional fee to make a site work with it.
In a few months it'll be IE11+ and the same fee structure will apply to anything lower (I'm just waiting for IE10 to wind down a bit).
[0] http://caniuse.com/#search=array.indexof
[1] https://www.microsoft.com/en-us/WindowsForBusiness/End-of-IE...
Sure, the code on Stack Overflow is licensed as MIT... but what assurance is there that the code which was posted is original property that the poster owns copyright to? What assurance is there that the posters won't claim patents on the methods being used?
The risk is low, but it's certainly not 0. Big companies caution their software developers to not even read code in SO answers to avoid lawsuits... how much risk do you bring to your company by copying and pasting (not to mention autocompleting) from SO?
I am not a lawyer, etc. etc.
1: https://meta.stackexchange.com/questions/271080/the-mit-lice... 2: https://meta.stackexchange.com/questions/272956/a-new-code-l...
Oracle has even refused to take a 3rd party's code and integrate it into their systems, even when all rights were assigned over.
I'd like to see a version of this based, say, on React examples. Might save me some time :)
These are key functionalities in order to actually be usable. With this you simply get hundreds/thousands of suggestions that are not related to what you are coding. And if by chance the suggested code is exactly what you want, then the variable names wouldn't even match. This is assuming that the suggested code is complete and functional.
https://www.gitbook.com/book/tra38/essential-copying-and-pas...
Found yesterday in this HN Story: https://news.ycombinator.com/item?id=11333448
It's funny because I thought "you can't just copy and paste SO, there's a number of things to think of to be able to do it correctly". Guess I'm not the only one who thought that.
Anyway, most intellectual work is just applying known recipes (copy-paste, renaming a few variables), so it's just a matter of granularity and sources. The hate for SO copy-pasters owes a lot to the few beginners who don't understand that yet and take copy-paste literally.
This post fills me with dread, it's come to the point during code review that I have to be sure the shitty devs aren't coping the bad answers from stackoverflow... I hope they don't find this.