The most copied StackOverflow snippet of all time is flawed
programming.guide
programming.guide
Then we run that against source repos, we could get update notifications for copypasta'd code.
"In file F at line L, it looks like you used some code from SO at revision R. In revision R', it's been corrected."
[1]: https://wiki.haskell.org/Hoogle#Theoretical_Foundations
It turns out when you can package snippets you use so many you can't possibly keep track and audit them all.
Just look at the Left-pad thing, or the event-stream thing.
Those prove that we could see the problem. Brokenness doesn't go away when you grab a snippet of code or reinvent the wheel, you're simply unaware of how much of it is buggy or broken.
At that point, the packages are essentially "indexed snippets" of code.
“I see you used “for i in...” and that copies this SO question about iteration...”
One technique would be to try and define what constitutes "trivial" code.
Another would be to prioritize sources. Documentation from standard or major third party libraries should take precedence over SO.
Another would be a feedback mechanism. If repo authors vote a particular snippet up or down, after a threshold it could be excluded from matching.
Or you could opt-in by means of a comment, though this might make it useless.
> I qualitatively analyzed the top 50 clones in that list and was able to identify the source (or at least a source) of the snippets in most of the cases.
[1]: https://meta.stackoverflow.com/questions/375761/how-to-handl...
The article even mentions:
> FWIW, all 22 answers posted, including the ones using Apache Commons and Android libraries, had this bug (or a variation of it) at the time of writing.
999,999 shouldn't be big enough to be affected by floating point rounding.
If the comment is correct then Java will evaluate ( 999999 >= 1000000 ) as true, which simply cannot be correct.
In the end, the number will be displayed to one decimal place, eg. "1.1 MB" or similar due to the "%.1f" format specifier.
If the input is 999999 bytes, the loop will see that it is less than 1 MB and so will format the number 999.999 into "%.1f kB". When this number is rounded to one decimal place as part of that formatting, it rounds up to "1000.0 kB". This is the wrong output.
There's a bit of a catch 22 because ideally, you would be able to do the rounding first, and then see what units need applying, but you can't do the rounding until you know what the divisor will be.
The article solves this by first manually determining what the cut-off should be (it will be some number like "...999500...") but personally I'd probably just decide to round to significant figures instead so that you can cleanly separate the "rounding" and "unit selection" steps.
With the loop at least it's easy to adjust the thresholds if that is desirable, although a comment will be necessary to explain why you are making the cutoffs in weird places.
Interestingly enough in the past I've used a loop similar to this:
static char suffix[] = { ' ', 'k', 'm', 'g', 't' };
magnitude = 0;
while ( value > 1000 )
{
value /= 1000;
magnitude++;
}
printf("%.1f%cB", value, suffix[magnitude]);
Which is bad because it's subject to repeated rounding from the division, but avoids the problem you described.That's actually not true (as somebody mentioned on reddit) https://github.com/openjdk-mirror/jdk/blob/jdk8u/jdk8u/maste...
It is probably slower anyway, but I was surprised to see no loops
You run into the same problem if you are writing something like the C 'itoa' function (integer to ascii); if you want to write the digits out front to back you need to know what divisor to use for the leading digit so you need to either look it up in a table or take the log.
Taking the log is a lot slower than the table lookup, I found that out the hard way.
People convert so many integers to ascii and it is shocking how slow ascii <-> binary numeric conversions are compared to binary numeric operations, so it's not a matter of "premature optimizations".
Now you can write an itoa which generates the digits from back to front and not have to worry about copying the results because you return a pointer to the middle of the result buffer but then memory management gets more complex...
edit:
The point being that you lose all connection with a snippet after you copy+paste it. I can clearly see benefits when you centralize its development, make use of the collective mind to harden it, and get notified about possible updates whenever an edge-case is found.
Rust has a nice solution for this though: tests can also be embedded in documentation comments:
/// Adds one to the number given.
///
/// # Examples
///
/// ```
/// let arg = 5;
/// let answer = my_crate::add_one(arg);
///
/// assert_eq!(6, answer);
/// ```
pub fn add_one(x: i32) -> i32 {
x + 1
}
From: https://doc.rust-lang.org/book/ch14-02-publishing-to-crates-...More languages could adopt that idea, and a good StackOverflow answer would include those tests in the snippet. StackOverflow might even automatically run the tests and add a passing/failing badge!
(Though not in the stdlib til v2.1, April 2001)
https://elixir-lang.org/getting-started/mix-otp/docs-tests-a...
Rather, I think this is an argument that this kind of functionality should be in the standard library; perhaps in the equivalent of `*printf` for each language.
https://www.zdnet.com/article/two-malicious-python-libraries...
Ironic that this was published today. :-)
Like a web browser does?
Random code snippets from the internet are obviously completely unsafe. There is therefore basic "due diligence" to apply when considering using one such snippet:
1. Very carefully read the code to understand it.
2. Test it (corner cases/threshold values are the trivial things to test for such a piece of code doing conversions)
In general I do not copy-paste code snippets. I use them as examples of how to perform a task or how to use an API, then I write my own code. This also avoids IP issues.
On the other hand, when I download and use Openssl (for example) I am reasonably confident that the code was developed and scrutinised by people who know what they are doing.
I agree I don't like the risk but is npm more risky?
C and C++, which haven’t made the decision to bundle a package manager with a programming language (which is dubious IMO because they are almost completely unrelated concerns), and for which you’re normally supposed to get dependencies from your curated, maintained, OS-provided repositories.
The NPM community is well known for their one liner packages with most of the work done in the dependencies. If that one liner has 5 dependencies and your project transitively depends on it 5 times through react or something then you end up creating 5 times as many files than are needed. It's very easy for a trivial application that uses NPM to have a million files in the node_modules folder.
"These licences have been described pejoratively as viral licences, because the inclusion of copyleft material in a larger work typically requires the entire work to be made copyleft."
Now how much code uses something copied from SO? And I wonder how copyright even applies to "code snippets"?
The Software Freedom Law Center has a substantial article on copyrightability of code: https://softwarefreedom.org/resources/2007/originality-requi...
The important point here is that there's no minimum length for code to be copyrightable. It simply needs to be original and at least minimally creative. Since at least thousands of other developers have found the snippet to be useful enough to directly borrow rather than writing an equivalent, it sure looks copyrightable to me.
"In particular, the laws stress that it is a programmer’s expression of some functionality that may be protected by copyright, and not the functionality itself. If code embodies the only way (or one of very few ways) to express its underlying functionality, that code will be considered unoriginal because the expression is inseparable from the functionality. Similarly, if a program’s expression is dictated entirely by practical or technical considerations, or other external constraints, it will also be considered unoriginal."
Sounds like a case that at least some snippets aren't copyrightable.
You could reasonably argue that every piece of code is completely and only expressing functionality, because it's all inherently directing the computer to do stuff. So only comments would be protected.
On the other hand, you could instead argue that every piece of code can be translated into another language, and in fact is, whether interpreted or compiled, so the source code is exclusively expression only as the functionality is never tied to it.
But it doesn't seem to me to make any sense to say that some part or aspect is expression and another is functionality. It's all or nothing.
Also seems that copyright regarding code and the whole world of copyleft is still a grey area in the courts.
I wonder what standards colleges and research journals have now.
I was just thinking about how everyone ignores the license for the code on SO. The code and way it's used is flawed.
I'm not a professional developer, but I've copied snippets from SO before. I've always included the answer URL in a comment next to it though. But mostly because if I ever had issues with it, I wanted to know where I got it and on the off chance anyone else looks at my code, I wouldn't want them thinking I wrote code I didn't write.
In legal reality if no license is attached to code you have almost zero rights to copy, use, or distribute it.
So if a casual browser of SO doesn’t see any license terms, they should assume that doing almost anything with the code is illegal.
It’s a testament to SO that those URLs have still worked when I clicked on them years later and often provide valuable context to some esoteric bit of code.
The way I understand it, if you fix a bug in an SO code snippet, then in theory you own that bug fix back to the crowd.
EDIT: The more I read about this the more cringeworthy it becomes. CC BY-SA is not designed for source code. The language in the actual license is extremely imprecise for trying to reason about using modular bits of CC licensed code in a complex system.
“Adapted Material” is defined simply as;
Adapted Material means material subject to Copyright and Similar Rights that is derived from or based upon the Licensed Material
... and ...
in which the Licensed Material is translated, altered, arranged, transformed, or otherwise modified in a manner requiring permission under the Copyright and Similar Rights held by the Licensor.
Is my 100,000 LOC application “derived from” or “based upon” a function which converts a hex string to a byte[]? That’s a question that can only be decided in a courtroom, and to my knowledge CC BY-SA has never been litigated in the context of source code.
My inclination is that the terms “derived from” and “based upon” must necessarily carry a stronger meaning than, for example, “incorporates” or other similar terms which would not imply a central shared function or feature.
[The Error of Our Ways • Kevlin Henney] https://youtu.be/IiGXq3yY70o?t=1368
[slides, page 32] https://gotocon.com/dl/goto-berlin-2016/slides/KevlinHenney_...
It would be fun for someone to make a standard library that consisted of all the highest voted SO utility functions.
Although including numbers such as 10^N - 1 isn't out of the question either.
You'd be violating the ToS.
ls -l --block-size=MIt should always use the highest relevant prefix to avoid confusion imho.
Have you ever used the "examples" section of a man page or do you prefer to scroll through the (sometimes long) list of parameters?
In my opinion, code examples are an important tool for teaching real-life usage patterns.