Cybercriminals pose as "helpful" Stack Overflow users to push malware
bleepingcomputer.com
bleepingcomputer.com
It’s not like we haven’t got higher quality corpora, on average. They’re just poorly annotated.
Installing packages (i.e. source code) for programming languages should not execute arbitrary code.
Actually it is even worse because JS is more popular and average JS dependency tree is so much more massive. Any tiny forgotten transitive dependency of a dependency (or dev dependency) 15 layers deep can pull this off. Total leftpadization is not a thing on pypi due to different culture
There is the reductive argument that basically everything you download is arbitrary code, but throwing away the code that is run seems uniquely silly.
curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh
curl | tee saved.sh | sh
thenSeems like prohibiting arbitrary code in installation scripts would only help with issues like Bumblebee's `rm -rf`: https://github.com/MrMEEE/bumblebee-Old-and-abbandoned/issue...
True, but in the case I've mentioned, if you've mistyped the name of the package, you can safely uninstall it without any issues.
Furthermore, code is often ran in (sort of) sandboxed environments like Docker during development, in which case arbitrary code on runtime is less dangerous than arbitrary code on install time.
> Furthermore, code is often ran in (sort of) sandboxed environments like Docker during development, in which case arbitrary code on runtime is less dangerous than arbitrary code on install time.
Wouldn't the package install then also happen in a Docker container anyway, negating the problem? Or how would you install a package to your host environment and then use it from within a container?
It can give people also a second change to notice, e.g. the typo in the package name.
That way you could pick a repo based on your risk appetite rather than needing to trust PyPI. Debian style python for the cautious and AUR style python for the bleeding edge and reckless.
But setup.py is part of the Python's legacy, from before PyPI even existed.
I can install python-matplotlib from my package manager, and it appears in my python path. I never bother with venvs.
Maybe my needs aren't advanced (or exotic) enough.
You also shouldn’t be able to push code directly to anywhere other than a centralized repository. All the build stuff should happen in a dedicated and independent process.
From a security point of view you need to accept that the idea of running into a malicious package as a direct or indirect dependency is not only not zero but fairly realistic and you should try to limit the blast radius as much as possible for when that does happen.
But an attacker could infiltrate them by manipulating magnetic fields to generate keystrokes / mouse moves so also faraday cage them.
Tell that to JS, Python and Rust people.
But hey, we are secure. /s
If you have one person asking a deliberate question, someone else answering with a backdoored package that does actually does solve the problem, then a few accounts upvoting and adding comments it would lot a more more convincing than just a random answer with zero upvotes that doesn't actually solve the problem.
Extra credit: have the infected user’s computer also start posting these answers to SO
Users have been pushing their work in answers for years. Not a surprise that some are malicious.
In my case, I never bother looking at answers that prescribe a dependency. I find that it’s very important that I completely understand any code that I use from SO (or any other source), and I usually change the code to better fit my specific use case (and it’s usually a pretty small code sample).
If you want to criticize Stackoverflow, there's plenty of solid ground to do that. No need to say things that aren't.
at what point do you stop being a newbie? I have been on SO for 14 years and have found "lately" or the last five or so years to be extraordinarily unfriendly. I am in the middle of solving a complex problem so I posit a question but it's stripped to bare bones to make it easier to answer -- and the answer is almost always is not what I seek but instead "why are you doing this". Aaaaargh. I would need to post War & Peace to make you understand why so I just go and delete the question when this happens which is too often. Example: https://stackoverflow.com/q/77202800/308851 lots of comments asking why I am doing this were deleted but downvote remains. I kept this one up for whatever reason although downvotes usually are enough to make me delete a question.
My advice: if you can't help then stay away from the question. Alas, this likely won't reach the hopped-on-SO-power idiots but it's worth a try.
Of course, once an LLM scrapes that answer and regurgitates it, then you'd never know.
Since Swift allows extensions of fundamental types, these hallucinations can be quite convincing.
Hell, you could create a library, not have malware in it, answer the questions, wait six months, add the malware. Oh wait, I'm not a criminal. Forget I said that.
Patience is unlikely for individuals, but it becomes a real concern when you come to the point of state-based actors.
I regularly use Python libraries in my work that emulate Excel. and they're so often out of date, they're so abandoned/unloved for the most part because the people that write them eventually move on and they don't have a big enough user base.
That means that they're ripe for somebody inserting stuff. Worse Excel is usually used by big companies, so it's a target as well.
Solving the problem, so the real user would make it the selected working answer, which would help to cascade the malicious package for others that come to solve the same issue.
E.g. Here's a simple solution to the problem:
’’’
from setuptools import setup
from badtoolkit import setupnotmslicious
setupnotmslicious()
setup {
Name = "custom_log",
Version = "0.1"
# rest of your setup }
’’’
This reads like yet another major attack vector through LLMs. A threat actor only needs to provide a few mentions of the package on reputable sites and it is almost guaranteed that your backdoored package will be mentioned by the LLM once it is retrained with new data, or even in cases like RAG.
I am guessing I am not the first to think of this though. Wouldn’t be surprised if this kind of attack vector is already being set in motion for all kinds of other purposes; product reviews, etc.
While there are definitely ethical issues with the data that they are using. Trainers need to get a handle on this kind of thing because for LLMs to be truly useful they have to consume large parts of the web.
I mean, if you're manually curating what you feed to the LLM you end up doing one of the web directories of old and might as well skip the LLM part... or use it just to gain funding...