It is the intermediary subshell in system() that is the root of all evil, not composing unix programs together; that is what our OS forefathers intended.
It is the intermediary subshell in system() that is the root of all evil, not composing unix programs together; that is what our OS forefathers intended.
The error: you should never pass unescaped/untrusted input into a subshell.
In Python, it’s the difference between the safe:
def make_dir(path):
subprocess.run([“mkdir”, “-p”, path])
Instead of: def make_dir_unsafe(path):
subprocess.run(“mkdir -p “ + path, shell=True)
The latter is easier to type because you don’t need to use the annoying list syntax, but it’s unsafe. If you are writing a library function a subshell is not a sound approach. def make_dir(path):
subprocess.run('mkdir -p'.split() + [path])I’ve seen bugs where some forgot to add quotes and commas to split some arguments and the program did the wrong thing. Parent’s approach avoids that. However, things like spaces in quoted strings are not correctly handled.
I assume you meant something like subprocess.run(['mkdir', '-p', '--', *path.split()]), which in that case using shlex.split will fix your problems with quoted strings.
subprocess.run([“mkdir”, “-p”, path])
This does not work for all values of path. And I'm not just referring to blorgle's complaint. I'm being deliberately vague to make you pause for a second and think.The snippet breaks for paths starting with a minus, as mkdir will interpret those as another option.
The solution is:
subprocess.run(["mkdir", "-p", "--", path]) os.makedirs(path)- performance. Obviously that depends on the circumstances but if you are working on a function that is used in many places you don't know the circumstances.
- error handling and reporting. Are you going to capture stderr from mkdir? If this fails can you provide a clear indication of the problem to the caller?
- paths. If you run mkdir without specifying an absolute path your program may fail if $path is not what you expect. That can also be a vulnerability.
- dependencies. Now you won't run in a distroless container or other cut down environment
- you lose type safety
Not saying "never shell out" but I've been burned by most of these so I am cautious about it.
Certainly the worst of it, but even if you pass the untrusted data as a discrete arguments you're still vulnerable to parameter injection.
Let's say, for example, that you're running ["find", USER_SUPPLIED_PATH, "-type", "f"]. If the user supplies "-delete" as the path, the resulting behaviour would not be as intended.
Would it be noticeably slower to the point it's measurably relevant, though? Or are we talking about something well within the domain of premature optimization?
I mean,if you're doing IO, how often does creating a dir end up being the performance bottleneck?
std::vector<int> numbers = Func();
may or may not make a copy depending on the signature of Func. So it's not clear at the callsite how expensive that line of code is. This is unlike Go (which has this explicitly as philosophy) where you need to call a copy function/use append to make a copy of a slice.In that example, there would always be a move or copy there, though there are scenarios in which it might be elided.
When you read the code, you know what the types of a, b and c are. You know whether they're built-in types, and therefore whether those operator expressions are actually function calls.
Nothing is hidden, so long as the programmer doesn't make invalid assumptions based on his experience with other, different languages.
In your case, if a, b, and c are all the same type, then b+c can return a completely different type, and you'll need to track that down before you know what type a+(b+c) will have. And in that case, (a+b)+c can have a different type from a+(b+c). If you've memorized the associativity rules for the language, you'll know what to expect -- but it's not at all obvious from the notation.
And, sure. If you know absolutely everything about the language and everything about the codebase, you can make accurate predictions about the behavior of any given line. But if you're coming into a new codebase, there's no way to know at a glance what a given line of code is going to cost. That's what folks are calling hidden complexity.
Oh -- and thanks to the auto keyword, you don't always know the types of a, b, and c without hunting things down.
Yes, you need to look at the declarations of names to know what the resulting type of expressions involving those names is. I don't see that as being surprising or unusual. The same is also true in C.
This is, of course, an extreme example. But there are definitely cases where it matters a lot.