Never run ‘python’ in your downloads folder
glyph.twistedmatrix.com
glyph.twistedmatrix.com
Don't get me wrong, I LOVE Python! It's my go-to language for both small and big tasks. But sometimes I do wish Python had a function like PHP's require()… (Obviously for package-local imports, not for importing other packages.)
project
├── main.py
├── package1
│ ├── module1.py
│ └── module2.py
└── package2
├── __init__.py
├── module3.py
├── module4.py
└── subpackage1
└── module5.py
As long as the file containing the main() function resides in the top-level folder you can use absolute imports in every file, e.g. from package1 import module1
from package1.module2 import function1
from package2.subpackage1.module5 import function2The only way this can be accomplished is by designing the import system from the start to use a function instead of a statement. It's too late for Python to redo things and change its import system, unfortunately. It really is the one language feature you have to get right the first time or you'll miss out on those kinds of enhancements forever.
[0] https://en.m.wikipedia.org/wiki/Live_coding#Notable_live_cod...
The docs for Python's importlib say as much:
> If a module imports objects from another module using from … import …, calling reload() for the other module does not redefine the objects imported from it — one way around this is to re-execute the from statement, another is to use import and qualified names (module.name) instead.
But Lua isn't actually different here since you can still bind the members of an imported module to local names manually, it's just that there is no "from" keyword that does this for you. My personal theory is because of the lack of something like "import * from" in Lua, it became more common for library authors to use late binding and access the returned module by indexing into it. However this is just my personal theory.
Also, now that I think about it you still have to make adjustments to your code in order to have hotloading be feasible, even in Lua. I deliberately chose to use late binding in the code I wrote with "qualified names" (really just table indexing). And all the third party dependencies are versioned in the project's repository, and don't have dependencies on other external modules, so there was no need to worry about a dependency breaking hot reloading because it uses imports differently. I think this might be a cultural thing. A lot of Lua libraries are written in a single-file, no-dependency manner which makes them much easier to integrate in the ways I want. Python seems to have a much larger ecosystem with many external dependencies depending on others.
So yes, I think it makes sense that Python could have module reloading, but the problem is that the language encourages the use of local binding from imports which is incompatible with it, and there is already a lot of code that imports things like that, so it's not as practical if you're using libraries from pip or use "import <...> from" anywhere in your codebase.
Edit: the below is probably wrong about what you're doing. It sounds like your require-replacement merges the new return value of a module into the old one, so code that accesses stuff through the module's top level table will get the right result.
Wrong stuff: Maybe the hot reloading system you use relies on modules mutating the module object rather than making a new one. This is pretty unusual among Lua modules though. I usually see people unconditionally make a table and return it.
from importlib import reloadWhy not just replace the `__import__` function at program startup, when you haven't imported anything? Monkey patching `__import__` at any other time seems like a bad idea.
https://docs.python.org/3/library/importlib.html#module-impo...
So in many instances, there's no need to replace the require builtin: one may simply add a new function to `package.loader`, anywhere in that order that's useful.
I've done this, and draw most (basically all) modules out of a SQLite database. It works like a charm.
If one has need to invalidate the cache for a given module, as comes up in live coding environments, simply set the require string in `package.loaded` to `nil`. Small caveat: the `package.loaded` table is 'special', in that the C which `require` is written in hard-codes a reference to it: one may not simply replace `package.loaded` in the global namespace and expect the new one to be referenced.
Python’s import goes through the configured finders (sys.meta_path) and invokes them in-order with the name of the package being imported until one of them returns an import spec. Sound ‘bout the same to me?
I’ve written loaders which created “virtual” modules on-demand in the past.
I regularly have to spend some time on pandas source code just to untangle the mess of import aliases magic they do.
But the standard library is no exception, importing os will dynamically import os.path.
Now that I do everything in a container, life is much improved.
Try Docker as documented in the AllenNLP README:
$ mkdir -p $HOME/.allennlp/
$ docker run --rm -v $HOME/.allennlp:/root/.allennlp allennlp/allennlp:latestIf I started a new job or new project, it was the same story every time. “Do these five things and then you should be able to run the code” after half a day of back and forth I’d have a list of ten things, and the next new person would discover it is actually 12.
Docker doesn’t just change your experience, it forces the author to actually capture most of their prerequisites. To the point where even if you don’t use docker, you benefit from a project having at least tried to use it.
https://github.com/tensorflow/tensorflow/blob/master/tensorf...
Inside the source code files, they then use a decorator (@tf_export) to define where in the package/module hierarchy a function/class should get placed, see e.g.
https://github.com/tensorflow/tensorflow/blob/master/tensorf...
For instance, the write() function in that module actually gets exposed as `tf.summary.write()` even though it's contained in the file tensorflow/python/ops/summary_ops_v2.py.
Needless to say, this even brings auto-completion of an IDE like IntelliJ to its knees at times.
To make things even worse, things like Tensorboard also make use of global state everywhere, see e.g. the function _should_record_summaries_internal() in the above file. But now the question is: Which module's global state does the function actually access if it gets "moved" to a different module by TF? At the end of the day, I gave up hunting down that bug (_should_record_summaries_internal() always returned False, no matter what) and resorted to mocking the function out whenever I used Tensorboard. To do that (i.e. to find its parent module in the first place), I in turn used the `inspect` package.
This is what I meant by "import magic". :)
I’ve never had a problem with python’s import system either.
I ask because from what I understand from Google's monorepo, it should be easy for engineers to take direct dependencies on whatever they want.
Of course TF is an exported out library so things might be different, but surely it won't then be a good example of Google's DNA?
Not correct. The dependency system supports visibility restrictions - you can take direct dependencies to anything only if you're in your own experimental tree. Otherwise package authors can control access.
You can see that even in the Bazel: https://docs.bazel.build/versions/master/visibility.html
(google employee)
It's not like most people would want more complexity (at least I hope not), but somehow it kept happening, and I feel at least some of it is cultural.
1) They hire a lot of smart but inexperienced people. I've noticed that those kinds of programmers often write clever and overengineered things, which means overly complex.
2) They get promoted for creating complicated things. They don't get promoted for writing simple things.
This article is too long for what it is trying to say. Although I will go the author a kudos for being able to write so well about nothing - it took me to scroll about a third of the article to realise that what they were saying was blatantly obvious.
PS: the only reason I clicked on it was because I thought python by default would import files with some “magic” names in the current directory.
python -m pip install ./totally-legit-package.whl
in a folder will execute a file pip.py in that folder jupyter notebook ~/Downloads/anything.ipynb
is a potential risk and may run “modules” in ~/DownloadsIt's really far fetched to go for python or jupyter when there are easier options like system DLL. The downloads directory is fundamentally unsafe, that we can agree on. Thankfully browsers don't let website push files unlike what the author may imply.
That's true on Windows but NOT on Linux (which I actually find a bit annoying for deploying applications to users). On Linux you could explicitly add "." to your LD_LIBRARY_PATH variable but even then I think it would look in the working directory rather than the executable directory.
Linux ELF binaries can have a library search path embedded using the -rpath linker option. Inside that search path, '$ORIGIN' refers to the binary location so you can have the Windows behavior if you want to.
[1] https://stackoverflow.com/questions/6324131/rpath-origin-not...
You're talking about the current working directory, which can be different, and you're right it was moved later in the search order in XP SP2 (also mentioned in [1]). I had no idea it was ever searched at all, so thanks for that!
[1] https://docs.microsoft.com/en-us/windows/win32/dlls/dynamic-...
/usr/local/bin/notpython ~/Downloads/legit.file
will load libraries or execute code in ~/DownloadsThat seems something specific to python/jupyter
Ruby apparently does not though (the internet shows many asking how to do it). Good for them.
When you run a script, the location of the script is added to the path.
# cat > /tmp/test.py
import sys
print(sys.path)
^D
# /usr/bin/python /tmp/test.py
['/private/tmp', .....]
And when you don't, the current directory is added to the path. # cd /tmp
# /usr/bin/python
>>> import test
['', .....]python -m pip is often advised over pip or /usr/bin/pip, because it uses the version of pip associated with that particular major+minor python install.
But in 99.9999999% of cases it will do what you expect (whereas "pip install x" will be easier to use the wrong one if you have multiple Python versions).
Also the same goes for modules. Since I've installed youtube_dl in my global python 3 I'll always run it like py -3 -m youtube_dl and update it using py -3 -m pip install youtube-dl -U. Or to create a new python 2 venv (for whatever reason) py -2 -m virtualenv venv.
Nobody does this with a downloaded file. It’d just be ./install from the unpacked directory.
But I don’t think it’s at all a stretch to imagine someone invoking python in the way the article describes.
The parent is talking about having such extremely well done sandboxing so that you can automatically run a compiler on source files (even when downloaded from untrustworthy sources) just to get more info about it.
Well, also because you can open a text editor and have a script with very straighforward code in 10 minutes...
As an exercise, I used Bash to build scraping and download scripts for about a dozen websites. I found that I only needed to dip into Python when doing complicated operations on strings, or I needed access to data structures.
I actually liked the approach of using Bash first for scraping, because it gave me direct access to Bash's superior IPC and concurrency constructs. I could plug in a Python script into the pipeline and it wouldn't be any different than interacting with another Bash process.
As everyone else replying to you indicated, I use jq for it. But when I type `curl http://example.com | jq .`, I'm programming in fish.
Granted, when I want to do something more complex and durable, I'm likely to reach for, yes, python. But nearly as often, I build up a moderately complex jq command, and set an alias to it.
It normally doesn't matter that it's magic because most people don't need to interact with it. And it causes a very intuitive result in the part of the language that most people actually do interact with: that x = obj.foo; x() is the same as obj.foo() (not true in Javascript!). But that doesn't change the fact that method binding happens in a more complex way under the hood than you'd expect.
LOL
When dealing with beginners, a strict indents policy makes sense.
Also the standard data model - makes tons of sense as a way to introduce to OOP.
Java's "this" is just kind of implicit magic in comparison, and it didn't accomplish the same understanding.
I think Ruby's actually a very nice general-purpose scripting language; as much as I like Python in principle, I find myself reaching for Ruby more often.
i had to learn PHP to do WordPress stuff (after programming in other langs for many years). what's so bad about their usage of PHP?
(off the top of my head: in WP's "everything happens in some hooked callback" model, debugging can be a pain, lots of messing with globals and gluing strings together)
(Also, WP's style guidance is super weird compared to, well, just about any other PHP style that I've ever seen anywhere, or that of any other C-style language, for that matter. I know that's subjective, but it's Just. So. Weird.)
class A(object):
def __init__(self):
self.x = 1The pitfall is only part of the problem. The problem is much deeper than that. In Python, the solution lies in the base class, which another author may have written in a totally different time/place. Yet the problem only manifests itself when you, another poor soul, try to derive from it later. This kind of spooky action at a distance pierces abstractions, which is just about the last thing you should want from a programming language... and in a sense it literally violates causality (for the lack of a better word). The poor soul that notices this in multiple-inheritance will have to go very much out of their way to work around it: after spending a while trying to track it down (and understanding what's even going on, which is not easy), they'll have to either monkey-patch the original class at runtime, or introduce superfluous wrappers. That is a very high price to pay, and that's on top of the silent failure that led your code to crash and burn.
But anyway, I was just illustrating Python can be quite complicated (even moreso than C++ in some ways) and has flaws despite its simple and straightforward appearance, and it can catch even experienced developers completely off-guard. Just answering the question you asked, basically.
class B(object):
def __init__(self):
self.y = 2
class C(A, B):
pass
print(C().y)Also, the new snippet has a syntax error, the last version where this was valid reached end of life January 1st, and has been deprecated for 10 years prior to that. Not sure how much stock to put in your python opinions given that context.
Wow, that's a really low blow. I'm arguing in perfectly good faith. Failing to call the base class initializer introduces misbehavior in multiple inheritance and the way this happens in Python is completely unexpected. Not every language is like this.
> Also, the new snippet has a syntax error, the last version where this was valid reached end of life January 1st, and has been deprecated for 10 years prior to that. Not sure how much stock to put in your python opinions given that context.
Er, what syntax error are you talking about? https://ideone.com/3hI4Wj
Isn't this a bug in class C, not your original snippet?
I'm also not sure what you mean by "failing to call the base class initializer introduces misbehavior in multiple inheritance", this seems like an issue related to the MRO of the inherited classes.
If you write class C with class B inherited first, the code runs.
It's not. The bug is in A.__init__. It needs to call super().__init__(). C would work fine in that case.
I say qualified because whether or not the bug warrants fixing is another matter. It's more warranted for public-facing APIs than internal code, since it's less practical for downstream users to modify your code. In your internal code, if your team knows about the issue or just avoids multiple inheritance altogether, or if you have some kind of static analysis to check class hierarchies for you, it might be safe to avoid. (Just listing some considerations that I can think of. There might be more.)
My overall point though was just to illustrate one particular example of a flaw that catches even experienced Python developers off-guard, let alone beginners.
For this to work, should every class also have an __init__ method that accepts zero arguments?
The solution is just to avoid inheritance. It’s full of terrible pitfalls in most languages, but Python more than most.
C also works fine in the case where you understand how multiple inheritance works in Python and structure your classes accordingly though...
I.e. inherit B first instead of A like I showed.
You would not encounter this issue in e.g. C++. It's strictly due to Python's idiosyncratic linearization of the inheritance hierarchy, which can make your sibling class (which you have no knowledge of) your superclass. (!!)
struct Base { Base() { cout << "Base" << endl; } };
struct Derived1 : Base { Derived1() { cout << "Derived1" << endl; } };
struct Derived2 : Base { Derived2() { cout << "Derived2" << endl; } };
struct Join : Derived1, Derived2 { Join() { cout << "Join" << endl; } };
And then you get two calls to Base: Base
Derived1
Base
Derived2
Join
There's virtual base class inheritance but then that's a different problem that's hardly any more straightforward than Python."2 calls to Base" is a misleading way to put it. You in fact get exactly 1 call per every instance of Base in the hierarchy (and they act independently). Contrast that with Python where you can accidentally call the same constructor multiple times for the same instance of that class, and Python doesn't care to prevent you. (And notice how others' proposed solutions in another comment did exactly that, and they didn't realize it either.)
> There's virtual base class inheritance but then that's a different problem that's hardly any more straightforward than Python.
This is debatable (the Python behavior is quite unintuitive) but "straightforwardness" is actually besides my point. At least with virtual base classes, the solution (whether complicated or not) is in the same place as the problem: both are in the derived class. With Python, the solution is in the base class, which another author wrote a long time ago. Whereas the problem only occurs when you, another poor soul, tries to derive from it a long time later. This kind of spooky action at a distance makes it impractical to abstract things away... in a sense it literally violates causality (for the lack of a better word)! That's not something to take lightly, to say the least.
Anyway, your argument seems to be that you should be able to use naively written classes in a multiple inheritance hierarchy. I suppose that's a defensible opinion, but it's also not the case in any language I know of; C++ solves the problem of not calling the "correct" superclass by default by foisting the responsibility of choosing the correct subset onto the derived class author, but doesn't solve the "I called my superclass' method too many times" problem, only the "I have too many (n>1) copies of my superclass' data" problem (the solution to which is also opt-in). Python resolves all three of these, but in so doing requires careful construction of multiple inheritance hierarchies and use of the super builtin that is religious to the point of fanatical. That's a tradeoff made in respect to a feature that, IMO, is inherently sharp.
Aside: "idiosyncratic" is an interesting word to use given that the C3 MRO was originally intended for Dylan.
I know what's going on and my comment was sufficiently accurate to get the point across in 1 sentence. I wasn't exactly intending to get into a lecture on MRO here.
> C++ solves the problem of not calling the "correct" superclass by default by foisting the responsibility of choosing the correct subset onto the derived class author
Which is much better than leaving a ticking time bomb.
> doesn't solve the "I called my superclass' method too many times"
Er, yes it does. Every class instance in a hierarchy is initialized exactly once. Trying to call it twice will produce an error (e.g. "class has already been initialized").
> Python resolves all three of these
It does not. It's perfectly fine letting you call a class's __init__ multiple times. Unlike with C++.
> Aside: "idiosyncratic" is an interesting word to use given that the C3 MRO was originally intended for Dylan.
It is a fine word. "Idiosyncrasy" does not imply there exists only 1 person in the whole world engaging in the given practice:
"Her habit of using “like” in every sentence was just one of her idiosyncrasies."
https://www.merriam-webster.com/dictionary/idiosyncrasy
I'm tired of the pointless arguing so this will be my last comment.
i.e. child logic is handled in the child, not in the parent.
Essentially, why not do it like this:
class A():
def __init__(self):
self.x = 1
class B():
def __init__(self):
self.y = 2
class C(A, B):
def __init__(self):
A.__init__(self)
B.__init__(self)
print(C().y)https://www.python.org/download/releases/2.3/mro/
But; multiple inheritance is a beast in any language due to its complexity and bringing it up in a discussion of "magic" that python exposes beginners to is specious. It's true that C++ doesn't have this problem, because it has no super keyword and you must explicitly delegate everything, but also introduces the problem that you can have multiple instances of a base class and so has added complexity, virtual inheritance, to solve THAT problem. And tons of languages have a problem where omitting a call to "super" is logically required is not statically diagnosable.
It's emphatically not "having a super keyword" or not having it. That's a red herring. Visual C++ has __super and yet it doesn't suffer from this: it errors with "ambiguous call to overloaded function" when there are multiple candidates.
There's a lot more I could say about this (frankly the issue I outlined just scratch the surface of a deeper problem in Python), but this is diverting the discussion. Nobody even said C++ is simple or somehow easier to learn than Python in the first place; I was just pointing out an insidious issue in Python, and frankly, it could've been done better without any need to emulate C++'s approach at all. (Exercise for the reader: suggest improvements.)
What's your point about this nonstandard extension? My point is that C++ doesn't have the problem of classes not written to be part of a multiple inheritance hierarchy delegating somewhere unexpected because, by design, it doesn't attempt to solve the same problems that Python is. If anything that this nonstandard extension doesn't work in multiple inheritance is agreeing with the points I'm making.
> There's a lot more I could say about this (frankly the issue I outlined just scratch the surface of a deeper problem in Python), but this is diverting the discussion. Nobody even said C++ is simple or somehow easier to learn than Python in the first place; I was just pointing out an insidious issue in Python,
You were the one who brought up C++! (EDIT: this was in a cousin comment.) My point is that different languages navigate this thorny landscape in different ways, but the way in which Python does isn't as haphazard as you are suggesting.
> Exercise for the reader: suggest improvements
Perhaps being part of a multiple inheritance hierarchy is sharp and uncommon enough that it should be opt-in. What are yours?
My point was what I said: you can have your cake and eat it too. You claimed the issue was due to a lack of a super keyword and I showed you it would not occur even if C++ had super.
> You were the one who brought up C++!
I "brought it up" to illustrate it does one particular thing better than Python. Not to claim it does everything better than Python. If you want a simple, beginner-friendly language, avoid C++ like the plague.
> What are yours?
Idk, for starters maybe produce a warning when multiple inheritance occurs and one of the __init__s in the hierarchy doesn't have an obvious call super().__init__.
I made no such claim. I said (in my opinion) C++ doesn't have the issue you're raising with Python because it has a completely different design that doesn't attempt to solve diamond problems in superclass attribute lookup. Are we not in agreement?
> I "brought it up" to illustrate it does one particular thing better than Python. Not to claim it does everything better than Python. If you want a simple, beginner-friendly language, avoid C++ like the plague.
You can validly think that C++'s choice here is better. I don't think there's a counterargument to personal preference here. My point was that it does not solve all of the problems Python's super builtin is designed to. And that's fine! It avoids the issue in question! I actually don't understand what you are arguing with here. Personally I think the design goals here are just so wildly different that I don't have a preference, but I admit that's kind of a cop-out.
> Idk, for starters maybe produce a warning when multiple inheritance occurs and one of the __init__s in the hierarchy doesn't call super().__init__.
In general I think having more "warning labels" on multiple inheritance is the right idea, if a language is going to persist in offering it at all. I think this would be a good start, especially for a linter.
Note that `object` being the root of the class hierarchy implies that you must have careful coordination between all classes in a multiple inheritance hierarchy to determine, for a particular method, when subclasses may/must delegate to super, what the arguments to the method are, and even what the argument names are in some cases! You can't just go glomming classes together and expect it to go well unless you know the classes either don't use the same method names or agree very carefully on what the delgatable interfaces are.
But, to be fair deployment and package management of python is a mess when you try to scale up.
Such as? The only thing that comes to mind is significant whitespace - but some other languages are much more sensitive to whitespace than Python.
Or keep their Perl at 5.24.3.
Perhaps the sharp edge would be reduced if Python skipped empty entries in PYTHONPATH. This might require some people to explicitly add "." in their configurations, but they should have done that anyway, and it wouldn't require code to change. PYTHONPATH is also getting less used anyway, so this might be a relatively unnoticed change by most.
What do you think? Would that be worth proposing?
E.g., people will often add a directory to PATH by doing:
PATH="$PATH:my_new_dir"
This is practically always fine for PATH, since it's practically always set. If you use exactly the same syntax with PYTHONPATH, you have an instant security vulnerability. Now the first directory searched is the current directory, and if you ever wander into directories with untrusted fils (like "Downloads/" or "somebody_elses_container/") it could become serious fast. This is a sharp edge that was never intended.
I'm going to look into proposing this as a PEP. The change is trivial, but there may be surprising uses. A PEP would mean that it'd get more scrutiny (so the change will be less likely to be reverted later).
What it would break is running a script and trying to import from a utility module next to it.
> cat "import bar" > foo.py
> touch bar.py
> python foo.py
> python -I foo.py
Traceback (most recent call last):
File "foo.py", line 1, in <module>
import bar
ModuleNotFoundError: No module named 'bar'
and this is exactly the feature which makes it risky to run `python` from an untrusted location.So packages installed with, say, "python3 -m pip install --user" wouldn't be importable.
I don't think I've never encountered this behavior and any time I've downloaded a file I've had the opportunity to see and change the name it is stored under.
So as far as I can tell it would be hard for a site to cause a download to occur without the user noticing.
The file could stay in there for years just waiting for the system’s user to slip up. This is obviously not a targeted attack under time constraints.
Way more problematically, Safari will also open “safe” files (which includes zips and PDFs) by default after having downloaded them.
Hierarchical file systems were invented for a reason. For the same reason, I want browsers and chat apps to ask me where to save each file I download. Fortunately, most of these apps have an option to enable such behavior.
As a nice side effects, this prevents security issues like the one described in the article.
Now the question is: Is SELinux the sawstop the author asks for?
It's a major operating system flaw that no solution for it is in the horizon.
In my view, the focus on low friction and getting new users quickly productive with minimum fuss (shared by many scripting languages) sets the stage for these inevitable discoveries. And I'm not just hand-waving with the vague "these" characterization. Focus on ease-of-use creates roving best-practices factions that wax and wane over the years, with consequences perfectly captured in the famous xkcd "Python Environment" cartoon.[0] The fact that the empty string interpretation in PYTHONPATH hasn't been a bigger deal much earlier raises interesting questions. (Are bad guys not trying as hard as we imagine? Will newbie-friendly languages always end up with these kinds of exploitable surfaces? etc)