Python Best Practices for a New Project in 2021
mitelman.engineering
mitelman.engineering
At the lightning pace of only 12 years since the release of python 3 (and 15 months since python 2's deprecation), some linux distribution have ALREADY switched to it.
There's this tiny, very-very vocal minority of people who continue to mystify the 2/3 changes and act as-if migrating code were some massively difficult herculean task that almost ended Python. Sometimes a little bit of "string/bytes split is ALL WRONG and utter non-sense!" mixed in.
The truth is, that these were not actually big changes. The truth is that some orgs just wanna sit on their lazy asses doing jackshit to maintain their stuff and are surprised that, after 15 years, they have to do MAINTENANCE WORK??? ON THEIR CODE??? This is an outrage! The truth is that if porting was hard, or difficult, the code most likely had no meaningful tests, and so you couldn't do any changes anyway. The truth is, that changes like these expose bad engineering. And guess what, people don't like being exposed.
I think you hit the nail on the head with "lack of meaningful tests". The entirety of "locked to 2" python code I've dealt with was a serpentine mess with little/no tests and so brittle even going from print to logger risked things breaking.
"But it currently works!"
Miracles happen, but people shouldn't count on them.
The only difference between the Python3 migration and other projects' major version changes (PHP5-7, .NET4 to newer versions, etc) is that migrating to Python 3 had the equivalent of a couple of major version's changes all rolled into one. Yes, it's painful; major version migrations are not easy.
But it's really no different from not upgrading JS code to keep up with the latest browser security issues, or not upgrading your JVM (there are _tons_ of Java codebases stuck on ancient JVM versions).
If you had a problem with migrating to Python 3 after a decade you would've still had the same problem with any other major software upgrade in your infra.
Also versioning is not an issue, the problem is that Python didn't provide any viable versioning strategy. You just had two separate executables now, and you, as the user, have to decide where to put them in your $PATH and how to make sure the right code gets executed with the right version. That should have been baked into the language or the tooling itself.
That's not really a problem because the Python maintainers conveniently provided a package to do trivial 2-to-3 migrations.
> Also versioning is not an issue, the problem is that Python didn't provide any viable versioning strategy.
There were explicit backporting libraries created to manage the transition. Django depended on these for many years to succesfully support both Python 2 and Python 3 packages.
> You just had two separate executables now, and you, as the user, have to decide where to put them in your $PATH and how to make sure the right code gets executed with the right version. That should have been baked into the language or the tooling itself.
If you're executing a script, Python supports shebang notation to determine your executable.
Sorry but none of these are real problems. The only place where you had migration issues were with very large packages that did complex string manipulation stuff, and that's the sort of code the requires maintenance in any language anyway. It's not like there isn't an enormous amount of precedent in these kinds of migrations (Ruby minor versions break. PHP's 4->5->7 migrations were huge).
That's not the point. For example, when ubuntu dropped python 2.x as the default, some of my workflows got broken because of scripts which I didn't write and didn't even know about. There's plenty of old code which people still depend on. System upgrades should not break your projects, or require you to go in and do surgery on scripts you didn't write.
> There were explicit backporting libraries created to manage the transition. Django depended on these for many years to succesfully support both Python 2 and Python 3 packages.
That's a workaround. A solution would have allowed python 2 & 3 to coexist with no additional effort.
> If you're executing a script, Python supports shebang notation to determine your executable.
Again, you're depending on the people who wrote the python code you depend on to handle this in the correct way. It's not something which is built into the ecosystem.
I'm sorry, but all your arguments seem to boil down to the fact that there are ways to make python usable despite the extreme fragility of the toolset.
On Linux land projects that depend on system libraries and `*-dev` packages break constantly across major versions.
> That's a workaround. A solution would have allowed python 2 & 3 to coexist with no additional effort.
You can do that if you make your python interpreter version explicit in your shebang.
> I'm sorry, but all your arguments seem to boil down to the fact that there are ways to make python usable despite the extreme fragility of the toolset.
A toolset that gave you a decade to upgrade with ample warnings is the opposite of fragile.
The only solution to that is neve breaking backward compatibility because either someone might prematurely upgrade (not what happened woth Ubuntu) or someone might not maintain software (what seems to have happened with the third party tools in question, whether it was the original maintainer or some packager or...) and might also not vendor dependencies, on the assumption that external environments will never change.
> A solution would have allowed python 2 & 3 to coexist with no additional effort.
PEP394 allows that, if you depend on a particular python major version use “python2” and “python3” to refer to it: both side by aide installation and continuity of operation over the time “python” switches from 2 to 3 is provided.
While some linux distros broke the recommendation on when to switch “python” targets, that would be transparent to anyone following the recommendation.
This is an especially bad problem with an interpreted language, since changing the system version of python can break code at a distance. I can run a program today and it works, and tomorrow it won't because somebody changed the symlink to the other python version. That's why there's this huge mess of environment management and containerization.
I will give you an example of the problem. When ubuntu finally removed python 2.7 as the default, some of the workflows broke because scripts were depending on it. Scripts that I didn't write, or even know about.
This is only a problem because nobody bothered to solve it. I think one correct solution for example would have been to have both python versions live inside the same executable, with some kind of flag or something added to run against python 3. Just some standardized way of handling the different versions would have made a huge difference and saved a decade of pain.
Not depending on the system version of anything is normal advice for other interpreted languages too.
Linux doesn't split shebang arguments. So you can't use env and pass a flag.
IIRC, there were a few linux distros that bucked this; at least one (Arch I think, but it wasn't one I used) switch “python” to Python 3 quite early, as I recall, and I think there were some others that did there own thing.
I usually keep a CI branch with dependencies unpinned (I made pip-chill for this use case) precisely to make sure I'll be the first to know when something coming from the future breaks my code.
Being lazy is a virtue, but not if it cause an engineer to avoid work that needs to be done.
Ie someone into Unicode on 2 with u'' got hosed on 3 despite 3 going on and on about importance of string handling. How hard would it have been to support u so someone could support both versions more easily? It was insanity.
Exceptionally hard. Py2 str objects are tantamount to py3 bytes but with the py3 str apis. What this amounts to is every "str" in py2 is an untagged union of str/bytes. Python is exceptionally dynamic. This means if you allow different behavior in different modules, you risk silent, customer-data-corrupting bugs, among other headaches, as things continue to "work" but do the wrong thing.
At least with a hard 2/3 switch, you are on your toes and know there is a (mostly) finite transition period.
from __future__ import str basically guarantees you'll have "3-compatible code" causing headaches years into the future.
That's not even to touch on the difficulty of switching the engineering difficulty of interop of encoding-oblivious strings with unicode ones.
Python 3 has supported writing u"foo bar" for a very long time, I think starting with 3.2 or so.
Similarly, Python 2 accepts b"foo bar".
`curl https://pyenv.run | bash`
hmmm...
The script is served over https, so it's not going to be tanpwred with (unless you have a malicious cert, but at that point you can't trust anyone), and curl | bash isn't any worse than downloading a script and just running it, or running a precompiled binary you don't trust.
They should offer a download with signature validation instead. Signed by Apple, Microsoft, etc if possible.
The safety is in reviewing the code there, not in avoiding curl | bash. Running pip install or npm install is just as dangerous.
> They should offer a download with signature validation instead. Signed by Apple, Microsoft, etc if possible.
If the host is compromised, the attacker will just get Microsoft to sign their malware instead; see [0]. If the host is compromised, and you run the code without reviwing it, you're hosed regardless.
[0] https://arstechnica.com/gadgets/2021/06/microsoft-digitally-...
What if your distro package repository was?
[0]: https://www.idontplaydarts.com/2016/04/detecting-curl-pipe-b...
If you have a fear of your source maliciously serving you different code over curl, don't run their code at all.
> You're better off piping curl to a file, reviewing the file and then running it manually.
Right, but the safety there is reviewing the code. Running brew install, npm install, pip install, or a binary could all run malicious code too.
Something along the lines of:
curl https://pyenv.run | pass_on_through_sdtout_if_hash_matches md5 8bffaf30c9ba21393329d531063056fe | bash
That way, someone who validates the file locally, can be sure that what's piped is the same thing.It requires a few extra steps to be actually secure. You actually need to verify the hash from a trusted source for it to be actually secure. If the delivery has been tampered with, you need to ensure that the delivery of the hash has also not been tampered with. In practice, codesigning is the solution, but certs are expensive, and impractical for a small project.
1. (local) Download the file from the URL.
2. (local) Review it locally, in a text editor.
3. (local) Get its hash locally, from the file in your file system.
4. (SSH) Feed this hash into the fictional tool above.
5. (SSH) If what curl gets is the same as the file that you've reviewed, it gets piped further into bash, otherwise the execution stops and an error is output.
Of course, that's only applicable to this particular case, where a compromised server could detect that a bash pipe is used and return different file contents. That would only be useful in situations where you want to review it on a local device, such as a desktop and run it on a remote one, such as a server.Edit: If you want to review it remotely, there's nothing to prevent you from using less or something to view it before manually opening it with Bash. That just requires the discipline to not use one liners that both download and run it, as long as no such tool like the above exisdts.
You could replace blockchain with checking if it's signed, and the key matches an owner on keybase/github/some other federated identity provider too.
In reality, I'm a more traditional Unix person and prefer MacPorts, where you can do `sudo port install python36 python37 python39` in a very BSD way of doing things.
Homebrew has broken my computer one time too many.
[1]: https://www.idontplaydarts.com/2016/04/detecting-curl-pipe-b...
You can't know that curl and your browser get the same data - but you can for example split it up:
curl https://pyenv.run -o install.sh
#examine install.sh
bash install.sh
Ed: or just "save as" like with an installer.Piping straight to bash can be especially bad if you've cached sudo credentials for the current session - some of these scripts call sudo "inside".
Otoh - the connection is signed (it's https)-unfortunately it's often quite easy to compromise a web site. Obviously, listing gpg signatures on the same page doesn't add much unless it's possible to verify the gpg key some other way.
Ed: another problem is that you really should check exactly what's in you clipboard before pasting to a terminal.
Well, yes. The safety is in doing something between "acquire potentially malicious payload" and "running payload". I don't see how "safety [is] not in avoiding curl | bash" when, avoiding the direct pipe to bash is exactly what I suggest.
If you look at the url, then curl and pipe that url, you have no idea if bash sees what you just reviewed.
I generally run make, setup.py, cargo build etc in the context of cloning a source repository. I certainly could do a better job of sanity-checking those things, but I do try. And I definitely try to avoid having sudo credentials cached when I do - to foil "sudo cp artifact /usr/sbin" and other awful things people do, because they found it convenient.
> Also, if you are untrusting of the source enough to verify their install script is safe, why would you install their template to run on your machine without verifying all of that too?
I generally trust people more to write "left pad" than install scripts. Many sysadmins are good programmers, few programmers are even remotely decent sysadmins in my experience.
> Finally, 10 line bash script might (as tbis example does) just call out to another curl | bash
In which case one has to chase down the rabbit, or give up.
Sometimes one will discover that the end game was downloading a gpg signed tar archive with the release artifacts - and one can go and do that.
> or to a pip install/npm install.
People do do awful stuff in makefiles and package install scripts, but for vanilla python/Javascript - the lazyness of programmers tend to work to our advantage - there be little extra madness/magic in there.
Sute, running pip install -r requirements.txt can do almost anything - but it's unlikely to run your package manager under sudo and mess up your system packages, or add something questionable to your package sources.
True, it does not. I don't recommend downloading (random) binary installers and running them either.
With eg Linux isos, you typically already trust the signing key for your os updates.
But unless you are vigilant about your ssl root certs, you'll easily allow a lot of malicious and incompetent services to potentially intercept most of your ssl traffic... (due to there being many trusted roots by default).
> if someone has overtaken a host and replaced the binaries
This again depend on who and how the binaries are signed, and how the signatures are trusted. Typical windows (and Mac?) setups will gobble up any signature. But if you do check who signs the binaries - then the signing key will easily be the most secure part of the system - a compromised ftp/web site allow hosting malicious binaries, but typically not grant access to the signing key.
With letsencrypt a hacked web site will typically have access to a valid ssl cert - no need to further compromise mx/mail records or gain access to a business phone number etc.
A ascii-armor signed shell script can be distributed safely via a paste-bin. Unfortunately there's no good automatic/standard way to do so. Or rather no standard tool to prompt to trust the signing key - and then run the script - beyond basic gpg --search-key --key-server.. + gpgv.
Maybe signed git repos would be easiest - but I don't know how easy it is to limit which keys are trusted - if it's possible at all?
The helm project does a little dance to try and verify downloads - but for all the effort it pretty much amounts to trusting the script, not the keys/signatures:
https://github.com/helm/helm/blob/v3.6.2/scripts/get-helm-3#...
I was hopeful sequoia might help - but apparently its sqv tool is even worse than gpgv - neither can handle an ascii armored public key, and sqv can only handle detached signatures.
And just for completeness - a reminder that any cut'n'paste in the terminal is a bad idea: https://nakedsecurity.sophos.com/2016/05/26/why-you-cant-tru...
Yes.
You should inspect what you download.
Also, you should probably use the Python interpreters provided by your Linux distro, that stay in directories you usually can't write to and come in signed packages. On a Mac, the next best thing would be MacPorts.
The article specificially goes into why not to do it.
> pyenv allows us to set up any version of Python as a global Python interpreter but we are not going to do that. There can be some scripts or other programs that rely on the default interpreter, so we don’t want to mess that up.
I used custom Python interpreters a lot and it's nice to be able to rely on the system to provide a sensible environment instead of forcing myself to build my own.
Every time I fail to find such a section.
This is the reason why I consciously attempt to move away from Python and choose Go or Rust for new projects if possible. Of course, on existing projects, Python deployment is a pain.
which isn't docker.
From what I recall, spotify were the first company with a large footprint to use docker. However for some reason they skipped VMs and went screaming into docker when it was _very_ new. Personally that seemed like a mistake, but you know, each to their own.
If you're wanting to get into a bun fight about containers, then IBM 360 and JCL has some time for you.
On the other hand, OP’s setup makes it very easy to publish packages. So you can create a tarball and have a user pip install that, then run your app with a simple CLI. Or publish to pypi if you want it public. The downside is you assume the user has the right version of Python and knows how to switch versions if need be and all the weirdness that comes with that. But for web apps, packaging your app also makes it easy to wrap in a simple dockerfile that basically just installs the package and then runs it.
Have you tried Nuitka or mypyc? (Haven't tried them myself but I've heard very good things about Nuitka.)
In every modern corporate environment I've been in, you either have the ability to run whatever on your box, or it is locked down tight.
One should probably run their own package server like https://github.com/pypiserver/pypiserver
All of that said, containers are nice because you have a log of what is running, easy to transport and coordinate.
When you use Go and Rust over Python, does the use of Docker disappear? What replaces it?
Never used pypiserver but I’ve had a good experience with https://github.com/devpi/devpi
Elastic beanstalk is just a horrid dev environment. Lots of waiting, lots of non-obvious options, and very little reward. I would personally push for lambda and zappa (https://github.com/zappa/Zappa) for python, as it seems to be much easier to deploy and debug.
So, seein' that you've coded in both Go and Rust apparently, my question to you is: Which do you prefer of the two (and why)? I personally lean a bit toward Go, but I haven't learned enough of either to decide absolutely which of the two I should learn first.
Rust on the hand, feels like a better C++... with a steep learning curve, and useful when you can/need to get something correct/efficient/safe at the cost of increased developer time and cognitive load.
They both have their place IMHO.
I've spent many many hours (years) trying nix, rust, Haskell, go, spring framework and all sorts of other things which are a lot of fun but not so good for getting shit done.
For other domains this doesn't apply of course; lower-level network stuff, portable CLIs, CPU-intensive workloads and so on are much better in go/rust but you can either integrate them with a network call, spawning an OS process or an FFI in the case of rust.
Its really not. You either tailor your dependencies to match the host (clue, you should be doing this anyway, it makes things much easier in the long run), or use virtual envs.
pyenv also gives you some (non obvious) flexibility as well
However, having said that, the chances are, your go or rust binary is going to be in a docker env as well(so is your python), so its basically all the same.
1. Put files on a server,
2. Start a process.
How much ancillary scaffolding you put around that is a matter of taste.Isolate the project from the python runtime it uses, and you'll always have the right set of packages installed.
None of this is perfect, but an in-tree .venv/ and convenience scripts seems to be the least-worst option.
If this is anything else, or if you work for a shop that hasn't embraced containerization, then you use PyInstaller (http://www.pyinstaller.org) to bundle your application. Either into a directory that contains your full Python virtual environment (only 5-10 megs!), or into a single executable file.
The latter is most convenient for a Go/Rust type experience. But the former will startup faster, because that single-file executable has to first uncompress itself to the system temp directory.
I meant to say "I look to see how they are going to deploy and use the virtualenv they have created."
Hint: what works in development -- just telling people to type "source .venv/bin/activate.sh" or such, doesn't fly in an unattended environment.
All that it requires, of course, is a bin/venv-python wrapper (bash) script to reference the created .venv/ directory, so this is hardly ground-breaking stuff, but as I mentioned originally, this (crucial) section is missed every time.
I use a minimal-but-complete pairing of venv and pip, and a couple of location-independent wrapper scripts, and I can run things the same across all environments.
> This is my very opinionated attempt to compile some of the best practices on setting up a new Python environment for local development.
So as I see it there are two types of tool, formatters and checkers. A formatter doesn't alter the Abstract Syntax Tree (AST). In other words it doesn't alter the meaning of the code, just how it looks (eg. the formatter Black). A checker on the other hand looks only at the AST and just gives warnings in your editor, and then it's up to you to change it or not.
The import order is fixed in the AST, so it falls outside the scope of Black (which never changes the AST). So I felt there was a need for a tool that worked as a checker that just gives warnings in your editor if your imports don't conform to PEP8, and hence Flake8 Alphabetize.
It's worth mentioning that Flake8 Alphabetize follows Black's philosophy of having only one way of doing things, so a project can standardise on Flake8 Alphabetize and everyone's imports will look the same.
Anyway, any feedback is welcomed:-)
Anyway, that's just my feeling at the moment, I'm sure others have a different take.
Even if the answer is "Poetry handles it" you certainly want to explain why they're important just like is being done in the rest of the "Why use..." sections.
I use this with my m1 mac, and it works great.
[1] -- https://docs.conda.io/en/latest/miniconda.html#linux-install...
You can conda install the CUDA dependencies and then install the required Tensorflow version via conda pip. But that's not much different to installing CUDA manually and then installing tf from system pip.
It's much faster and easier to pull a tf Docker image as it's their "officially supported" way to get up and running.
So... as a tf user and a sysadmin... Nah. No conda for me thanks.
E.g. https://github.com/conda/conda/issues/8087, https://www.anaconda.com/blog/understanding-and-improving-co...
Oh and if there is a conflict somewhere, it goes into some conflict detection routine that will take hours and not produce anything useful.
I could go on, but I have come to really dislike conda.
This is why mamba [0] was created. It is a C++ reimplementation of conda for much better performance. mamba is a drop-in replacement of conda and can operate on the same anaconda, condaforge (and mambaforge) repositories.
I use Gentoo and its package manager is written in python. Even though it is more complex (IMO) it doesn’t have nearly the same slowness when it comes to dependency resolution and conflict detection.
Just right now I'm trying to fire up a new instance on GCP. With a completely clean image, doing a conda install hangs for 30 minutes while it's trying to "solve" something.
It's very disillusioning to see how sheer twitter-followings and "popularity" type metrics drive development these days by forcing alternatives to be de-facto neglected. Everyone does what's "hot", so all the tutorials and bug reports and tests and SO questions and new libraries and and and all go towards that framework or language or tool or method. You can't even argue technical merits towards the neglected options because yes the popular tool is better, but only because we have a metric boat load (millions) of man-hours being pumped into making it better instead of all the alternatives. It's like the tech-equivalent of fashion fads in that it's self-reinforcing. Not to take away from some of the actual and technical achievements that some of these things have made, of course.
I supported a research cluster, and the amount of times conda caused an issue was a lot.
With all scientist packages, conda is pretty good but it is sooooo slow, it is just insane!
EDIT: Mentioning `pyproject.toml` and the relevant PEPs would also be great (i.e. PEP 621, PEP 517, etc.). Fortunately Poetry is compliant with these PEPs.
It's still early days in the Bazel ecosystem, so you won't find a ready-made solution, but the fundamentals are solid.
Using Cython adds another layer of complexity to the packaging/distribution that that link does not address. Fortunately now that you can specify build requirements in `pyproject.toml` Cython has become significantly easier to use on that front, but there are still some less than obvious bits to say the least.
Maybe I should publish an overview of Python best practices for publishing C extensions (with or without Cython).
"python.venvFolders": [
"~/.cache/pypoetry/virtualenvs",
"~/.pyenv/versions",
],
This will make VSCode automatically recognize virtualenv interpreters, so you can select them without starting code from an activated shell.The one thing I dislike about Python projects is that Python plasters the compile cache files all over the place. Is there a reason to change that? Currently I use the -B flag for all my scripts. But that makes it slow. I wish Python would have an option to perform like PHP and keep cached compilations in memory instead on disk. Or at least somewhere in /tmp/.
Also at what point do people just realize that all of this overhead is a gigantic waste of time and just use a better language?
However what makes Python a failure is that people feel they need this to dependably run a python program which only has pure-python dependencies.
Compare this to a language like Rust, or the NPM ecosystem. In those cases, the tools have managed to dependably encapsulate projects such that you only need the package manager to make a project fully repeatable.
With either of those ecosystems, there's basically one system dependency, and you can find any repository online and dependably do `git clone ...` then `cargo build` etc. to make it work. With Python, you effectively have to reproduce the original developer's system, and that is a failure.
Because if you don’t rely on Python packages with extensions that farm out to external libs it’s as easy as git clone, pyenv virtualenv, pip install -r, and python -m build.
1. virtualenv shouldn't be necessary. This is more or less the same concept as containerization. This is only needed because python has a fractured ecosystem, and setting up your environment for one project can break another.
2. you also have to know which environment encapsulation and package management solution the library author is using - this is not standardized
There is not much overhead in running a project in a container. The project has a setup file that turns a fresh Debian 10 into whatever environment it needs. And thats it. Run that setup script in your Dockerfile to create a container and you are all set. Want to run the project in a VM or on bare metal? Just install Debian 10, run the setup script and you all set.
Probably some time shortly after your developer time costs less than your cloud compute time. Until you hit that point (if ever) there are few options as cost-effective as Python.
Poetry does I expect a package manager to do, and does it well, especially when working with a team of developers on an application versus individually. There's not a compelling reason for me to use pip directly as a less functional alternative.
Can you describe an issue that you had by not locking transitive dependencies?
Lock files help solve for these. You can build software without solving them, but it makes my life easier.
Advantages:
- Separates development and production dependencies.
- The dependency version is specified separately from the lock file. In practice this means that the version in pyproject.toml generally only needs to be set to anything other than asterisk if and when it becomes necessary to use a specific version range.
- The lock file includes SHA-256 checksums by default, and these are checked during installation.
Disadvantages:
- More complex configuration than Pip.
- Python package managers come and go, and this one is likely going to suffer the same fate eventually.
- Introduces poetry.toml simply to specify that the virtualenv should be in the project directory. The default is to put virtualenvs in ~/.poetry, which is a non-standard location and therefore might interfere with typical IDE setups, mounting the virtualenv in containers or VMs, and the like.
[1] https://github.com/linz/template-python-hello-world/pull/106...*
That. The simple fact that a Pip file mixes both the packages you want and the dependencies required by this package, is a valid reason to switch to Poetry IMO.
Here is the config https://python-poetry.org/docs/configuration/#virtualenvsin-...
[0] https://github.com/samuelcolvin/pydantic/blob/master/setup.c...
Not only it's not needed, it creates a layer of "magic" the user has to understand on top of the environment.
And, if you really want to use brew, you can still continue using it, knowing it only has python 3.9 at the moment. Plus, it can coexist mostly peacefully with Macports.
Is incorrect, they just move older versions to their "@" nomenclature, for example https://formulae.brew.sh/formula/python@3.7
IIRC the trick is that the "at" versions don't get "brew link"-ed by default, since they'd almost certainly smash on top of the non-at binaries or manpages or whatever. Using `PATH=$(brew --prefix "python@3.7")/bin:$PATH` or ones favorite context switching gizmo will help, or (with the python ones specifically) using `$(brew --prefix "python@3.7")/bin/python -m venv ...` is a great way to avoid having a lot of special env-var silliness
I also add Docker or Songularity containers as needed for deep learning deployment but that is quite computing platform and application specific.
https://github.com/jazzband/pip-tools
And you can still use standard setup.cuff and pip install -e unlike Poetry. Also, much faster.
Please help me understand why I would want to lock specific versions of libraries in my Python projects. I’ve worked extensively with Python for a number of years and have never needed a lock file for my dependencies.
However, this is how you would setup for a new project only and in a piecemeal manner, if you didn't already have some existing template.
Compatibility issues with your main dev tool would rule out using Poetry, in the same way running the latest versions of any software without reason is a rookie mistake. Chasing a higher version number is a jr / intermediate folley.
If you're developing python professionally, save your money and just pay for pycharm. The morning of fluffing with a dozen new tools which then have their own maintanence overheads is less cost effective than buying a product that gives you >70% of what the author has recommended and it does so in a generally consistent manner.
This blog post was a great reminder of how much pycharm gives me on a day to day basis, how people get lost in their +1 more tool mindset and how switching languages is an intial cognitive overload.
There's also a simpler template repo[2] with almost all of these.
Then I disable the use of virtual environments in Poetry, because it's a useless overhead.
Then for updates, I just change versions in my Dockerfile or .toml and rebuild the container from scratch in a few time, which is cleaner than manual updates for everything IMO.
I'd prefer to only use the pre-commit versions of these libs, but then I'd sacrifice editor integration.
So if I want to use Python3.9 i go python3.9 -m venv .venv, if I want to use python3.6 i go python3.6 -m venv .venv etc.
Yes, lockfiles are great, but while everyone can execute the above commands, it is a lot harder to convince everyone in my organization to switch to poetry.
I think it's not the only reason. Limiting columns forces developers to read more vertically, which is less tiring for eyes. For example, try to read a book on a 27' monitor with no max width. Going to the next line will quickly become painful.
Less so for CLI style projects that read and write to the local file system. Yes I know you can bind mount directories but it’s clunky and you also usually have to fight with file permissions issues.
Which can be seen by this hilarious sentence:
> By default, Python packages are installed with pip install. In reality nobody uses it this way. It installs all your dependencies into one version of Python interpreter which messes up dependencies.
Why hilarious? Because everyone I know uses `pip install` if they're not using poetry. And because poetry is kinda new, that's a lot of people.
In Ruby you've got rbenv in place of pyenv.
Pytest is excellent and I'd recommend it.
But literally everything else is optional or completely insane. This is not "best practice," its the entire kitchen sink. You might use a linter, or you might just try to write clean code. You might use Poetry, but frankly, the default pip installer is fine for most people most of the time. Installing and configuring something to sort imports for you is just silly, and pre-commit hooks are a terrible idea that just get in the way.
My recommendation would be to avoid adding tooling until you have a problem it solves, and focus on building something useful instead. I imagine that's true of most languages.
I thought you just needed poetry?
I've been using conda, but have begun contemplating a switch to poetry. So what am I missing?
* Keep test frameworks out of production (in case they've busted something; the two points below can happen to test libraries too.)
* pip recently changed the way it's resolver works[1], breaking numerous projects. Yeah yeah, it's a major version bump, but lots of containers and environments installed pip latest just sort of assuming that it would always handle requirements.txt the same way. With that contract broken, now using pip directly isn't a non-decision anymore.
* Locking specific versions of dependencies can you roll back Bad News in production. Let's say you have a project with library A which has dependency B. If A asks for the latest version of B, you can be in a situation where a new version of B breaks your project and you won't know about it until you do a release _and_ if you try to roll back you'll still be in trouble. We recently had to deal with a similar issue and had to fall back on an older container until we could figure it out.
[1] https://pyfound.blogspot.com/2020/11/pip-20-3-new-resolver.h...
* if your code is under src/ and your tests under test/, you use a requirements-test (or better, tox.ini) your tests including the dependencies like pytest dont end up in a wheel
* if a container depends on 'latest' and not some semver major number, it's 100% the containers fault when they blindly update to a new major version
* 'pip freeze' is your friend
You can choose to build your own dependency management practice around pip, or use one someone else has already created. I think that it is easier to get a team on the same page with something like Poetry, especially if they're used to bundler or npm.
You're right about the docker containers as well, but upstream does what it wants and downstream has a strong tendency to not mess with upstream's choices.
So it’s back to plain old venv + requirements.txt for me
my problem is that venv + requirements.txt is just plain simpler to use.
I've found this is almost great, but when it breaks down, perhaps because two packages request a transitive dependency in different ways, or if the resolver really wants to install a new version of something that won't compile on your system, it's a giant pain and sometimes impossible to work around.
These days I just use pip with venv unless I have a good reason not to.
Given the current situation in Python, where there is little development, the old boys have totalitarian control, and new contributors are smart enough to avoid that mess:
Try out Rust, Go, Elixir or Lua instead. It might save you a lot of trouble. Heck, if you are willing to put in a lot of time to create carefully written objects, C++11 code can look a lot like Python (if you are into that).
"What data is collected. The GitHub Copilot collects activity from the user's Visual Studio Code editor, tied to a timestamp, and metadata."
PyCharm, Emacs, Vim, are all better and don't spy.
The language server Microsoft built for type checking and completion (https://github.com/microsoft/pyright) is excellent. It makes Python feel like a first-class statically typed language, which isn't something I was able to replicate in Pycharm (tried this very recently). Meanwhile, it's just baked into VSCode, no configuration needed.
I use Jetbrains tools for several languages, but for Python and frontend (Typescript with React/Vue/Angular) VSCode is hitting the perfect notes.