How to create a Python package in 2022
mathspp.com
mathspp.com
If I can be forgiven one nitpick: Poetry does not use a PEP 518-style[1] build configuration by default, which means that its use of `pyproject.toml` is slightly out of pace with the rest of the Python packaging ecosystem. That isn't to say that it isn't excellent, because it is! But you the standards have come a long way, and you can now use `pyproject.toml` with any build backend as long as you use the standard metadata.
By way of example, here's a project that's completely PEP 517 and PEP 518 compatible without needing a setup.py or setup.cfg[2]. Everything goes through pyproject.toml.
[1]: https://peps.python.org/pep-0518/
[2]: https://github.com/trailofbits/pip-audit/blob/main/pyproject...
However, I took a look at PEP 518 and failed to understand what was wrong with Poetry's default configuration. Can you help me out?
IIRC, there's work ongoing in Poetry to allow it to support PEP 621.
Using pyproject.toml with pip / flit still has many rough edges such as pip being unable to install deps locally for development or not generating lock files. Poetry is way more mature IMO.
The font size is extremely small to the point of being unreadable on a mobile phone.
If you want to keep a similar style: https://fonts.google.com/?category=Handwriting
Also: https://www.pagecloud.com/blog/best-google-fonts-pairings
We start with learning that we absolutely need this Poetry thing because… it's what everyone else uses. It's refreshing to see author who can skip usual badly argued justifications and just plain admin that he does not know shit and is just following rest of the herd. Then we continue by "solving" depependencies by usual way of ignoring them and just freezing whatever happens to be present.
Then there is inevitable firing up of virtualenv, because that's just what you have to do when dealing with messed up dependencies.
Next one is new to me. Apparently, one does not just set up git hooks nowadays but use separate tool with declarative config. Because if you ever happen upon something not covered by the Tool, that would mean you are no longer part of the herd.
Then we push our stuff straight to pypi, because of course our stuff can't possibly have any dependencies outside of python herd ecosystem. It's not like we knew our dependencies anyway.
Then comes the fun part, pulling in tox, because when you have special tool to handle dependencies, what you just need is another tool with different environment and dependency model.
Code quality section I will just skip over, seeing what pass for code quality these days makes me too sad. What follows is setup of several proprietary projects that modern opensource seemingly can't exist without. What is more interresting is "tyding up" by moving code from git root to subdir. Now, this is of course perfectly sensible thing to, but I wonder why is it called 'src'? Maybe some herd memeber saw compiled language somewhere and picked it up without understanding difference between compiled binary and source code?
Now don't take this as if I have problem with the article content in itself. No, as a primer to modern python packaging it's great. It's not authors fault that his work is so comprehensive it lays out bare all the idiosyncrasies, herd mentality, cargocultism and general laziness of python ecosystem these days. Or is it?
you'd think herd mentality might help it but it only creates more packaging solutions.
Now days, I've stopped using python outside of tiny scripts and I will never touch it for a large project.
There is exactly one thing I miss from Python packaging tools: developer mode. I can factor out parts of an application into a library and develop both at the same time by installing the library in editable mode and pointing to the library's local directory. This is something I've always wanted but never had in every other language I know. Only god knows how much time I spent trying to do exactly this with git submodules.
https://pkgdocs.julialang.org/v1/managing-packages/#developi...
It’s packaging tool is top notch. Maybe because it’s developed as a core part of the language.
> Maybe because it’s developed as a core part of the language.
Always thought it was strange how libraries and packaging never seem to be considered part of the language. My favorite example is Scheme: a beautifully minimal language but with no library support and the result was endless fragmentation due to unportable implementations.
pip freeze doesn't pin transitive dependencies and so you have to pick something and Poetry is fine and actively developed.
> virtualenv, because that's just what you have to do when dealing with messed up dependencies
No that's what you do when you have multiple dependency trees for different projects on your system. Somehow people got the message that global variables were bad but still think that "random bullshit strewn on my specific system" is a great way to make software that works on other people's machines.
> Because if you ever happen upon something not covered by the Tool
You write your own hook because it's entirely plugin based.
> tox, because when you have special tool to handle dependencies, what you just need
A tool that doesn't pollute your development environment with testing packages and doesn't run your tests in your development environment, hygiene that before this tool basically nobody bothered to do because it was tedious.
AFAIK it lists all installed packages and hence pins all dependencies.
I guess it's great if you're just looking for a shortcut to push something up to pypi, but my guess is someone new to it won't really understand what's going on other than some vague sense that they're following "best practices".
And then I imagine that same person will go on to write another article like this, and on and on we go!
The sentiment behind your comments is shared, but I don't see the need to sarcastically rant about it and rail all the suggestions OP made.
If anything, I'm surprised someone with more experience didn't see the post for what it is, and attacking someone's post like this just shows immaturity when you could have easily taken those opinions and formed a constructive argument or given good advice.
Pre-commit is one of the most annoying tools that have come into existence in recent years that everyone seems to be cargo-culting. It doesn't play well with editors since in order to find the actual binary path, you'd have to open up a sqlite database to fish out the virtualenv pre-commit created. Pre-commit also increases maintenance burden, since its configuration is completely separate from your usual requirements-dev.txt/pyproject.toml/setup.cfg etc. If you have dev dependencies in one of these files because making your editors to find the pre-commit created binaries are hard, now you have to keep both versions in sync.
I really don't see the point of any pre-commit hooks unless you are the one guy that doesn't use a modern CI/CD platform.
Running the tools locally is basically about tightening the development loop. Many of the commonly used tools (e.g. black, isort etc.) actually make the changes to the files so you'll never even commit failing versions. Do you really want to push changes to some remote CI system only to be told it's failed some boring QA check? There's nothing at all stopping you from doing that. Pre-commit is completely optional for each developer. I would just recommend it for sanity reasons.
Yes, that's what CI/CD is for.
> Pre-commit is completely optional for each developer.
No it is not, it's installed at your git pre-commit hook, and it gets run every fucking time I commit, even if I've set up my editor correctly to auto-format and lint everything at I develop, meaning 99.9999% of the time sure the code I commit will pass all these linter checks. I can use git commit --no-verify to bypass pre-commit, but then again, what's the point of using pre-commit in the first place if you need to bypass it? There is absolutely no point of linting twice locally and thrice in total just to hit the first stage of your deployment.
rm ./.git/hooks/pre-commit
Or just do not install the hook in the first place. Choice is still the developer's unless your team is doing something else here that limits your ability to delete/rename files locally.2. pre-commit install installs the hook by default.
3. If you happen to fail linting a few times a year, there's always an anal coworker telling your boss in a 1-on-1 you are not following some BS ways of working.
No, it doesn't. Pre-commits are just a DX bonus so you won't have to wait for the CI/CD server.
2. pre-commit install installs the hook by default.
Am I missing something? If you run "$ pre-commit install" of course it will install the hooks. Do you mean that poetry install, setup.py themselves are making the magic of installing the hooks themselves? If so, it's indeed a bad practice.
3. If you happen to fail linting a few times a year, there's always an anal coworker telling your boss in a 1-on-1 you are not following some BS ways of working.
I don't get why anyone would care, unless you're polluting the CI/CD history with tons of tiny commits all day along just to run linting.
Of course I don't know GPs specific context so they may have other details, but generally it's
Over my career, I see that every company has some version of:
1. New guy joins a team, wants to push code ASAP
2. He sets up his favorite editor and punt configuring linters and formatters
3. Code pushed, CI fails at linting
4. Some well-meaning coworker or new guy suggests some variations of husky/pre-commit/fancy git hook scripts
5. Team agrees and put that into the repo
6. 6 months to 1 year later the entire team realizes its benefit does not out-weight the cost, and unanimously agree to rip it out unceremoniously.
Maybe you have experience with some weird tool? Or running black and autoflake messed around with emacs and it needs to reload all the buffers?
If you have it all sorted out from Emacs maybe check the repo they posted which tries to do all this.
Not really. CI/CD is a methodology where a team immediately integrates new changes into a trunk/release. The CI/CD pipeline is there to give the team the confidence to merge your stuff. When I see people constantly pushing breaking code into a CI pipeline I see an incredible amount of wasted time and shared computing resource. Especially if it's some trivial formatting check that you should have already done locally. You do what works for you, but tools like pre-commit were invented to save time and effort and they work well.
> I've set up my editor correctly to auto-format and lint everything at I develop, meaning 99.9999% of the time sure the code I commit will pass all these linter checks.
Then what is the problem? Have your team installed hooks that take a long time to run? Even on larger codebases pre-commit adds a negligible amount of time to each commit, unless perhaps you've touched every file in the codebase or something. Honestly your gripe with pre-commit seems mostly irrational.
That's why linting is a part of every CI stage. Linters check your code for bugs.
> Have your team installed hooks that take a long time to run?
Yes, they are called tests.
If your pre-commit takes more than 3 seconds to run it’s set up incorrectly IMO, and should belong to a more manually (and CI of course) invoked test suite instead.
One thing that's really annoying these days are CI/CD that can't be replicated locally, generating quite annoying delays in the development. Jenkins seems particularly problematic in this regard: the steps get encoded in some cryptic pet Jenkins server, and then you have to wait minutes until an agent picks it up and reaches the step you actually care about. Other tools are a little quicker, but still...
So, I think at the very least pre-commit hooks help with this "over-reliance" on the CI/CD server. It's so much better DX when you can run parts of the pipeline instantaneously.
pre-commit just runs your linter, formatter and tests. Surely you can fully replicate this step locally. Just run make lint both locally and on CI.
Anything else that's hard to replicate has to do with the distributed systems that you are probably working on because these systems are all probably proprietary stuff that live on the cloud.
> Although I'd indeed prefer just before push, or just make display warnings
Wholehearted agree. This is a compromise I can live with.
And, keeping things separate from setup.cfg or pyproject.toml is optional: The tools still look for configuration in their usual places, so it's still possible have your black options in pyproject.toml and just a bare-bones entry to call black in your .pre-commit file if you prefer.
You also don't have to run the hooks at pre-commit time. Just don't hook pre-commit into your checkout. The pre-commit tool can also be configured to run its checks at a different stages than, well, the pre-commit stage:
https://github.blog/2022-02-02-build-ci-cd-pipeline-github-a...
Strictly speaking, any tool or set of tools that allow you to trigger building & deploying/publishing artifacts in response to source control commits can be used to build a CI/CD pipeline. One could write bash scripts linked to a cron job that pulls a remote repository every n minutes and then performs some scripted actions to integrate changes between branches before building & publishing the artifact to a local SFTP server.
If you prefer a more mature solution with better documentation however, there is a (non-exhaustive) list of CI/CD tools on this awesome-devops list:
https://github.com/wmariuss/awesome-devops#continuous-integr...
---
edit: I saw gitlab on the list and realized it is probably the closest self-hosted equivalent to the github option I mentioned previously
You can also set things up the other way around, having the CI/CD system install and and run pre-commit's checks as a build step. Pre-commit provides a nice framework, I find, for running these checks.
The advantage of pre-commit is that it catches mistakes before they are committed. Most CI/CD systems are setup to only validate the tip of the branch (i.e. the last commit in the PR), not all the commits along the way. Yes, you can configure a CI/CD system to test each commit, but it's usually swimming upstream to do so. And yes, you can squash all of a PR's commits into a single commit, but there are good reasons NOT to do that. So assuming you have multiple commits in a change, it's nice to know they have likely all been validated in a project using pre-commmit.
I'll make an appeal to authority here:
I've been a professional developer for decades. I've worked with a variety of VCS's and build systems. I've written plenty of Makefiles. Nonetheless, I still find pre-commit useful. I use it even in combination with a Makefile sometimes. You'd be horrified, I guess, to know some of my Makefiles have a rule which runs `pre-commit run --all`.
As an example, earlier this week I setup a new repo which build packages an AWS Lambda written in javascript, and deploys it using the AWS "sam" CLI via a CI/CD system. So the deployed code is Javascript. The repo contains shell-scripts to assist with deployment. And there are multiple yaml files. There's a yaml file to configure the CI/CD system, and there's the CloudFormation template files.
Here's what I configured pre-commit to do:
1. Check for whitespace nits.
2. Check for syntax errors in the yaml files.
3. Validate the CloudFormation templates.
https://aws.amazon.com/blogs/infrastructure-and-automation/u...
4. Run shellcheck against the shell scripts.
5. Run eslint and prettifier against the javascript.
6. Run "npm test" as a local step.
Normally I'd leave (6) out because in most projects its too time-consuming, but in this project the tests run quickly enough that I just made it a pre-commit check. The "npm test" step runs jest, which is installed as a dev dependency in the project's package.json. The other tools are all installed by pre-commit itself.
> I really don't see the point of any pre-commit hooks unless you are the one guy that doesn't use a modern CI/CD platform.
I've tried to make an argument for why to use pre-commit above. Nevertheless, you don't have to install pre-commit's hooks in your checkout. They are ideally there to save you time having to correct mistakes after the fact. It sounds like there's an impedance mismatch between your personal workflow and how you've seen pre-commit set up. Perhaps by resolving that mismatch by adjusting the pre-commit configuration, you can enjoy pre-commits benefits w/o experiencing the issues you've run into.
However, in my experience, I've seen plenty of pipelines setup to run lint on every push on a PR branch, which is effectively only checking outgoing changes before merge, it's just in this case it's merging to your feature branch. My point still stands - as long as linting is done on CI, and you've set up your editor to lint as you edit, you don't need pre-commit.
I'm not entirely sure what point you are trying to make. It sounds like all you've done is moved all of the tools you'd call anyway from a Makefile to a YAML file.
Having done it both ways several times I lean pre-commit for now
Thank you for linking it! Yes, this will be a huge convenience and security win for the large number of packages that use GitHub to release new versions.
However, it drinks the code coverage cool-aid that started like 30 years ago when code coverage tools emerged.
Management types said "high test code coverage == high quality"; lets bean count that!!
A great way to achieve high code coverage is to have less than robust code that does not check for crazy error cases that are really hard to reproduce in test cases.
Code coverage is a tool to help engineers write good tests. One takes the time to look at the results and improve the test. It is a poor investment to be obsessed with code cover on paths where the cost to test them greatly exceeds the value.
10% coverage and 100% are both alarm bells. Don't assume naive, easy to produce metrics are the same as quality code.
Otherwise, and excellent article.
Combined with thoughtful use of `# pragma: no cover` a 98% code coverage nowadays is an immediate warning that something was rushed. With this and type checking I feel RuntimeErrors much easier to avoid these days.
And typing, not even a mention?! :) But otherwise a great article, thank you!
Coverage is a decent (among other things) measure unless it becomes a target. Once it becomes a target you get shitty rushed tests that act mostly as cement surrounding current behavior - bugs and all.
I can see it's a non goal then you have access to deployed code and Sentry. But as library author or author of customer apps there is no other way around.
... Have to admit: We recently ended up contracting conda packaging out because it was nowhere near clear enough to make sense for our core team to untangle. Would love to see a similar tutorial on a github flow packaging & publishing to conda. Still no convinced we're doing it right for subtleties like optional dependencies: equivalent of `pip install graphistry` vs `pip install graphistry[ai]` vs `graphistry[umap-learn]`, etc.
Personally I never use either of these tools.
The contractor still took a couple weeks to figure out and get up. I assumed it'd be one evening for initial bulk as we already had setup.cfg etc, but after searching conda tutorials.. not surprised. Our ~final meta file is pretty simple, so no idea why the docs are so indirect.
Also, can't believe everyone let me get away with not writing about documentation! I'll see to it that it gets done and added to the article.
For Node, it's quite simple and even built into npm. Also the version is only part of the package.json file. For Python you probably have your version somewhere in __init__.py, and I always end up writing ugly bash scripts that modify multiple places with sed.
Is changing the number 2 places really that big of a deal?
You should have a release checklist anyway, with steps like sending an announcement email or tweet etc. How much time does this really save, at the cost of so much complexity?
I maintain half a dozen small Python packages. I don't do emails, tweets, etc. I just want to create releases easily when there's a bug fix. It not only saves time to automate this step (I can use the same script to release each package), it also means you can't forget things. Before I had a script, I always forgot to push the tag, or run the changelog, etc.
The term "python package" means something entirely different (or at the very least is ambiguous in a pypi/distribution context).
To add to the confusion, creating a totally normal, runnable python package in a manner that makes it completely self-contained such that it can be "distributed" in a standalone manner, while still being a totally normal boring python package, is also totally possible (if not preferred, in my view).
(shameless plug: https://github.com/tpapastylianou/self-contained-runnable-py... )
Python is a mess.
Just because this mess happens in some other languages doesn't mean it's the right thing. Having a very fragmented community is not a good thing for a beginner. Also, npm is far more of a de facto choice than poetry, which is still better than the state python finds itself in.
I started using tox-poetry-installer[1] to make tox pick up pinned versions from the lock file and reuse the private package index credentials from poetry.
[I have no ties to this company and have never applied there.]
[tool.tox]
legacy_tox_ini = """<tox.ini content here>"""I don't have a blog post but you can see the process on my personal project https://github.com/DontShaveTheYak/cf2tf
Check out the merged PR's and the GitHub actions.
I even do alpha releases to test pypi.
Another good tool (which was endorsed by the PyPA) is Hatch - https://hatch.pypa.io/latest/environment/
I currently use PDM because it supports conda virtual environments for isolation, but am keeping an eye on Hatch.
Are there any similar resources for setting up internal packages that you don't intend to publish publicly?
I can think of a number of situations where I would have benifitted from it, but the process of configuring a package to publish, hosting it and then pulling it when necessary is a mystery to me.
https://packaging.python.org/en/latest/guides/hosting-your-o...
https://python-poetry.org/docs/dependency-specification/#pat...
In other words, this a more thorough explanation of my point from [3]:
> Yes, if you ran poetry install you could get editable mode, but that requires every single end user of your package to install poetry and explicitly invoke a poetry install just for your package. And if someone else's package uses a different package builder, now they need to install and invoke that one. And on and on, and you end up with a six-hundred-line install script because of all the one-off "must install this developer's favorite package manager to use their package" stuff.
[1] https://peps.python.org/pep-0660/
[2] https://github.com/python-poetry/poetry/issues/34#issuecomme...
[3] https://www.reddit.com/r/Python/comments/t3p3ub/comment/hyum...
[1]: https://cjolowicz.github.io/posts/hypermodern-python-01-setu...
I got src directories in the Elixir dependencies written in Erlang, in about 12% of the node_modules used by a React project and in the few C extensions I'm using for Ruby.
May I conclude that src is uncommon at least in scripting languages? (Elixir is compiled.) Maybe the reason is that there is only source code and there is no need for a separate directory for a build / dist.
https://blog.ionelmc.ro/2014/05/25/python-packaging/#the-str...
It describes some of the outdated motivations for the other layouts commonly seen with python, as well as the many benefits of the “src” layout.
So if you have a project called "splitter", your source code really lives under `splitter/splitter/`. I would agree that seems a bit redundant and `splitter/src` look better, but the source code is not in the project root.