Common Infrastructure Errors I've Made
matduggan.com
matduggan.com
100%. I stopped writing anything that had to be deployed (basically everything except Jupyter notebooks for data stuff) in Python because it’s truly a nightmare. Go and goreleaser is great for writing a CLI (and if it’s public, it can auto generate binaries and upload to GitHub, create a Homebrew/Scoop bucket, etc)
From my personal experience, there's so many things that can crop up, not limited to:
- Conflicts with system packages from e.g. apt
- Conflicts with system packages from easy_install/pip
- Mismatching Python versions (even among 3.x with syntax additions only available in newer versions)
- Different dependency management systems (pip's requirements.txt, setuptools's setup.py, poetry's pyproject.toml, etc.)
- ... and the list goes on if you start talking about C extensions!
Additionally, not everyone does things in the same way which means the easiest way to package Python tools tends to be just bundling Python + its entire standard library together. Contrast this with Go, where `go install` does basically the right thing 9 times out of 10. I'm not a huge Go fan, but the convenience of `go install` is really unmatched.
I mean if you upgrade from x11 to wazland and zour app depends on x11 then sure, but I still use binaries that were compiled in 90s (some because I dont have the source)
And windows is also very backwards compatible.
Python is hell for the latter. I support Windows and Linux, so I can bend pyinstaller to my will with enough effort, but I need to have a special Windows VM to "cross-compile".
For the former, though, it's a matter of creating a setup.py and telling the user to pip install the repo on GitHub. Or you can use Poetry, which makes things slightly easier.
There's some python magic for determining the version, but if you use a feature not supported by the version it raises an uncatchable exception...because python's parser has control before your code has control.
This isn't the only deployment issue, but a good example of what you're in for.
In short a package manager + git is the easiest way to distribute and maintain software ... once you know how to use them. Configuring in a new environment? There is nothing that works reliably, other than maybe a full blown human level intelligent agent that is a senior developer with direct access to poke the system until it works.
Further, developers with little time or experience packaging software for regular users may balk at what they rightly perceive as a route that is more difficult and creates more work in the long run because as a dev you can't just drop a "fix pushed" message your messenger of choice and go back to work.
Static linking seems like it is a solution until you have to distribute to multiple operating systems. Only things like cosmopolitan libc have been able to overcome this, but even that requires a binary blob.
In long ....
I maintain a dozen or so python packages. Small stuff, mostly used by myself and my research group. At some point I needed to make some of that functionality available for other devs and some non-technical users. It was and is a massive pain.
Why was it such an issue? I think it is because many of the quick internal scripts were originally written assuming that only technical users as advanced or nearly as advanced as the author would be using them, and that any such user would be able to use git, configure their environment, and generally keep their system up to date using a package manager.
Aside. Add to that the fact that my dev environment is Gentoo, which has one of the sanest python environment management solutions available, and every time I step outside it or get reports from the poor lost souls trying to run my code "elsewhere" I enter a world of sanity destroying chaos. The workarounds that I created to try and mitigate some of the issues were poorly designed because I tried to YAGNI without realizing that you inevitably gonna need it (YIGNI?) and that the technical debt surrounding issues with interfacing with the environment is 10x more time consuming than anything else, especially if it is in the form of reports from souls lost in the warp.
Even if a dev can use git and a package manager they can still get stuck configuring the environment. Seems easy until there are secrets to distribute or someone needs to run something in a sightly different environment and nothing works.
For non-technical users it can be a stretch to get docker set up correctly, especially if they need X11. If I can get ssh access to the machine then there is a chance, but that doesn't scale beyond maybe 2 or 3 users. I have seen enough of this to have started to seriously consider just learning how to create installers for macos and windows and build them via CI. The long term solution is to switch off python completely, but the legacy code will continue to haunt us for a long time.
I've learned docker to try to create controlled environments. I've looked at statically linking SBCL to musl libc. I've looked into cosmopolitan libc. I've explored trying to use Emacs, elisp, and org-mode as another possible route to bootstrap stable and predictable environments. I've looked into Gentoo prefix. I've wondered whether I could get users to install Pharo, but realized that it is already a tossup whether they can get docker running. I'm so sick of having to repeatedly do searches in high dimensional configuration space (fight with the environment) just to get some random code to run for a user who can't deal with the configuration space themselves (and it's not their fault!). At this point it would be easier to wire all my users up to a mitogen botnet ... except that some of them use windows!
Telling a user "install python" or "install x" and then run this code leaves out the "and resolve the 10 completely novel errors you will encounter along the way." This is because there are at least 5 ways to "install python" (oh, don't forget the even bigger nightmare that is "install pip") on any given operating system (forget distros [0]), and each one results in a setup where things like data or configuration files are installed in different places with no obvious way to figure out where those places are, or you are forced to resort to a runtime lookup that causes CLI program startup time to explode to hundreds of milliseconds, making the entire point of having a quick CLI script moot.
Hello world may be the easiest program to write but it is by far the hardest program to run. Dependencies? Data? Network? Good luck.
If I have to do weird things with data structures or complex logic, I reach for Python, but I do it one of two ways: 1) so stupid that "python foo.py" works (no external deps), or 2) publish it to pypi so the user can just 'pip install foobar'.
What I personally do is write POSIX shell, then add in things like arrays, $BASH_SOURCE, printf %q, and a few other handy features that have been around for a while. To test it, there is an official Docker image for Bash v3 (https://hub.docker.com/_/bash).
But! Almost nobody gets it. If you're into operations: "here we write our tool chain in Python". If you're into dev: "here we write our tool chain in js/go/php/etc/etc Bash would just be a +1"
+1??? You still need to write bash, or shell in general, for orchestration of build/ci/cd pipelines. If someone thinks that in 2021 they can escape learning shell, they are out of their minds. So no, it's not a +1.
Bash is the language that is not only available quite everywhere but it's also the one that allows you to write oneliner solutions https://twitter.com/andreineculau/status/1466107915450916867... and move on.
As devs we should spend 10x more on solid foundations that work as many of us, for as many years to come, than we do on hip-new-thingie-that-promises-to-fix-all-my-problems. It's as if we took the regex joke and our brains could only comprehend: I'm never using a regex. Keep it simple? Never.
Heck no. Bash is a terrible language by modern standards, with obscure syntax, many weird quirks, no type checking, no return values... How could developers program in Bash before Shellcheck!?
I think that internally, I draw the line at around 100 lines. Even refactoring a program shorter than that it's not a pleasant experience (not difficult, but requires a great deal of attention, when compared to any modern language).
You're right about them needing to be short. Ideally every program should be short. But scripts are especially supposed to be short, because they're never supposed to get complex (or be fast/efficient); that's when you need a "real" program. Even the bottom of the Bash man page says about itself: "BUGS: It's too big and too slow."
(fwiw, back in the day we didn't use linters, we just had to write it all correctly the first time. punch card typos are a memorable lesson in checking your work)
The fact that every language has a variable amount of a certain property doesn't mean that the amount is the same for all the languages. The amount of domain knowledge required for simple programs is considerably higher in Bash, than any related language; even iterating the lines of a string can be unintuitive to casual users.
Bash is weakly typed, and weak typing is considerably unsafe (bug-prone) than any other related language (e.g. Python, which is has taken has share of shell scripting in modern operating systems).
Returning values via stdout is a workaround. If other parts of the function send messages to stdout, one needs to use a combination of redirects, which the vast majority of the programmers don't know. Calling the ability to return a value from 0 to 255 support for return value is a big stretch.
I think that Bash is a necessity, and I frequently use it (and it's fun-ish for small scripts), but it definitely doesn't fit the definition of "easy", being considerably more error-prone and unintuitive to any comparable language, possibly even when compare to Perl.
Based on the idea choose the right tool for the job it seems python is not right for this job.
An alternative would be to write in Perl5, perl is pretty much fixed, these days (if you insist on not using containers)
Try to compile your go tools in year, or two. If think that you will be hard pressed to find the dependent library versions. Packaging is not the strongest part of go. A version in your package translates to a tag in the referenced git repository, now sometimes people do funny things with their repositories, and older tags disappear.
Python and Perl5 are better, at packaging. At least you have all of the older library versions on https://pypi.org/ I think, that the decentralized approach to packaging (as used in golang) is far from optimal.
Maybe this is all due to the fact, that google is using a mono repository, internally for their own code. So that they are not eating their own 'dog food'.
Sorry for ranting, judgements like 'right tools' really get me going..
I think GP was just saying that most times, fewer moving parts is better. Avoiding Docker is great if you don’t need it (especially for a CLI tool). Obviously there are trade offs that must be evaluated in the context of a given problem and domain.
you may be, but the point is that the time you save writing in python is paid multiple times when each person who wants to use your code struggles to install it and get it working.
Mainly because volumes are mounted as the host user UID, while the container might be running any other user, often root.
Quite baffling they haven’t made this easier since it would be a real killer feature unlocking even more potential from containers.
Another use case is running a http server locally, for something that has a graphic ui. Like one of my projects here: https://github.com/mosermichael/s9k
Using the zipapp module, it is possible to create self-contained Python programs, which can be distributed to end users who only need to have a suitable version of Python installed on their system. The key to doing this is to bundle all of the application’s dependencies into the archive, along with the application code
Venv+pip or pipx are also fairly good if you have a pip registry.
And there's the problem
Just keep them up to date with changes and always use system provided packages, and there is no issue.
A lot of the advice is good but I take an issue with this one. With poetry and docker, packaging Python apps for easy consumption is a non-issue. Same for Ruby. If you can get your team to standardize on poetry, you might not even need containers -- but, honestly, running these tools from CI or automation (anywhere) is so useful that you probably want container versions anyway.
Golang is not a good fit for exploratory CLIs that work with complex data structures and are written for one-off, low-CPU consumption purposes -- not for scale-up API services. Just having an interactive shell (or `pry` in Ruby, those two are identical for the purpose) saved me probably weeks of time. Trying to unit test any moderately complex API surface brings me to tears when I compare it to trivial object mocking in something like Ruby.
Python (or Ruby) are ideal for this and have excellent frameworks for CLI tools.
Python gets more and more complex the further away the people running the tool are from fancy recentish distros (Fedora, Ubuntu, Arch) or special constrained environments (Nix, running the CLI in docker). The moment you have to deal with unspecified RHEL version (6 is still reasonably common) or derivative, or Mac or Windows, kiss any expectation of python packaging being nice "bye bye".
Unless of course you have a platform team that can handle the packaging and distribution for you, but then it probably falls a bit under "constrained environment".
I really, really can't say that about Python.
This is why you shouldn't write CLI tools in Python, you need frigging docker to package them
Given how much easier that was than getting people to use conda/pip (conda is better for DS stuff as it handles C-level dependencies), I completely understand how people just suggest Docker.
Like, if you're doing pure Python lots of this may not be necessary, but as soon as you start having multiple C-level dependencies pip breaks (and pip didn't even do dependency resolution till last year(!)).
Yeah, apps it's worth doing this for, cli tools it's not.
Writing and deploying Python tools is easy if all your dependencies come in distro packages.
... I might have somewhat similar scars to TFA author, I guess...
Presumably this is known to all the people who've been doing Python for a long time, but it bit me in the ass relatively recently.
Problems start when going elsewhere :/
There's no obligation to support every potential platform your script might ever need to run on until that need actually arises. You just need to define what platforms you want to support and work within their constraints. Unfortunately for lots of software that platform is essentially "whatever the developer managed to install on their system".
Constraining yourself to support a particular platform might mean there will be libraries you will not be able to use even if they might be useful, but that's the tradeoff you make when building software against a stable platform.
The bad part is the Setuptools documentation, which is slowly improving, but is (and has been for years) so bad that almost nobody can learn from reading it.
>>Packaging python is an exercise in frustration >>Setuptools documentation [...] is [...] so bad that almost nobody can learn from reading it
and these
>> until you become a level 3 wizard >> It really isn't that hard to package for Python.
as being fundamentally equivalent statements.
I find that the majority of the people who find difficulty with packaging are really just struggling with documentation and ecosystem fragmentation, as you hinted at. For example, that page I linked weirdly focuses on one specific tool (Pipenv) and doesn't attempt to survey the landscape of packaging and dependency management tools.
There tends to be a poor balance of "explanation", "example", and "overview" in these documents. I would say that I don't quite know why, but I know that writing good docs is incredibly difficult, so I have sympathy for people who've tried and didn't quite succeed.
One thing I'll offer is that, in the Python world, "managing dependencies" is distinct from "packaging". This seems to be different from other language ecosystems, where there is 1 package management tool that does everything.
That said, I take serious issue with the claim in the article that you quoted:
> Nobody knows how to correctly install and package Python apps.
This simply is not true. There are lots and lots of Python applications that are packaged and that you can install without a problem. You can complain about the need for virtual environments instead of having "app-local" packages by default, sure. But that's not the issue in most cases.
I think ultimately this amounts to FUD. Not because your experience is invalid, but because the blame is put on the wrong thing. The fact that Python packaging documentation is generally dogshit doesn't mean that the actual process of making a Python package is hard. If someone sat with you and showed you how to set things up the first time, I am confident that you would have no trouble cranking out Python packages that you can distribute to your engineering team without any problems.
> I wasn’t even able to find some kind of best practice to manage friggin dependencies
Again, this is a documentation problem.
> There’s like a myriad of package managers, all of them work differently.
Not really, Pip is still mostly the only package manager in town. But there are several tools like Pipenv and Poetry that replace the old Setuptools when it comes to actually building and installing packages, as well as generating lockfiles for dependencies (which Setuptools does not do at all).
I don't find it particularly bad that there are a few different tools to do similar jobs. I do however find it upsetting that the Python packaging document I linked doesn't even attempt to describe them or explain when/why you would use one or the other.
> nobody seemed to had something like redistributing an app to other people on their mind.
That's just outright untrue. Otherwise, PyPI wouldn't exist. Again, this sounds like a documentation problem.
> machine learning Jupyter notebook
The other issue here is that you're looking at a machine learning research script, not an "application" as such. Machine learning libraries tend to have complicated dependencies with bindings to C and C++ libraries that might or might not be bundled with the packages themselves. This would be and is a problem in any language ecosystem, e.g. Node or Ruby or Perl.
These are NOT easy applications when it comes to packaging.
Another issue with your scientist's code is that they were very likely using Conda, which is kind of its own thing. Conda is somewhat language-agnostic, and largely bypasses all the established Python "developer" tooling, in pursuit of somewhat-reproducible computing environments in support of reproducible research. Conda is very popular precisely because machine learning libraries tend to be so complicated to package correctly, and Conda solves that problem as a kind of portable alternative to Deb, RPM, et al.
Unfortunately, Conda poses some challenges if/when you need to take a project that lives in a Conda environment and transfer it to a more typical Python package, mostly because Conda environments can control things like the C compiler, while Python package managers cannot.
However, I will insist that this is not a Python-specific problem. If they wrote their application in Clojure, for example, you would have the exact same issue in migrating it from Conda to the usual suite of Clojure tools.
Oh, and speaking of Clojure: Python isn't the only language with the "which tool do I use and why are they all so complicated?" problem. I don't have a goddamn clue how to package for Clojure (I tried!), but I'm not going off a forum about how nobody should use Clojure for internal tools.
To be honest, having been in this situation it generally makes more sense to implement the pip dependencies inside the conda venv (by running pip), as conda can handle the C-level dependencies and pip is mostly perfect (modulo back when it didn't actually manage dependencies) for pure python stuff.
Also, it sounds like you know a lot about python packaging, please by the love of all that's holy, write this stuff down somewhere other than HN. There really isn't a lot of good docs around this.
> Also, it sounds like you know a lot about python packaging, please by the love of all that's holy, write this stuff down somewhere other than HN. There really isn't a lot of good docs around this.
Believe me, I want to! I am pretty bad at managing my "project/writing wishlist".
I peaked at the documentation, but found it both overwhelming and unfocused, thus hard to follow. As you say, it's focused on specific tools and feels more like a reference than actual instructions for what I was aiming to do.
> One thing I'll offer is that, in the Python world, "managing dependencies" is distinct from "packaging". This seems to be different from other language ecosystems, where there is 1 package management tool that does everything.
That might be the actual thing that stumped me: I don't intend on publishing my package somewhere publicly, but simply enable other devs on the team to install the required dependencies with a single command, without any assumptions on their environment, which seems awkward to me.
> That said, I take serious issue with the claim in the article that you quoted: >> Nobody knows how to correctly install and package Python apps. > This simply is not true. There are lots and lots of Python applications that are packaged and that you can install without a problem. You can complain about the need for virtual environments instead of having "app-local" packages by default, sure. But that's not the issue in most cases.
That is because we understood that quote differently. In your post you mention Pip, Pipenv, Poetry, Setuptools, PyPI, and Conda. Additionally I came across Virtualenv, and Pyenv. All of these tools are somehow related to python packages and dependency management, and everyone on the internet seems to have their own opinion one which combination is the best. There seems to be no consensus on one of the most foundational tasks in software engineering, namely sharing code with other people. To me that counts as nobody knowing how to it correctly. That said, reading the docs again, which clearly recommend using Pipenv - and that seems to be simple enough - I probably should have headed for the official docs right away.
> These are NOT easy applications when it comes to packaging. I noticed that :) Caused me to hit the wall multiple times, for example while using Alpine Linux/MUSL and wondering why the C bindings had to be compiled on every install. But you're right, that would bring its own complexities in every language and might be partially responsible for my general frustration with python dependencies. While the scientist didn't use Conda (I think he just happily installs everything he needs globally), that was the recommendation I found on one of his primary packages' website (I think it was pandas), so I followed through and tried to set things up using Conda. It was a nightmare, frankly.
I'm really not opposed to learning new things. And I get that my assumption that packages live inside an application directory by default might be just that, and colored by other ecosystems I'm more used to. And, as I said earlier, I'm definitely going to check out Pipenv and try to refactor that thing into something more standards-compliant. Still, I think Python places more roadblocks than necessary in your way, and the combination of bad docs, several opinions on how to do basic stuff, as well as the bad habit of global dependencies, make it difficult for newbies to deal with Python. Which, all in all, confirms the statement in TFA for me: Don't use Python for internal tools.
And frankly, as someone affected by Python packaging, I'm definitely going to complain on forums as much as I want if things are harder than they have to be. Just because other systems have their warts too doesn't mean there's nothing to improve.
"exploratory CLIs that work with complex data structures and are written for one-off ... purposes"
Sounds like a maintenance nightmare if you are actually "deploying" Python scripts that fit this definition. Yeah, golang will usually require a bit more boilerplate up front, but it is going to make your workflows infinitely more maintainable & flexible in the long term due to type safety + easy (and docker-less) packaging.
IMHO, if it's not a literal one-liner in ($SHELL|curl|awk|jq|*), it should probably be done right the first time in golang/etc or go back to the drawing board.
By the way, OP specifically mentioned "one-off", which you also quoted. Why mention deployment?
Because the article does.
As for complex data structure, I don't understand exactly what do you mean, as if a dynamic language would model that easier than a strongly typed language.
When you get a single binary for CLI it's hard to use anything else, the pain of using pip for anything python based.
Chef?
I don't know, I've generally had a good time with. You have to know how to do it and where the foot guns are (I'm still learning), but it's better than being comparatively gimped in development velocity using Go, never mind Rust or C++.
1) it gives your company a stronger bargaining position with the cloud provider.
Granted, my companies tend to have extremely high spend- but being able to shave a dozen or so percent off your bill is enough to hire another 50 engineers in my org.
2) you may end up hitting some kind of unarguable problem.
These could be business driven (my CEO doesn’t like yours!), technical (GCP is not supported by $vendor) or political (you need to make a China version of your product, no GCP in China!)
Everything is trade offs. AWS never worked for us because the technical implementation of their hypervisor was not affined to CPU cores of the machine, meaning you often compete with other VMs on memory bandwidth. — but AWS works in China (kinda). So my solutions support both GCP and AWS as slightly less supported backup.
It's neat having a serverless single page app that is hosted in s3, served through cloudfront, with lambda's that post messages to sqs queues that are read by god knows what else, but what happens when there's a bug? How do you test it? You can throw more cloud at it and give each dev a way to build their own copy of the stack, but that's even more work to manage. Maybe localstack behaves the same, but can you integrate it with your test framework?
I never took a hard "we must never use aws-only services" approach, but having the ability to run something locally was a huge plus. Postgres RDS? Totally fine, you don't need amazon to run postgres. Redshift? Worth the lock-in given the performance. Lambda? Eh, probably not, given that we already have a streamlined way to host a webapp.
One of the main reasons we never used google spanner was that we can’t test it locally.
Every startup I've worked at (and I've been at this for 15+ years) has moved hosting providers, but I still wouldn't put it high on the list of reasons to avoid vendor lock-in. If you make sure someone(s) know how the app actually runs, and you try to pick stuff you can run locally, the vendor lock-in stuff won't be your biggest challenge in the move.
I hate to tell you this, but there was a thread on HN about this exact topic not too long ago. While becoming less common with AWS, Heroku, Cloud Run etc, there are still companies large and small, that get it working on their machine then just run it off their machine until it breaks.
In fact, one of my favorite stories is a guy I know who does ML work gets a crazy multi CPU, multi GPU workstation for each project, and, when the project is running on his machine, they slap readyrails on it and ship it to the data center to run in prod.
looks like google spanner have plugged that workflow gap since you evaluated it for your project:
> The Cloud SDK provides a local, in-memory emulator, which you can use to develop and test your applications for free without creating a GCP Project or a billing account. As the emulator stores data only in memory, all state, including data, schema, and configs, is lost on restart. The emulator offers the same APIs as the Cloud Spanner production service and is intended for local development and testing, not for production deployments.
https://cloud.google.com/spanner/docs/emulator
although there are https://cloud.google.com/spanner/docs/emulator#limitations_a...
At my $WORK we run all of our backend on AWS but anything I've touched has to also run locally.
We use serverless framework, and there are plugins for running the lambdas locally as well as for dynamodb, sqs, ses and eventbridge.
I think it's a case of choosing your dependencies carefully though. I would be wary of integrating against an AWS service which does not have an API compatible offline or provided-elsewhere alternative.
A case where we fail at this is Cognito. Even our 'offline'/local stack has to connect to our dev environment for that one.
Then he goes on to advocate against designing for cloud flexibility.
This almost feels like AWS marketing.
Another example: there's multiple countries (for example, here in Russia) where personal data must be stored in data centers located in the country's borders and not every country has a AWS datacenter on its soil .
Reading the actual text of this one I get a different impression, but I'm still not sure I agree with this one. Applications can be radically different from each other in terms of how they are run.
At one company, we ran a simple application as SaaS for our customers or gave them packages to run on-prem. We'd stack something like seven SaaS customers on a single set of hardware (front-ends and DB servers). The cloud offering was a no-brainer, you can just migrate customers one by one to AWS or whatever, or spin up a new customer on AWS instead of in our colocation center.
Applications have a very wide range of operational complexity. Some applications are total beasts--you ask a new engineer to set up a test environment as part of on-boarding and it takes them a week. Some applications are very svelte, like a single JAR file + PostgreSQL database. The operational complexity (complexity of running the software) doesn't always correspond to the complexity of the code itself or its featureset.
I've only participated in a single on-prem to cloud migration. Some parts of the migration were easy, e.g. moving a postgres DB that was running on some on-prem linux server to run in AWS RDS. Some parts were rather unpleasant: where you discover that a bunch of the application code that runs in worker processes assumes it has access to a shared CIFS network share that can be used for communication throug the filesystem, and absolute file paths to on-prem CIFS network share locations are stored in metadata throughout the database. So then your available moves for how to migrate the application code and migrate the CIFS network share and migrate the data in the database all become somewhat tangled together.
Well, it's not about AWS shutting down at all! It's about them having complete control over your infrastructure, so they dictate the terms. This has many consequences: (1) they can raise prices and you can do absolutely nothing about it, (2) since you chose AWS with its dynamic pricing instead of flat-rate dedicated servers, each expansion (traffic, new services) is a cost for you. This means at some point you will realize you will save sick amounts of money if you switch to bare metal (as several notable companies have done). Except that at this point it's really difficult because you have to basically start from zero so the inertia basically pulls you into continuing this vicious cycle.
So this is just a straw-man argument. Really, I haven't heard anyone saying "but Amazon can go out of business", it's just ridiculous.
They may be startled, but they almost certainly won't listen. The purgatory nature of IT work culture ensures this repetitive pattern.
It’s not that I’m a jerk about it. Most of the time it’s business types saying “you don’t need to do all that stuff - just deploy it like that, it’ll be fine”.
And then it’s not fine - and also, it’s somehow my fault.
It seems to be more about personal aggrandisement (“I’m the boss and my word is law”) rather than trying to build a great business together. I’m pretty over it.
It is easier to talk than to listen.
It is easier to build new than adapt.
All three of those statements are false. But all three are seductive
I just wanted to point out that I found that particular error really funny given the context.
Designing for portability is important. Otherwise you expose yourself to dreadful uncertainties.
"AWS will not disappear". That is probably true. The average business can take this risk (and if you are huge you are not listening to me). But AWS might raise its prices to a point they are getting all your profit. DO you trust Amazon? Really? The particular AWS feature you depended on with the tight coupling "Don't Design for Multiple Cloud Providers" implies may get deprecated. What then?
This is as old as the hills: Design in layers. Have an AWS layer. If AWS goes away, quadruples their fees, deprecates your services, or you are hit with USA sanctions then there is a layer that has to be rewritten.
Old wisdom. Use it.
That is a feature! Bleeding edge bleeds.
And just a side note, Parler was toxic and I shed no tears about their demise...
Kara Swisher had a very interesting interview with their CEO, John Matze. Its a good listen.
https://www.nytimes.com/2021/01/07/opinion/sway-kara-swisher...
Now you've given your users 2 issues completely unrelated to the problem they're trying to solve. If you can't give your users a single simple command, or a single file to download your tool is too complicated to install.
My point being to the users of the products we create, software licensing is not what they are thinking about.
>Yes. Except. If I have a Python2 tool I paid $X for, and now I need to to change it for $Y because, because, because no reason. Tough luck! Stop "whining" and write a cheque!
If you're concerned about compile time -- Go does a pretty reasonable job of caching (including caching unit test results). If you're working on a Python project with a large number of unit tests, because Python is so slow to execute anything, and the go build and test tools are quite fast, it wouldn't be that surprising if it was actually faster to compile and run the test suite in Go vs running the test suite in Python -- particularly for incremental work where you make a change in one library and rerun the impacted tests.
If you're concerned more about the workflow of needing an additional compile step, go has `go run script.go` which lets you use go like a scripting language, assuming you're in an environment with the go toolchain installed.
pkg.go.dev and gopls also automate away most of the "exploration" I would be doing in Python REPL.
For packaging python stuff, it kind of depends who or what the end user of the CLI tool is. I used to work in a small business that shipped windows desktop software to customers, a lot of the software was written in python. From memory we packaged it with py2exe -- so you end up building a zip file containing a self-contained executable, the version of the python interpreter your tool needs, along with all the python packages as well as native libraries. That worked quite reliably, but it'd seem rather distasteful to use that approach for sharing CLI tools with your colleagues or CI machines in a dev team!
edit: there are pitfalls to deploying go application binaries if you try to put them in scratch containers and use libraries that assume they can find data such as timezone data provided in the usual place in the filesystem by the distribution (there will be no such data files in a scratch container unless you explicitly copy them in), or build a go linux binary with dynamic linking assuming glibc, then try to run it in an alpine based environment with musl. So it's not magic. But it's mostly pretty nice.
If you know that, this is not the problem for you.
Often you do not know that. Then this is a huge problem.
Hey you can always use docker for shipping /s
> Don't write internal cli tools in python
Disagree completely with this. This has been probably the biggest overall boost for both engineers and operators at a few companies, I worked at.
You deliver fast, it's easy to debug, and requires no compilation -- which is usually a bigger hassle than any Python-specific problem. It gets really important if you have operators on Linux/Windows/Mac.
If we ran our cluster in the cloud we'd be on the hook for hundreds of thousands of dollars of additional costs due to the high throughput of our service. There are always exceptions to any list of rules.
What do you do instead? We use Ansible to bootstrap k3s. It works, but we have to build a lot more stuff related to monitoring, alerting, routing, etc.
I kinda agree on his first point about migrating stuff to the cloud but if you've done your deployments on like-like platforms (on prem containers to cloud containers) its not that bad.
Well, maybe it's the latest in a long line of options...
There's not much wrong with just installing the tool to the global (user) interpreter, though.
Sure, you can do a lot with venv, Anaconda and whatnot, but if you have a well written package, it can be very portable without the need of environments.
There’s always going to be odd edge cases where things break, no matter what language you build your tooling, but I’ve installed thousands of Python packages on CentOS7, Ubuntu, OSX, Debian via pip and haven’t had issues. Ergo, I disagree with the writer.
This article is basically saying “things are complex, complexity is hard, so don’t do complex things”. On that point I agree. If you don’t have the time to invest in the tech, don’t build unsupported systems. But don’t throw Python under the bus because you don’t have time to build a proper package.
There is a very nice comment talking through some of the reasoning here that is worth a read: https://github.com/aws/aws-cli/issues/4947#issuecomment-5860...
I am quite annoyed v2 isn't in PyPI as it makes updating to v2 a sizable project in some cases and v1 does not support all services, but also quite understand their reasoning.
Granted the tradeoffs may be weigh out differently for an internal tool as the post describes, or if distributing the tool to a more constrained audience.
That is one area where I think you want to outsource that to specialists.
Build some toy security… but don’t deploy it.
Glad of that.
If you can get away with a zero dependency Python script then there's no struggle. You can download the single Python file and run it, that's it. It works without any ceremony and just about every major system has Python 3.x installed by default. I'd say it's even easier than a compiled Go binary because you don't need to worry about building it for a specific OS or CPU architecture and then instructing users on which one to download.
Argparse (part of the Python standard library) is also quite good for making quick work out of setting up CLI commands, flags, validation, etc..
There's a number of tasks where using Python instead of Bash is easier. I tend to switch between both based on what I'm doing.
If you say "Python3 of course, is this 2002 or something?" what do I do with all my Python2 scripts?
A lot of Python 2.x scripts will work with Python 3. If they don't then it's on you to fix them since Python 2.x was officially labeled end of life almost 2 years ago from today. On the bright side, your Python 2.7 script had a good run. 2.7 was released back in 2010, so having ~10 years of not having to worry about anything is pretty good! Chances are we'll get the same experience or longer with Python 3.6+, we already at the 5 year mark for Python 3.6.
Does not inspire confidence!
Eh, the salesman told me it would be seamless while we were watching the football game from his company’s box. And they are the experts: it’s their cloud!
I’m gonna tell the team to do it this way when I get back to the office. I think they just like running hardware and aren’t thinking of our balance sheet.
$ cat dogclock/__init__.py
import arrow # Example dependency
def bark():
print(*(["woof"] * arrow.get().hour))
$ cat scripts/dogclock
#!/usr/bin/env python3
import dogclock
dogclock.bark()
$ cat setup.py
from setuptools import setup
setup(
install_requires=['arrow'],
packages=['dogclock'],
scripts=['scripts/dogclock'])
$ pip3 install .
…
$ dogclock
woof woof woof woof woof woofWhat if your team is Python-based? Why would I write a CLI tool to be used by other Python programmers in Go or Rust, when some of them know neither?
It doesn't matter that you know Go and can generate all possible binaries; eventually, someone else will have to make a change in your tool. It will already be difficult for them to understand a new codebase, so you don't need to make it harder by also exposing them to another language.
Anybody tried PyInstaller ? it packages the whole Python project, dependencies included into a single exactable binary
This has its own sub-antipattern: “Just put your application in a container, then it will run anywhere!”
Isn't the reason people do this to make sure they have leverage in case AWS increases prices in the future? I can see how cloud providers have probably made it extremely difficult to design for multiple clouds and so this effort might not be worth it but the reason at least seems justifiable.
I wonder if a useful middle-ground is to have lint checker rules to enforce using a blessed subset of a cloud provider's SDK. So that some thought/effort must be put into using a new feature?
Thank you.
I am not sure what will happen if you will not?