I would love to see a move towards "static" binaries that package everything together into a single, self-contained unit.
I would love to see a move towards "static" binaries that package everything together into a single, self-contained unit.
* insert lines of code into asciidoctor Ruby to debug how the command line interface differs from the library interface;
* read through Raspberry Pi SenseHat code to figure out that the drivers won’t load via SSH (you have to plug in an HDMI cable!);
* rummage through Python’s PY Sequence List to answer the question “what kind of sort does Python actually use?”.
The problem with shipping binaries is it lowers the standard for being able to build from source.
If a package is easier to ship as a binary because the source is hard to distribute reliably, we give up hope of end users being able to modify and build their own versions.
This is obviously bad for freedom. It’s also hard for debugging and learning.
I always _try_ to show how one can answer problems from first principles before resorting to “go look it up on stack overflow”.
It boosts ones sense of usefulness as a teacher if one can occasionally give more in depth answers as teachable moments.
For one thing, demonstrating deeper understand is how you score high marks in exams.
Open up Python's source code and check - this is first principles-first hand.
As a matter of fact I am going to go as far as saying that reading source code is the _only_ way to actually be any good in software engineering because it's the only source of truth that never gets out of sync with the de-facto.
Did you expect to find a blog post about it (you can find it on the first page of Google search I mentioned)? Or a YouTube video (didn't pay attention to that)? Books are reasonably good but usually get out of sync quite fast.
And going back to first principles: the reason most of the literature may not mention what sorting algorithm Python uses is because a) it's irrelevant as a first principle knowledge b) it gets quickly out of sync.
An important part of school is learning how to learn. It is a constant struggle with pupils to get them to actually think instead of googling for an answer.
Sometimes it feels a bit like I am a grumpy oldie telling them things like _in my day we had to look things up in books instead of using Google!_ but most of the time it’s valuable in and of itself to go through things from first principles.
(The sensehat bug was a case in point.)
As far as source code is concerned, it turned out to be quite tractable to follow the Python source code at least far enough to find the timsort.txt document.
If any of these kids go on to have technical careers, there will be plenty enough time for googling answers later on in life. Hopefully some of them will be writing the answers too.
No one in tech support believed me. No one was willing to help me. They told me to make a new account.
I've done my best not to buy from Amazon since.
Please don't be hostile to users.
Possibly after you provide some kind of validation like the last few digits of credit card tied to that account or something?
I'm really curious about this. Why would a driver care about whether you load it on a "real" tty vs. over the network? How does it even know?
We didn’t dig further than that, though the error was quite obscure until we read the source code, at which point we realized that the “FB” the error mentioned meant the Linux console framebuffer.
I assume the Python module for operating this device ships with a GUI, which is couple to some Linux framebuffer initialization code, which in turn only works if the framebuffer is enabled, which presumably it isn’t if the host boots without an HDMI device connected.
The LED matrix is an RGB565 framebuffer with the id "RPi-Sense FB". The appropriate device node can be written to as a standard file or mmap-ed.
https://www.raspberrypi.org/documentation/hardware/sense-hat...
You might be giving up hope too soon, because what you describe has in my experience been the life story of source code from the start, to some extent. And I'm not blaming it for that. Distributing source code (well, not the distribution itself, but getting it to build) on all possible systems out there is hard and impossible to get 100% correct. It's simply inevitable to encounter systems on which it doesn't build because of their specific configuration, or user error.
Despite that, end users have been building their own versions and still do. Users who really want to don't let themselves hold back by a compiler error. Granted it's possibly that in a perfect world where everything would just work there would be somewhat more users doing that, but there's no such thing as a perfect world.
Moreover I do have the impression things are actually much better these days, mainly thanks to CI which makes it easy to get out of the bubble and build on multiple platforms. At least I don't have that feeling anymore of 2 decades ago where I sort of was prepared to waste hours getting anything to build which got handed to me on the internet. Still, it would be very interesting to go and see the percentage of projects out there which just build according to the instructions.
What? I can assure you this is not true.
My R Pi has never touched an HDMI cable and runs a sense HAT just fine. What OS?
If you needed to dig through source code to find out that python use Timsort, I genuinely am curious about how bad you are at search engines.
I believe in reading source code to answer some things, but that example is really odd for anyone who has researched sorting algorithms since Python literally invented a sort used by many other languages now.
If you need a billion dollar search engine and an internet connection to make sense of the source code on your local machine, I am genuinely curious how bad you are at reading source code!
But as I said, that’s pretty uncharitable. Because of course I didn’t have the source code for Python locally.
I downloaded it after Googling “Python source code”.
+ Explaining something that is "obviously common knowledge".
+ Reducing motivation of participant to share their experience.
I think we have HN bingo right here.
It isn't always possible, but Nix places a lot of emphasis on building software from source (or downloading a pre-built copy from a binary cache), and it has fairly good support for overriding/customizing the expression that builds one or more pieces of software.
A few months ago I was trying to debug a test in the Oil shell project that broke only on macOS for very unclear reasons. I eventually figured out that, in a specific mode, bash (which both Nix and macOS use as sh by default) was filtering out a specific environment variable that this specific test was trying to set, before running a target program with a sh shebang. I only needed to add a few lines to the project's build expression to patch bash and use the patched version for the tests (all without affecting the system/user bash).
The key benefit that nixpkgs has over other Linux distributions' package sets on this topic are:
1. It allows you to _program_ your package set, thus adding static-linking overrides to 1000s of libraries/executables in a go, instead of having to modify each package manually.
2. It allows the end user to choose what linking they want, as opposed to the Linux distribution making that decision for them upfront.
3. It allows to do this from any Linux distribution.
In my experience, the best way to read source code and experiment are to get the source from the upstream repo (eg. GitHub) and read/build from that. I do this with code written in C or C++ all the time.
The code is always out there (unless you're talking about proprietary software, which none of your examples above are).
I've suspected that for a long time without confirming it personally, so thank you for your time.
Sure people may not use the tools they have but then again people also don't upgrade their OS packages or maven packages. So for those people packaging everything doesn't hurt any more since they've got rampant CVEs already.
Like if building a pet app is your company/team's job then containers more than fine. If you're a distro maintainer that has thousands on thousands of packages then sharing libs starts to make a lot more sense.
They're screwed. Hopefully you aren't.
Hopefully.
The problem is that distributions get their software from upstream developers. And if upstream decides to go static or use embedded code copies, the maintainers have much more work trying to untangle everything.
In my experience, packaging Go software for Linux distributions is very frustrating and tedious because of that.
Packaging Python or C/C++ software is much easier.
Wondering since the application developers I know seem to prefer static linking.
I can't help but wonder if we'd all just benefit from better build systems, so that "rebuilding all downstream consumers" wouldn't be hard.
Perhaps this "untangling" isn't the best idea, at least not always?
- the bundled things might have unclear licensing
- bundling stuff that is available in the repos is sub optimal, as the repo version will be updated for security issues while the bundled version will not
- you can use some mechanism to mark what stuff is bundled in a package, but then you need to make sure the bundled thing is patched and rebuild each thing bundling it, where with system wide dependency you just update that and you are done
I'm fairly sure that almost all projects are receptive to updates for dependencies with security issues; I don't see why a distro needs to "untangle" anything for them.
> with system wide dependency you just update that and you are done
Unless the system-wide update won't work with your program.
The entire point of the Go tooling is that you can reliably and consistently build your software, and replacing random parts to create a weird unsupported (by the actual developers anyway) version is something the Go tooling was never designed to support: it goes exactly against what it was designed to do. Hence my previous comment: perhaps this isn't the best of ideas.
- Outdated packages
- Having to build from source when the package is not in the repos
- Failed builds for no good reason
- Unability to have many versions of that same packages
- having to install log(n) build tools if I want to build n packages
- etc
while in half-developed OS'es everything just works (application-wise, at least)
There's a fundamental unresolvable tension in automatic update systems between "not updating breaks things" and "updating breaks things"; you cannot solve for both at once and either one will get users mad at you, although not upgrading at least gives a pushback argument. See how all this has played out with Windows and Mac updates, for example.
You really should check what APIs your dependencies provide and how stable that API is, then set minimal dependency version. If you set the dependency to a specific version, you are creating a headache for anyone trying to use your software later on (by potentially forcing them to use oudated and insecure software) and for distro maintainers trying to keep distro software up to date and secure.
What if the dependency is so unstable that you have to pin the version or even use custom patched version? Well, thenaybe its not somethong you should be depending on & rather use something with less features but more stable. Othervise you are really being unresponsible - using quick convenience at the cost of long term usability of your software.
The thing is that in general you don't want to have many version of libraries available in the distro in parallel as each has to be separately maintained, security patches applied, build issues fixed, etc., eating valuable maintainer time. Also in some cases libraries of different version can't coexist on a system cleanly.
Then imagine every piece of software just pins versions of their dependencies to a specific random version that happened to work at the time. To satisfy those arbitrary dependencies the distro would have to maintain all these versions at the same time, which is simply impossible resource vise, not to mention incredibly wasteful (both in maintainer time & system resources).
As for stuff breaking if you don't pin dependency version - well, distros have mechanisms to handle that. For example for Fedora, there is a stable release very 6 months & stable releases are not expected to get major changes in libraries, just bug fixes and smaller enhancements.
And at the same time there is a rolling version of Fedora called Rawhide, where all the latest package versions land and where integration issues are addressed. So any breakage would happen on Rawhide and be addressed by maintainers (of the library/software affected or both) long before a new stable release is cut from Rawhide and users will actually use it.
For an example I'm maintaining the PyOtherSide Qt 5/Python bindings on Fedora. A while ago the build failed due to Python being updated to 3.9. I reported the issue upstream, which quickly fixed it and I've built an updated version in Rawhide. All this long before a stable Fedora version will get Python 3.9, but I can be sure that when this happens, all will work fine.
[0] Remember when Debian generated predictable random numbers because a maintainer wanted valgrind to shut up?
Now in comparison people using NPM just blindly pulled random stuff directly from upstream without anyone doing any sanity checking at all - no wonder one package vanishing made the whole thing fall over, often directly in production.
Especially with Python, and especially with the 2/3 split, people got used to assuming that the distro version was something broken to work around (e.g. Redhat, OSX), and that all "real" work happened in one's local language-specific package cache or venv.
I'm increasingly of the opinion that it's a mistake for distros to ship Python or Ruby packages in their distro-specific package format, but I can also see that's going to be a holy war.
As Linux user who sometimes build packages (tinker around) I need all kind of dependencies - python, ruby, perl, haskell. It is much safer and faster to use distro packages.
Breaking changes on major version is awful for any consumer. Python 2/3 story is a shame. These can't be arguments against distro packages.
I can't help but wonder if there's a better way that could scale to open source. I personally like monorepo-based development a lot - where you have one version of every library for the whole repo, and a total ordering on changesets (and a strong test suite to catch regressions). But organizing open source into a monorepo or even a virtual-monorepo seems tricky. Might be the sort of practice that can work well for companies but not so much when decision power is more distributed.
And users would find outdated versions (maybe unmaintained). Current system update dependencies automatically or allow collaboration on life support (patched application, dependencies, config). It would be good to have alternative like container or virtual machine with snapshot from the days long past. But it should be clear this is not safe.
Em, like all other packages?
Here are the scenarios:
- dynamic responsive remove bug: Positive/neutral. Team X would have done it anyway.
- dynamic unresponsive remove bug: Positive.
- dynamic responsive add bug: Negative. Team X will see the bug but only be able to passively warn users not to use Y version whatever.
- dynamic unresponsive add bug: Negative. Users will be impacted and have to get Y to fix the error.
- static responsive remove bug: Positive/neutral: Team X will incorporate the change from Y, although possibly somewhat slower (but safer).
- static unresponsive remove bug: Negative. Users will have to fork X or goad them into incorportating the fix.
- static responsive add bug: Positive. Users will not get the bad version of Y.
- static unresponsive add bug: Positive. Users will not get the bad version of Y.
Overall, dynamic is positive 1, neutral 1, negative 2, and static is positive 2, neutral 1, negative 1. Unless you can rule out Y adding bugs, static makes more sense. Dynamic is best if "unresponsive remove bug" is likely, but if X is unresponsive, maybe you should just leave X anyway.
If the install is on a consumer machine for regular usage then the right answer is shared libraries for the machine. It adds a lot of complexity for the packagers for that OS but you get a lot of safety for that consumer.
If the install is an inhouse app deploying to servers a company controls/rents whatever then the right answer is probably a static application. Consumers are more likely to want to chase the latest version of everything. Company developed software is far less likely to want that. The problem then becomes less about packaging and more about deploying fixes quickly. A static binary that is rebuilt, tested, and deployed is going to be a smoother path to fixing that error than figuring out how to deploy your new shared library and avoiding any issues in our production environment caused by clashes to that library. This is made more problematic by the likelihood that you will need multiple versions of that shared library with the fix to meet the needs of applications that can not yet be upgraded. Static binaries make things easier in this scenario and thus quicker to resolve.
> "static" binaries that package everything together into a single, self-contained unit.
Ummm, except debugging "static" binaries without symbol file is hell.
If there's a .py source somewhere on the disk, I can at least try edit the source file from the package and debug internal private variables as last resort.
> Trying to load source files from all over the filesystem at runtime is hell.
This sounds like a mess unrelated to Python. Why are you trying to load source files from "all over the filesystem" at runtime?
> I would love to see a move towards "static" binaries that package everything together into a single, self-contained unit.
Compiling source files from all over the filesystem would be equally annoying.
It's basically a huge mess. There's even an XKCD about it.
Consider the alternative: Go compiles programs to a single statically linked executable with no dependencies. You literally just copy one file to your target machine and run it. It basically can't fail. That's part of the reason Go is so popular for server stuff.
And it's not just because Python isn't compiled. Other scripting languages handle this much better. Even JavaScript - for all the hate node_modules gets for being enormous, at least it works reliably!
> Those libraries are installed by Pip... somewhere on your system.
They go into your virtualenv.
> And then Python has to find it somehow at runtime.
If you use a virtualenv, this works 99.9% of the time.
> That relies on environment variables being set correctly
What environment variables? Just use `./venv/bin/python` directly without environment variables.
> a huge mess even before you consider things like virtualenv
There's nothing to consider, always use a virtualenv. It's the same for thing for Node, except it implicitly handles it for you.
> and the fact that non-Linux systems usually have multiple copies of Python installed (and Python 2 and 3!).
How is that a problem? Just pick the interpreter you want to use when creating the virtualenv:
virtualenv -p /usr/bin/python2 venv
virtualenv -p python3 venv
virtualenv -p python3.7 venv
> Consider the alternative: Go compiles programs to a single statically linked executable with no dependencies. You literally just copy one file to your target machine and run it. It basically can't fail. That's part of the reason Go is so popular for server stuff.I agree that Go got it mostly right and it just works, except for that fact that it didn't even have a package manager for like 10 years so pinning dependencies was impossible unless you forked repositories.
> Even JavaScript - for all the hate node_modules gets for being enormous, at least it works reliably!
In what way is Python + virtualenv + requirements.txt less reliable?
> It's the same for thing for Node, except it implicitly handles it for you.
Indeed.
Funnily enough node_modules is one of the main regrets Ryan Dahl, the creator of Node.js has: https://www.youtube.com/watch?v=M3BM9TB-8yA&t=755s
Python has the most unpythonic package experience ever. As elegant as the language is, packaging is a convoluted nightmare. It still might work for some project, but certainly is no fit for the average python user.
That being said, I'll admit I don't know much about what it means to be pythonic. What do you think the drawbacks are to pipenv-type approach and what would be more pythonic in your mind ?
I wish I could just as easily say: ahh this script I made should become a package that I can reuse on my server. But then the server has Python 3.6 while your script has 3.8, so you end up installing another python, painstakingly watch not to install over the old one etc. When installing modules you have to do the same or set up a venv/pipenv.
All along the road there are possible ways to shoot yourself into the leg. Meanwhile in Rust you do a cargo new foobar, work on it, add dependecies to the cargo.toml, copy it to the server, build a binary with cargo build --release and copy the binary into the PATH. You spent zero brainpower on not breaking things and had capacity to think about other stuff.
The closest thing we have in those terms in pythonland is poetry, which is a very good start. But this should be part of the language, not something one has to do extra.
Thanks for the insight and clarification.
Isn't that why Docker is so popular? If we had Python static binaries (or the ability to have several versions of the same Python library installed) would we use Docker so much?
We still used wheel archives etc for managing some python dependencies during our dev and build process, but it was never a complete solution: some python packages depend upon non python shared libraries so you need to install them using an operating system level package manager. Then when shipping the software to users you cannot assume the user has a working version of python, or even if they do you don't want to increase your support burden by having your application running inside some arbitrary python environment.
E.G:
- shiv leave files hanging around, but you can address those files directly. With PyOxidize, you will need to use https://pyoxidizer.readthedocs.io/en/stable/config_api.html#... a way to read a non Python files. Because most projects don't know this and just use open(), they won't work out of the box and you may need to monkey patch open().
- pyoxidizer requires a compiler, which is easy on linux, but a higher requirement on windows. Compared to shiv, which is just a pip install away.
There is no perfect solution yet, because you have really often many things to balance: bringing in the python vm or not, allow compiled extensions or not, provide support for open() or not, etc.
IMHO python, node.js, ruby, and other scripting languages share a historical confusion of concerns related to development tooling and attempts to counter the unfortunate choice of installing dependencies globally that tend to leak to the server side in inappropriate ways. With Docker.
There are other things docker can do (like mapping volumes) but I would be much happier if we solved it at the level of the executable.
Most python packages I have seen will put their python source in the python site-packages directory. Where were the source files you were trying to load located on the file system?
https://stackoverflow.com/questions/3430400/linux-static-lin...
One need not depend on the filesystem for such things. And as bad as python packaging is, others get it better. One can have less-than-static-binaries without being quite that bad.
> I would love to see a move towards "static" binaries that package everything together into a single, self-contained unit.
Plenty of problems with that too!
You mean Docker containers? :-)
And yesterday I spent at least an hour fighting with cargo build. Without knowing exactly why we fought both our anecdotes are quite useless.
I wonder if there have been any surveys on poetry, and whether it's trajectory (presumably exponential) suggests that it's going to become a dominant package management utility in the next 10 years or so?
tldr; Do not use Python. It makes certain tasks seem easy, but all it truly does is make you look like an idiot developer. Pick literally any other language, as Python has screwed up too many basic concepts (related to POSIX and expectations from C) to be taken seriously.
I stand by my opinion that Python is not well-suited to applications written by a professional entity, intended to be shipped to 3rd-party clients. Python plays too loose with the fundamentals provided by the operating system to be a professional language.
For me, most of the pain with the Python packaging went away after I started using Pip-tools[0]. It's just a simple utility to add lockfile capabilities to Pip. Nothing new to learn, no new filosophies or paradigm's. No PEP waiting to be adopted by everyone. Just good old requirements.txt + Pip.