Python 3.11.0 final
discuss.python.org
discuss.python.org
But remember that while it's great to play with it freshly out of the oven, and that you might want to test your projects/libs with it, we should wait a bit before migrating production.
Indeed, every first release of a new major version of Python eventually important bugs that get ironed out in a later patch. Also, some libs on pypi may simply not be compatible with it yet, breaking your pip install.
I usually wait until the 3rd patch myself, after many years of paying the price of greedy upgrades.
We wouldn't get there if everyone does that though.
Not much room left for integration/development and so on
I pick these things up early in lower environments, as I suppose you do, so that the real production upgrade isn't scary... and we don't fall behind
Waiting forever for the community to QA doesn't work, you probably aren't running exactly what they are
> After some conversations with Yury, and encouraged by the SC's approval of PEP-654, I am proposing to add a new class, asyncio.TaskGroup, which introduces structured concurrency similar to nurseries in Trio.
I have never used but have been told that Trio's nurseries make it much easier to handle exceptions in asyncio tasks. Does someone more knowledgeable can tell if this will help? Looking at the docs*, this only seems to be a helper when you want to await several tasks at once, so I am not sure this changes much for exception handling.
* https://github.com/python/cpython/issues/90908
** https://docs.python.org/3.11/library/asyncio-task.html#task-...
> this only seems to be a helper when you want to await several tasks at once
Sort of. It's a helper for if you want to run multiple tasks at once, not necessarily awaiting them. And you're definitely running multiple tasks at once otherwise you wouldn't be using asyncio in the first place.
Task groups do require you to wait for the tasks - after all, you have to start the task in a task group, and then implicit await the tasks in it (by falling off the end of the task group context block). But you can always have an outer task group representing tasks that you indent to run indefinitely in the background. In that way, task groups force you to think about when a task would cancel other tasks, representing the overall structure of your program.
I managed to make a very very simple OTP-like framework with Trio: https://linkdd.github.io/triotp/
Nice!
> PEP 657 – Include Fine Grained Error Locations in Tracebacks
Hmm, what’s this?
Traceback (most recent call last):
File "test.py", line 2, in <module>
x['a']['b']['c']['d'] = 1
~~~~~~~~~~~^^^^^
TypeError: 'NoneType' object is not subscriptable
YessssI love writing chained expressions but debugging them is like visiting a special kind of hell.
Traceback (most recent call last):
File "test.py", line 14, in <module>
lel3(x)
^^^^^^^
File "test.py", line 12, in lel3
return lel2(x) / 23
^^^^^^^
File "test.py", line 9, in lel2
return 25 + lel(x) + lel(x)
^^^^^^
File "test.py", line 6, in lel
return 1 + foo(a,b,c=x['z']['x']['y']['z']['y'], d=e)
~~~~~~~~~~~~~~~~^^^^^
TypeError: 'NoneType' object is not subscriptableI sometimes sacrifice readability just because I hate creating variables. But then if it affects debugging times, my boss would be furious. As such, I use a full debugger anyway so I can trace quickly.
I actually was inspired to build the library after teaching a one week intensive pandas course to a couple of Data Scientists @ a Fortune 500... (pandas is really hard for beginners!)
While performance and OOM aren't priorities right now, I'd love to one day replace the pandas "backend" with Arrow (or something else) once I nail the API :)
The geometric mean of the 3.8 to 3.11b benchmarks was a 45% speedup.
My understanding is that it's based on the most recent attempt to remove the GIL by Sam Gross
https://github.com/colesbury/nogil
In addition to some ways to try to not have nogil have as much overhead he added a lot of unrelated speed improvements so that python without the gil would still be faster not slower in single thread mode. They seem to have merged those performance patches first that means if they add his Gil removal patches in say python 3.12 it will still be substantially slower then 3.11 although faster then 3.10. I hope that doesn't stop them from removing the gil (at least by default)
This is intended to supplement the current PGP signatures, giving Python distributors (and the packagers themselves) a similar degree of authenticity/identity without needing to perform PGP keyring maintenance.
The idea behind sigstore is to enable an end user (which, for CPython, might be someone who intends to build it from source for inclusion in a package manager) to verify that the artifact is the same one produced by a trusted entity, not just the same one downloaded from a server. A strong cryptographic hash only provides the latter.
It is not "bound" to anything cryptographically. Sigstore checks that you own the OIDC account, and if yes, it signs your public key and puts it in the log. Why not just sign your software's hash and put it in the log, "binding" it as you say?
> Why not just sign your software's hash and put it in the log, "binding" it as you say?
That's exactly what it's doing. Is the objection you have solely to the fact that it can be done with short-lived keys?
In this specific case, we can say that the artifacts in question were verifiably signed by owner of the <release-manager>@python.org identity.
Why not put a hash in that trusted registry instead?
A hash verifies integrity, but has no way to demonstrate any relationship to a signing identity. Signing is not just about integrity, but also being able to say _who_ generated the signature.
The way it seems to me that it works:
* You generate a keypair and use it to sign your Python installer (SP)
* Authority creates a signature of your public key (SA) and puts the whole thing in transparency logs (SP + SA)
* End users can check SA because they trust authority, can check SA is in transparency logs, therefore can trust your signature SP, therefore can trust your software
Why not just use a hash instead:
* You generate a hash of your Python installer (HP)
* Authority creates a signature of your hash (SA) and puts it in transparency logs (H + SA)
* End users can check SA because they trust authority, can check SA is in transparency logs, therefore can trust your hash, therefore can trust your software
What matters is that ultimately the contents of your software are signed by the Authority and a commitment of that is in logs. Why add this level of keys, that can't possibly be trusted since they are ephemeral?
It also requires the user to trust the signer to do proper secret generation, which is weakens the scheme. With Sigstore, the entire CA and CT infrastructure can fail or be compromised, but the certificates (and the ephemeral keys that they bind to) remain sound. That too is desirable, which is why the TLS PKI ecosystem is the way it is.
Edit: To be clear, the PGP equivalent for your scheme would be "trust Joe Public to sign for everyone on PyPI, he's reliable." If you can see why that doesn't work, you should also be able to see why your alternative to Sigstore won't work.
And whether the CA signs "identity information + public key" or "identity information + software hash", I don't see the different in "identities", no matter what that means to you.
Please, give a concrete example of information that is available/verifiable in one scheme and not the other. You both keep saying vague things like "it lacks identities" or "there's a binding" etc and I really don't see it.
For how this works specifically with CPython, see https://www.python.org/download/sigstore/ for details.
For how this works generally: it's the same public/private key cryptography you're used to elsewhere, and https://docs.sigstore.dev/ has more details.
There is no singular “root certificate”: there’s a trust root for the CA, a separate root for the transparency log, etc.
The certificate just binds the public key to the identity at a given point in time, in a public way. This certificate is generated every time you sign something, and is put in the transparency log.
There's a walkthrough of the process here that might be helpful: https://www.youtube.com/watch?v=jdf-gNYg0fw&t=494s
PGP has a web-of-trust aspect, allowing people to trust people. What is the point of something doing automated verification of identity, on top of the one done for HTTPS certificate issuance?
They’re different groups of people, which points to one of the potential benefits to sigstore verification here: people who download CPython from python.org can now additionally verify that the artifact was not tampered with on the server. They can, furthermore, mirror the artifact and its signing material on their own. In short, TLS provides delivery authenticity while sigstore providers publisher authenticity.
However, that complexity does not apply to simple cases like the CPython one: for this case, you can verify that the identity matches the one of the public email identities of the CPython release team. This is no more complex that PGP identity verification, and is much more resilient (since anybody can publish a claimant key for an identity in PGP).
You also linked to this random python.org HTTPS page which contains the list of people you are supposed to expect to have signed the Python releases. If this is the root of trust... it might has well have had PGP fingerprints.
The truth is that you login with OIDC and Sigstore signs your artifacts, giving you an attestation that the owner of that email/GitHub/... identity made that artifact, and publishes that to a persistent log. This makes the whole thing great for automation (both of publishing and verifying), but claiming that this adds to security is false.
Their Security Model page clearly outlines the limits of their system and is consistent with my characterization https://docs.sigstore.dev/security/:
> If an OIDC identity or OIDC provider is compromised, Fulcio might issue unauthorized certificates
> If Fulcio is compromised, it might issue unauthorized certificates
You have to trust OIDC providers, you have to trust the CA, and the presence of logs only allows those people to notice unauthorized issuance, not end users.
* Sigstore uses short-lived keys and short-lived certificates, eliminating an entire common risk class where maintainers accidentally disclose their signing keys. This property alone eliminates the single largest source of illegitimate signing events in ecosystems like Windows software.
* The logs in question are public CT logs. In other words: anybody can audit them for unauthorized issuance, including the legitimate publishing identity. It's not particularly useful for the end (installing) user to audit the log, but it was never claimed that they would find it useful.
For the specific case of CPython, you're missing the point: CPython is an easy case, since the email identities of the release managers are well-known facts that can be cross-checked across python.org, GitHub, etc. Python.org is not currently a root of trust for sigstore, but it is for PGP (again, because anybody can claim an identity in PGP).
There are, of course, limitations. But these limitations are no strictly worse than trusting CA and IdP ecosystems that you're already trusting, which makes them strictly better than mystery meat PGP keys.
It is also much more likely that someone managed to click one link in a developer's inbox once to complete the automated Sigstore verification, rather than they managed to steal their PGP keyring and passphrase.
I am not a fan of having to trust in developer's key-management abilities but this just shifts the problem very slightly, at significant cost.
The single advantage is obvious: this allows easy automated signing and verification, allowing enterprises to easily check boxes in their supply-chain-security checklist. This is valuable in itself, and I am all for automation, but I don't know why we have to claim that it is "more secure".
A PGP fingerprint is tied to a PGP key, which is tied to a claimed identity. Anybody can claim to be you, me, or the President of the United States in the PGP ecosystem. Some keyservers will "verify" email-looking identities by doing a clickback challenge, but that's neither standard nor common.
In theory, you trust PGP identities because of the Web of Trust: you trust Bob and Bob trusts Sue, so you trust Sue. But it turns out nobody actually uses that, because it's (1) unergonomic and doesn't handle any of the normal failure cases that happen when codesigning (like rotation), and (2) it's been dead because of network abuse for years anyways[1].
> It is also much more likely that someone managed to click one link in a developer's inbox once to complete the automated Sigstore verification, rather than they managed to steal their PGP keyring and passphrase.
That's not how Sigstore does email identity verification; it uses a standard interactive OAuth flow. Those aren't flawless, but they're significantly better than a secret URL and fundamentally avoid the problem of secure key storage. Which, again, is actually where most codesigning failures occur.
And again, you don't have to use web-of-trust. It is there, which is an advantage. If you don't/can't use that, you can find a PGP fingerprint on a random HTTPS page, which will be just as easy to copy-paste as the list of email addresses you showed me a couple posts up... with the advantage that I can use them for verification directly, rather than involving third-party authorities.
And the same can be said for PGP keyholders. There are very, very few threat models in which an open, logged-in computer is not a "game over" scenario (which is also why most password managers and authentication agents don't consider it a case worth guarding against). In other words: Sigstore is no worse than PGP key management in this manner, but is better in the other ways that matter.
Looking up PGP fingerprints on random HTTPS pages is not a scaleable or ergonomic solution, and not one that has ever succeeded. Remember: that is the status quo with both CPython and Python package distribution, and there is no evidence that either had gained any meaningful amount of adoption (either by packages or end users). The goal here is to enable users to sign packages without doing the things they've demonstrated they won't do.
(Also, we've focused on email identities. A separate goal is to allow GitHub Actions identities, which will require no interaction from a user's browser and has a threat model coextensive with the CI environment that many Python packages are already using to build and publish their distributions.)
> with the advantage that I can use them for verification directly, rather than involving third-party authorities.
I'm not sure what you mean by "third-party authorities" here. As a verifier, your operations can be entirely offline: you're verifying that the file, its signature, and certificates are consistent, that their claims are what you expect, and (optionally) that the entry has been included in the CT log. That latter part is the only online part, and it's optional (since you can opt for a weaker SET verification, demonstrating an inclusion promise).
So there's a) no long-lived private key for them to lose (because it's never stored after signing) and b) a consumer doesn't need to find the right key PGP ID, verify (somehow) that that key ID is associated with a given release manager -- they can just trust that the release manager is in control of their @python.org identity.
Additionally, with PGP, you have no idea if your private key is being used somewhere else to generate valid signatures maliciously. With Sigstore, in order for the signature to be valid, it must be published in a transparency log, which is continuously monitored. So in the event of if the key/identity is compromised, the identity owner can be made aware immediately and the signature revoked.
More details are here: https://www.python.org/download/sigstore/ and here: https://docs.sigstore.dev/
Must be great for the 5 people on the planet that maintain a personal web of trust with PGP while the rest of us just run whatever "curl|gpg --import" command the download page tells us to run, thus adding zero security on top of https.
But.. I am I the only one who struggles to parse the Exception groups?
*ValueError: ExceptionGroup('eg', [ValueError(1), ExceptionGroup('nested', [ValueError(6)])])
*OSError: ExceptionGroup('eg', [OSError(3), ExceptionGroup('nested', [OSError(4)])])
| ExceptionGroup: (2 sub-exceptions)
+-+---------------- 1 ----------------
| Exception Group Traceback (most recent call last):
| File "<stdin>", line 15, in <module>
| File "<stdin>", line 2, in <module>
| ExceptionGroup: eg (2 sub-exceptions)
+-+---------------- 1 ----------------
| ValueError: 1
+---------------- 2 ----------------
| ExceptionGroup: nested (1 sub-exception)
+-+---------------- 1 ----------------
| ValueError: 6
+------------------------------------
+---------------- 2 ----------------
| Exception Group Traceback (most recent call last):
| File "<stdin>", line 2, in <module>
| ExceptionGroup: eg (3 sub-exceptions)
+-+---------------- 1 ----------------
| TypeError: 2
+---------------- 2 ----------------
| OSError: 3
+---------------- 3 ----------------
| ExceptionGroup: nested (2 sub-exceptions)
+-+---------------- 1 ----------------
| OSError: 4
+---------------- 2 ----------------
| TypeError: 5
+------------------------------------
Would it not have been better to left or right align the exception group id? Centering them just clobbers them with the actual error output and makes it a bit hard to parse.Maybe it'd look better in the terminal, but to me it feels like the table formatting makes it HARDER to understand.
*ValueError: ExceptionGroup('eg', [ValueError(1), ExceptionGroup('nested', [ValueError(6)])])
*OSError: ExceptionGroup('eg', [OSError(3), ExceptionGroup('nested', [OSError(4)])])
| ExceptionGroup: (2 sub-exceptions)
+-+- 1 -------------------------------
| Exception Group Traceback (most recent call last):
| File "<stdin>", line 15, in <module>
| File "<stdin>", line 2, in <module>
| ExceptionGroup: eg (2 sub-exceptions)
+-+- 1 -------------------------------
| ValueError: 1
+- 2 -------------------------------
| ExceptionGroup: nested (1 sub-exception)
+-+- 1 -------------------------------
| ValueError: 6
+------------------------------------
+- 2 -------------------------------
| Exception Group Traceback (most recent call last):
| File "<stdin>", line 2, in <module>
| ExceptionGroup: eg (3 sub-exceptions)
+-+- 1 -------------------------------
| TypeError: 2
+- 2 -------------------------------
| OSError: 3
+- 3 -------------------------------
| ExceptionGroup: nested (2 sub-exceptions)
+-+- 1 -------------------------------
| OSError: 4
+- 2 -------------------------------
| TypeError: 5
+------------------------------------Perhaps now flake8 will finally add the support for pyproject.toml as a config file...
See https://github.com/PyCQA/flake8/issues/234#issuecomment-1206...
Besides… Between Black / Tan for cosmetic issues and Mypy / Pylance / Pyright for logical issues, flake8 has never since caught any concrete problem with my codebase and has solely been a source of things to disable or work around.
I have just three configs:
- ignore TODOs (but only in pre-commit so i still get the IDE squiggles)
- line length 88 (black compatibility)
- add pydantic to extension whitelist
> Simple "JIT" compiler for small regions. Compile small regions of specialized code, using a relatively simple, fast compiler.
https://github.com/markshannon/faster-cpython/blob/master/pl...
[1] Random google result as source: https://www.cyclonis.com/microsoft-edge-tests-disabling-java...
My opinion is that it's mostly a curse, since you are hampering the language's evolution and growth for a mere temporary benefit.
Not to mention the horror of distributing compiled libraries, which is one of the biggest reasons why packaging in Python is still such a nightmare.
Making CPython faster by getting rid of the GIL will do wonders for this language and it's community. It will make it much more portable, too. Think of Java-level portability, but in a much nicer package.
A couple of them have valuations in the trillions of USD, actually.
(Mozilla does support a ton of OSS, too, so read that as everyone else needing to step up rather than an attack on them)
All I’m saying is that asking why JS performs faster than Python in some cases is less about the languages and more what you could do with hundreds of engineers working for years. If the stars had aligned differently and that effort had gone into Python (or Ruby, etc.) I’d expect a similar delta.
I’ve heard this “no C API” thing echoed by a couple people and it’s baffling. Do folks really think three major JS engines all written in C++ wouldn’t have an interface to interact with C?
Python's C API exposes ref counting and the GIL. It's also very large
JS doesn't have that problem -- more code is written in pure JS, there are no C/C++ bindings in the browser.
There are C/C++ bindings in node.js for v8, but as far as I know they are discouraged and not used very much. The bindings are more "first party" in node.js than third party.
They have issues, but not the same ones as CPython, because the API is very different
JS VMs must be re-entrant because they're embedded in browsers. That was never the case for CPython
Fwiw, this a fully stable and well documented API. It’s not even v8 exclusive, Bun cloned it for WebKit.
JavaScript doesn't interact with random C libraries that the user might have, like Python does.
JavaScript does interact with random C libraries :)
Fond memories: I did use pypy or a predecessor in like 2004 I think when I took part in a student computer science competition and my algorithm searching for subgraphs wasn't performing well enough to terminate in time.
Pypy, as a practical software deployment runtime has been and will remain esoteric (I do absolutely think that they had a positive impact on the wider python community, both in terms of dissemination of ideas and also practical engineering artifacts). But what's their market share relative to CPython? A thousandth of a percent or less? Has anyone actually built a significant business on top of Pypy?
There is, IMO no realistic path at that point that Pypy could become a viable CPython alternative. They are effectively competing with a hostile platform they need to maintain an extremely high amount of compatibility with, and that can and does move in directions that invalidate some fundamental engineering choices they make. Practical end results include that they're stuck w/ a lot of crippling design decisions (GIL, FFI API etc.) and a core part that is still in 2.7 land (RPython) and have historically mostly only been compatible to very outdated versions of python. This has improved a lot, but the next time CPython throws another curve ball the same thing is bound to happen again.
The only chance they really had was to be compellingly enough faster or otherwise superior that community pressure would have forced the CPython team into adopting a much more collaborative stance. That seems very unlikely to happen, now after all this time, given that now CPython is catching up and they have the albatross of the C extension ecosystem around their neck. Almost anyone who cares about python performance outside of algorithmic programming competitions will be using C extensions where Pypy offers no compelling advantage, some disadvantages, and by the momentum of the existing eco-system is unable to develop a superior alternative.
Isn’t mypyc effectively an alternative (AOT-compiled) Python implementation? Guido doesn’t seem too hostile too it.
If PyPy had a similar mode where you could load it as a library it would have a much easier time gaining traction.
Being dynamic make it harder to be fast, but JS/v8 is as dynamic as python, and much faster.
Those features are not in common use. Python’s hard-to-optimize dynamic features are more commonly used.
Javascript never had this problem, as all code was always written in Javascript itself by necessity, so it was far easier to optimize as you did not have to worry about backwards compatibility of the internals.
For C/C++ extensions, I think there may be hope to support a slow/emulating C API and a faster, less internals-leaking new API that extensions could adopt. It would take years for the migration, but if speed gains was say 5x, I think it could be realistic.
pypy managed to emulate the C API fairly well, after all. E.g. you can build numpy and pandas on top of pypy and it actually kinda works.
Our sloppy container spec bit us today though. We had
FROM: python:3-slim
with a bunch of pip requirements following. Some of those were not 3.11 ready, eg scipy==1.8.0, and our build broke. Our answer was to not be sloppy and pin until everything catches up, eg FROM: python:3.10.8-slim
and we're good. Hope someone sees this that needs reminding.A related note is for any requirements files. Something like this bit me the other day.
Libraryname >=3.1
After a few years,the package was updated substantially and has lots of breaking changes in the recent branch. Fix was to so ==3.1 until we work out the next step
Any reason to not use python:3.10-slim? That seems to keep up-to-date on patch releases.
We are almost daily discovering upstream changes like this one that breaks something N components removed so our kneejerk response is usually to pin aggressively when found and periodically upgrade deps for a whole component.
What are the chances I have some dep somewhere that says python<=3.10.8 and is working today but when that 3.10-slim spec allows 3.10.8 turn into 3.10.9 it will break? That's what happened today for scipy but on the 2nd int and not the third one, because we had started with 3-slim.
Is a Monty Python reference, what follows is actually, completely different. They did not lie :)
https://www.google.com/search?q=inurl%3Ahttps%3A%2F%2Fwww.py...
I’d be so curious to hear how much faster it is than Python 2.4 (back when I first started)
Big ask, but can we wait to release until X number of packages are supported before releasing? And stay in RC until then? No pyarrow :(
That said, Python releases generally don't introduce syntax-breaking changes. APIs are sometimes deprecated, but these have large windows that give maintainers plenty of advance notice. Years, typically. Practically all 3.10 code should run fine on 3.11, including your pyarrow.
[1] https://w3techs.com/technologies/history_details/pl-python/3
If you go back to https://pyreadiness.org/3.8/ there are 20% of packages not explicitly supported, many from the top 100 list, anyone using these packages know this is not a concern. (Going back even further in versions you start to see dropped support instead)
Deducting these 20% should give a more accurate picture, and even then I’d bet most other packages work anyway, like you say.
No, that’s silly.
(1) There’s no reason for something that otherwise is ready for stable to stay in RC because an arbitrary number of edternal packages aren’t ready to say they are ready for it, and
(2) No user actually benefits from X number of packages being ready, they benefit from the specific packages they depend on being ready. So its best for them to track them, and upgrade only when all the specific packages they need support the new version.
I get it, 3.11.0 is "final" in the sense of "definitive" from the development team's point of view, the final one of the pre-releases. But 3.11.9 is also called "the ninth and final 3.11 bugfix update" in the schedule [1], the actual final one from the maintenance team's point of view, in the sense there will be no more.
Can't we find better terms, that work for everyone? 3.11.0 stable? 3.11.0 actual? For anyone but the dev team, this is in no way a "final" release, this is the "first" release.
If you want to be pedantic, 3.11.0 stable would also be a bad name - what if it's not actually stable but crashes all the time?
And 3.11.0 actual, what does it actually mean? Is there a non-actual 3.11.0?
Maybe "3.11.0 public" or "3.11.0 release" would be better suited to distinguish between the various dev-builds and the "final" release of the minor version
> A version identifier that consists solely of a release segment and optionally an epoch identifier is termed a “final release”.
Well, people write configuration in JSON, XML, and INI and process those and re-serialize to those formats too
It's also handy for making one-off scripts for changes to a big-ish configuration file - even if it's not a part of some automated pipeline.
It's meant to be human-writable, and offers multiple ways to express the same table data; what format should a TOML serializer use?
I can see that you might want to programmatically edit a TOML file, preserving the rest of its layout unchanged. That's a bit fiddly to get right, and needs a different interface from a pure serializer. If the correct design isn't obvious, better to leave it out for now. It could still be added later.
I never met any format where parsing it is useful but serializing to it is not. Maybe some binary format where I only care for e.g. playing a video or decoding an image. For a configuration format though?
The obvious use case is reading, altering AND writing back.
>It's meant to be human-writable, and offers multiple ways to express the same table data; what format should a TOML serializer use?
It should just pick one and stick with it?
If N versions of the format express the same data (deserialize into the same structure) then it doesn't really matter from a functional way which they pick.
Users would still need to pick a TOML lib to write data - so they already OK with that lib picking a specific way. Why wouldn't they be OK with Python do it?
If they don't want nobody to mess with their hand-written TOML files, they can always just not write them back from Python.
I don't care about specific stylistic options in serializing to the format that are all parsed the same way anyway.
If N versions of the format express the same data (deserialize into the same structure) then it doesn't really matter from a functional way which they pick.
To clarify my point, if you’re using TOML, presumably you want to be able to format things nicely by hand because that’s the whole point of TOML. If you don’t need to do that, just use JSON. So I don’t see a whole lot of value in a tool for serialising data into boilerplate TOML (or YAML, etc).
For me the whole point of TOML is as a stricter, saner, YAML-type (human readable that is) config first, and a format used in several places, including upcoming Python standards, second.
Couldn't care less about formatting things nicely by hand.
That’s actually (if what you are reading is human-edited input thar you might transform but expect to be human edited) a hard case, because then you want to keep as much as possible the format entered in, while accommodating changes, which is easy to subjectively evaluate but seems potebtially hard to express mechanically in a way which would provide generally optimal UX. You don’t want to hand edit config and have a UI, but have an “I touched it with the UI and made a couple small changes out of dozens of entries and now its unrecognizable”.
In the face of ambiguity, refuse the temptation to guess
There is one way to read toml, but a lot of ways to write it (do you preserve formatting, what do you do with comments, do you allow partial updates...). Therefor reading is easy to get right, but writing less so, and once it's in the stdlib, we can't make quick changes to the API.
Since reading is already very useful, and tomli author's is ready to provide that for free, let's include it now and see writing later.
Then don't use TOML for internal stuff?
Or just settle on a way your lib writes it, and if every parser out there can read it anyway, and get the same deserialized structure, it doesn't matter which way you write it exactly...
As for it was a good idea, this has already been debated and acted upon, there is no need to repeat exactly the same arguments.
You will find the detailed rational in the PEP, and link to the various discussions.
Well, the same debates happened over many other decisions debated, acted upon, and later regretted and reverted. At some point the GIL seemed like a good decision too!
No need to project psychology on it...
Similarly, TOML writing is a ton of completely. The org isn't opposed to adding toml writing to the stdlib fundamentally, they just aren't rushing and want to hammer out all of the grossness first.
Or we'll have to be content with 3.14.1 as our pi py.
https://anaconda.org/conda-forge/python
https://github.com/conda-forge/python-feedstock/pull/577
Using mamba to create a new encoding called py311 with python 3.11:
mamba create -n py311 python=3.11 conda create -c conda-forge -n py311 python=3.11 pyenv install -s
pyenv exec pip install poetry
pyenv exec poetry install
That incantation goes into my README.md.Inconvenient, but I’ve come to accept it. I only need to run it after initial checkout and when the .python-version changes.
pyenv local 3.11.0
poetry init
poetry env use $(pyenv which python)
poetry installAs a former Homebrew maintainer, allow me to stress another thing:
Don’t use your system’s (or even your system-level package manager’s) Python environment for your Python projects.
That Python environment isn’t for you. It exists primarily for one single reason: to make other packages work that happen to depend on Python.
The same goes for Node.js, Ruby and other fast-evolving platforms.
Now, if you _do_ use that environment, the packaging police isn’t exactly coming for you. Just be aware that maintainers are free to version-bump or even remove the environment at any time without notice. That’s why you’re going to be happier and safer if you use *env- (pyenv, nodenv, …) managed installations.
This is a relatively new and dangerous way of thinking about operating systems. Due to the massive futureshock of libs rapidly changing under-foot people have had to switch to containerization as a mitigation. And now people have been containerizing so long they are starting to believe it is the proper way to do things.
It's not. Giving up the idea of an operating system with system libraries is very bad. The idea that you have to set up an entirely new lib environment to run every single script is absurd and it has dire consequences for software longevity and portability. With no more OS system as a base an OS is fractured into literally innumerable possibilities. Gone is the idea of a distro being the same for everyone using it. Gone is the ability to just install things. And we're left with a pile of containers that make debugging things when they go wrong even harder.
Nix is taking this concept to the extreme and absurd, but using pyenv for every script is almost as bad. pyenv is not version management approach to Python. It's a bandaid that doesn't address the actual issue.
One could as well say: this is a response to faster-than-ever evolving platforms.
> Giving up the idea of an operating system with system libraries is very bad.
Good point. I don’t like the situation either.
I’m not sure there’s a good remedy. For example, how are system package maintainers supposed to know that your personal script is now ready to migrate, so they can finally bump system Python?
> using pyenv for every script is almost as bad. pyenv is not version management approach to Python. It's a bandaid that doesn't address the actual issue.
Not sure if I’m understanding you correctly here. You’re saying that using pyenv for every script is bad. Are you against venvs, too? Because you could apply a similar argument to those.
If your project requires dozens of installed libraries to use even it's base features, that tells me a lot about you and your project.
Mainly, it tells me you haven't thought about stability and the future of your application.
Firstly, this is about the core platform, not third-party libraries.
Secondly, I absolutely do keep all my projects in lock-step with the latest stable platform version. My point is that I do that deliberately and in a controlled way, decoupled from my system package manager’s decisions.
I certainly don’t appreciate waking up to dozens of unexpected compile errors and warnings due to some random system Python version bump.
I get the reasoning but I disagree with this as a blanket statement.
For my own stuff, I stick with certain OS versions that I know and trust. If I standardize on Ubuntu 22.04 for example, I know that it will only ever ship Python 3.10 (plus patch releases). If I ever need another version, that's what the deadsnakes repositories are for. Ubuntu is never going to spontaneously upgrade the `python3` package to 3.11, especially not in an LTS release.
I understand the situation is different on Mac (and possibly Windows?) as they have a history of shipping outdated (and/or broken) Python environments and any system upgrade has the potential to bump your Python version.
Python's virtual environments may be somewhat clunky but it's entirely possible (and not even that hard) to keep project libraries and dependencies completely isolated from the system's with the various venv tools that exist these days.
At work, we have a large in-house ecosystem of scripts, modules, and packages written in Python so it's easier to tell developers, "pyenv is our supported Python environment, use anything else at your own risk."
You’re right of course; I’ve never really used Ubuntu, so I’ve never come to enjoy that level of stability.
I used to have a Mac, which is shipping with outdated platforms, and used Homebrew on it, which has a rolling-release model. I’ve switched to Arch since, which also happens to be a rolling release. That’s the context I’m coming from.
> Ubuntu is never going to spontaneously upgrade the `python3` package to 3.11, especially not in an LTS release.
Even then, your teammate may use a different LTS or even different distro than you do, ending up using a different Python version on your project than you do. I wouldn’t be willing to deal with that drift.
> Python's virtual environments may be somewhat clunky but it's entirely possible (and not even that hard) to keep project libraries and dependencies completely isolated from the system's with the various venv tools that exist these days.
Absolutely. However, even when using venvs, I still don’t want my venv to point to /usr/bin/python. Hence the `pyenv install -s && pyenv exec pip install poetry && pyenv exec poetry install` instead of just `poetry install`.
I know pipenv caught a lot of flack a while ago because KR tried to push it out before it was ready, but I like it.
If you mean finding an entry-level engineering position, then it's the same as any other job search. Leverage your network to find openings, build something that shows you know enough Python to be dangerous and put it on GitHub, etc.
Backend/web programming? Make a personal website with Django.
Data science? Set up Jupyter/Conda and do your own analysis/follow tutorials.
DevOps/Sysadmin? Try automating something with Python scripting. Requests, file parsing, running system processes and capturing output, etc.
They are all fairly different skillsets and require learning different niches of the Python ecosystem.
Horrayy!!
Does anyone has any experience with working with python in a very large scale?
Most of the big tech companies I worked with use Java, I remember when working with large scale JS projects it was nightmare to debug, TS came along and really saved the working experience and scale of JS/TS projects. I've seen TS adopted almost everywhere in big tech companies, some even create microservices with node.
When choosing tech stack for a very small projects I tend to use Django, but when I want to create my own large scale business one day I'm quite afraid to use it because of the typing problems, I might choose node/nestJS, but just wonder how does it work in large scale businesses that use python mainly? Is it a nightmare to debug?
A large majority of packages are still missing typing hints, but the ecosystem is moving in the right direction on that issue.
Without the tooling it's a nightmare. My previous Django project with approx 500k LOC had linting and some typing and that was a mess.
[0] https://github.com/octoenergy/conventions/blob/master/patter... [1] https://github.com/seddonym/import-linter
Regex is very specific. It has added records, streams, virtual threads (preview I think) working with native better, pattern matching...
Java is like C++. Everyone talks bad about it but at the end of the day it is one of the most Getting things done language.
I am not sure why I should use Kotlin. Llokd nice! But for sure docs and support are not even close to Java.
One excels at helping you shoot yourself in the foot. The other gets things done…
One could argue C++ gets even more done, since it is appropriate for use in more contexts, like firmwares/drivers, native SDKs, games, etc. Java's main use cases have strong competition in languages like C#, Go, C++, Python, etc.
Of course some people just rant about it. But did you guys use it a span of time long enough to really know how good or bad it is actually?
After using Kotlin in Spring projects for a couple of years, I rewrote most of it to Java 17 and stick to Kotlin strictly in Android work. It's not that much different anymore with all the recent additions to Java, and as Java progresses (and the main JVM implementation which is first and foremost being developed for Java) and diverges in other directions in some very important areas (value types, Loom), Kotlin starts to feel more and more like a second-class citizen. I mean, that's what some old-timers were preaching right here on HN from the very start.
Things like Hibernate require some pretty ugly (IMHO) hacks to work properly (you need something like three compiler plugins, and spread keywords around like there's no tomorrow).
The compiler was also very slow compared to Java. The same amount of code written in the same style took 2× times to build. This may have improved.
Probably the only thing I miss are nullability annotations (after getting used to them in C#, not Kotlin).
Edit: after re-reading what I wrote, the tone feels way too critical. Kotlin is a nice language, and certainly easier & more fun to both read and write (although it's trivial to write unreadable mess by going too far with its features). It just doesn't go far enough (for me) to warrant using a separate language with all that entails.
Java has adopted maybe 20% of the things that make Kotlin better, but due to backwards compatibility it'll never catch up.
Personally I think it'd almost never make sense to adopt Java over Kotlin for any greenfield JVM project, unless the devs you hire are truly so low-skill they can't pick it up.
Kotlin supports neither structural pattern matching nor guarded patterns as they are available in both Java 18 and Scala. It's Kotlin that needs catching up nowadays.
It's an abstraction atop Java. Of course it plays the catch-up game. It inherits new features from Java. They can't have it instantly.
> They can't have it instantly.
Why not if Scala did it way before pattern matching was introduced into Java?
Null-safety.
Read-only interface by default - `MutableList` vs `List`.
if-expression, they just look nicer than `a ? b : c`.
---
You're right that Kotlin is not an "abstraction atop Java (the language)". It's meant to be a better Java that feels familar to Java. With that rationale in mind, it's reasonable that they left some features from FP languages out of Kotlin.
As I have said in another reply, it was a great decision to not follow what Scala did with pattern matching.
`if (o instanceof String s) {`
Not much benefits over Kotlin's smart casts.
Guarded patterns with this version of pattern matching should be translatable to adding another "and" condition. The Kotlin compiler should have no problem reasoning with that.
---
I find positional destructuring in Scala (which works similarly in Java 19) a bad design.
That's what I get for not reading about it more carefully.
But I didn't investigate on which change exactly is making it slower.
There were type gymnastics you could do using overloads to reach any arbitrary level of coverage, but it was ugly and always short of fully general.
For huge projects, I still think that Java is a good choice, and although I have only professionally worked on one Haskell project (medium size), I think that Haskell might be good if a team is in place who can use it. A new friend of mine in town is enthusiastic about OCaml, and after a few evenings of studying, I wish that about 8 years ago when I started Haskell I had chosen OCaml for a production typed language.
For Python: I really like Python for deep learning, reinforcement learning, quick and small semantic web apps, etc. The common thread here is that I am not writing much Python myself, instead I am exploiting large well tested libraries.
Do you mind writing a bit more about why? I have been a curious bystander in OCaml land but some of the differences with Haskell, like the lack of type classes, have pushed me toward the latter.
I’m not even sure it’s possible to have Django typed without reworking the ORM, I’m thinking about reverse relations, .annotate(), etc.
Yes, there are type stubs for these libraries but they’re either forced to be more strict, preventing use of dynamism, or opt for being less strict but allowing you to use all the library features, at the cost of safety.
I think in the end, new libraries built with static typing in mind, like Pydantic, FastAPI, and Edgedb, are the answer.
There are type stubs for Django that somewhat avoid these compromises: https://github.com/typeddjango/django-stubs
To be able to do this they have to use a Mypy plugin though. And even then it's still far from perfect.
Why not invest into good boundaries and turn your large project into a group of small projects?
500k LOC project should have plenty of natural boundaries. A team should recognize and draw those, regardless of the language being used.
I recently worked at ~200k LOC Django project: the code was far from perfect, and yet I had no trouble onboarding new team members and making them productive. Here's an isolated 20k LOC domain, you'll grasp it in a week, you'll ship on your first day and then almost every day afterwards, and eventually your knowledge will extend to other areas. Isn't that how every big project should be managed?
Sure, things like strong typing do make the monolithic ball of mud more maintainable. But how about not building big ball of mud in the first place?
It's always easy to say "just be better and more diligent programmers", but that doesn't work. If the language promote spagetti, spagetti will be written.
Oh, I completely agree and I would never say that.
But at the same time, Java promotes complexity and overengineering. I've seen 10+ nested classes for something that was a 5-line function in Python.
The big difference for me is when I talk to Python engineer they agree that their spaghetti sucks. They want to evolve out of it, they just haven't found a way yet.
It is much harder to convince Java folks that their class hierarchies are useless.
Fixing Python spaghetti is way easier than fixing Java folks mindset.
Nothing in Java inherently does that. It's actually improved quite dramatically since Java 8, with many features like records, pattern matching, lambdas, SAMs, etc.
We get by, but it's pretty awful. People will give you all sorts of arguments for why loose typing and an NPM-like sloppy ecosystem are advantages for Python. And for the domains where Python really shines (i.e. devops, and manual data-wrangling), maybe these are advantageous. But what they DON'T tell you is that best practices for large-scale backend development call for bolting-on so many extra things, and accepting so many constraints, that it wipes out the purported advantages.
You'll need to use at least one linter, if not two, and a style checker to enforce the use of type hints. This gets you to something that is maybe 50-75% as effective as a Java compiler, but will never be any better than that, and will never have anywhere near the same level of IDE integration.
If you have a lot of dependencies (and you will), then you'll need a proper build system like Poetry. Which puts you right back in that Maven world, that escaping from was supposed to be one of the advantages. And the dependency ecosystem is SO... FREAKING... BAD. Unlike Java, where most of the major libraries are professional in nature with corporate financial backing, most Python libraries are "pure" community and written by unpaid volunteers or hobbyists. Many of whom have NO understanding of semantic versioning, so you find yourself having to pin ALL of your dependencies to specific fixed versions so the spider web doesn't tangle up. Upgrading anything is a nightmare.
You'll be told that it's easier to hire for. But then you'll find that the candidate pool is largely ops or QA/QE people looking to transition into engineering, boot camp grads, and other people with no professional experience in backend API development or even Python itself. Most capable candidates that you hire will take the job for other reasons (e.g. wanting some exposure to data science or ML), and will complain constantly about having to work with Python for general backend services.
You can get by, but as the project or company grows, you'll probably find yourself either re-writing Python portions, or at least deprecating them as legacy and eventually migrating to some other new greenfield thing. I can't imagine any sane reason to migrate a Java or .NET project in the opposite direction.
I have years under my belt now, and what I ended up preferring is C# and TypeScript. I would only build a product on one of those two. This random person's opinion is that your company would probably be best served by using Python for the data team and TypeScript for front and backend.
For scripts and small projects or rapid prototyping Python is pretty much king. (Although Red/Rebol is a contender. Check out their GUI examples!)
For medium-sized projects (100K LoC, 1-10 devs) you might use Go or you could experiment with the usual suspects: Lisp or Nim or Haskell or OCaml or whatever you want.
For a large project (millions of LoC, 100's of devs) I would use Ada or Java unless the problem was very Erlang-shaped (in which case you would use Erlang (or Elixir.))
For front-end work I'll personally never use anything but Elm for the foreseeable future. Based on economics. Compared to Elm the entire JS ecosystem is a incredibly massive boondoggle, a total waste. (I'm cranky this morning and I'm practically trolling here. Apologies to those who are not entertained. I mean it though: in my considered opinion Elm mocks JS, brutally.)
What exactly happens between 20k and 100k LOC (or 2 devs and 10 devs) that makes Go more suitable than Python?
What exactly happens between 100k and 1kk LOC (or 100 devs and 500 devs) that makes Java more appropriate than Go?
How do you even find this boundary when same apps in Java and Python might have 2-3x LOC count difference?
I've seen this mantra repeated for decades (technologies X/Y/Z for small/medium/large projects) and could never make sense of it. For example Angular was always sold as a technology for "large" projects (as opposed to React). Where is Angular now? All those "large" projects are now legacy that everyone despises. And if you have 100+ developers on a "large" Angular project I bet you're are feeling fucked now.
Since everything breaks at scale, I'm actually into the opposite mantra: only the tech stacks that are great at small scale are capable of being great at large scale. Note: capable, not guaranteed.
I always prefer layering complexity on top of a simple technology, instead of praying that complexity inside of a complex technology will perfectly match my needs.
> What exactly happens between 20k and 100k LOC (or 2 devs and 10 devs) that makes Go more suitable than Python?
> What exactly happens between 100k and 1kk LOC (or 100 devs and 500 devs) that makes Java more appropriate than Go?
> How do you even find this boundary when same apps in Java and Python might have 2-3x LOC count difference?
I've been thinking about those questions for roughly thirty years, and I still don't have exact answers. (If I did I'd be famous and hopefully rich, but that's another story, eh?) Speaking in broad generalities, the main thing seems to be something like the number of separate concepts you have to keep in mind to make (non-breaking) changes to the system. The more that the system you're using can do for you, the more suitable it is for larger-scale projects.
> I've seen this mantra repeated for decades (technologies X/Y/Z for small/medium/large projects) and could never make sense of it.
I gotta say, it's not a "mantra", it's experience. Have you not worked with various languages on various sized projects?
(In re: Angular, to me that was an obvious shitshow from day one. I would never hire anyone who admitted to ever thinking Angular was a good idea. Same thing in re: Heroku, while I'm at it. Lots of people do lots of foolish things when computers get in the mix.)
> I'm actually into the opposite mantra: only the tech stacks that are great at small scale are capable of being great at large scale. Note: capable, not guaranteed.
I don't see how that follows? (I also don't see how that's "opposite" to the other "mantra".)
Or it can be slower… depends on your specific code. For me it's slower.