Tauthon: Fork of Python 2.7 with new syntax, builtins, libraries from Python 3
github.com
github.com
It was a fun project; I learned a lot about how the CPython implementation works and have a lot of respect for the people that built it. It was surprisingly easy to implement Tauthon based off the work the core dev team did on Python3: https://www.naftaliharris.com/blog/nonlocal/
For what it's worth, I do believe that Python3 is a better language than Python2. We use Python 3.7 at my work (SentiLink) and we've had a good experience with it. (If you're starting a new project or can migrate, I'd recommend it). But I do think that the ~10 year saga of upgrading to Python3 from Python2 wasn't necessary when the main benefit was really the unicode refactoring.
I no longer maintain Tauthon personally but there are others who are excited about the project who occasionally add new features or bugfixes.
the main benefit is actually sane exceptions. unicode is nice and all (i'm a native speaker of a non-english language) but it's a storm in a teacup IME.
When you launched it, you called it “Python 2.8”. You posted it everywhere to gain traction, and didn’t rename it until the PSF and Guido got you by the ear, so to speak. There was no mention of “everything but str” or whatever, as far as I recall.
It was an outright (and hostile) attempt to fork - something that, going by your words here, I guess you now recognise as a mistake. I guess saying “I screwed up” is hard.
* https://news.ycombinator.com/item?id=13144713
* https://github.com/naftaliharris/tauthon/issues/47#issuecomment-277081725(friendly note: code blocks break clickable links)
1: https://news.ycombinator.com/item?id=13144713
2: https://github.com/naftaliharris/tauthon/issues/47#issuecomm...
This matches the pattern of Python.org developers and python 3 aficionados being unnecessarily hostile and condescending to the concerns of Python 2 language users. You saw that in 2010; you saw it again in 2015; and you can see it in these threads today.
If people still want to work on Python 2 projects that's fine and it's their choice, but it's time to just let the rest of us go our own way.
It's certainly possible that there are parts of the migration that would be tricky, a quick skim of the file didn't give me any obvious ones, but it's also huge and hard to read, so I very well could have missed something.
Most of the truly challenging things to migrate involved some combination of extension modules, heavy metaprogramming (eval/exec), and apis which change significantly between 2 and 3 (most of which are string related, but some libraries also decided to do backwards incompatible things)
> gvanrossum commented on Dec 10, 2016
> Since I was asked: The project's name (and its binary name) need to
> change. They are misleading. The rest looks acceptable according to
> Python's license. This is not an endorsement (far from it).
> naftaliharris commented on Dec 10, 2016
> I don't mind renaming this project. Any other suggestions for good
> names? I personally like "Pythonesque (/usr/bin/pesque)" the best so
> far, thanks @dbohdan! :-)
>
> @VanL, not that I'm necessarily picking that, but would a name like
> that be acceptable?
When you're in a community or a space, actions that are unintentionally hostile towards that community or space, can be seen as intentionally hostile by the members, and if you only hear about it second-hand, or you spend a lot of time around the group, that belief can be reinforced through the discussions and gripes the group has about it(1).(0): https://github.com/naftaliharris/tauthon/issues/47 (1): That's not to say there aren't bona fide intentionally hostile actions that happen, but rather that's how actions that aren't intended to be hostile can be remembered and percieved as such.
I'm not saying it was entirely his fault - certain widely-heard voices in the community had been advocating for this to happen, in practice, for several months; he saw an opportunity and went for it. I just object to the rewrite of history to justify the mistakes of the past.
Have you considered that not everyone who forks something participates in the original community?
> I just object to the rewrite of history to justify the mistakes of the past.
So far there has been no evidence for the stated claim, just supposition and rumour. So as-is there's no reason for anyone here to believe that "history is being rewritten" aside from easily-mistaken word of mouth.
These days most discussions over the internet happen via the written word, so it's difficult to believe that you can't find records from IRC, Github, or Email to support your contention that it was hostile.
At the time, the author stated:
> I picked [the name "Python 2.8"] initially since when talking with friends about this project it conveyed pretty darn immediately what the project is and does. I'd be very keen to hear people's suggestions for alternate names!
This was no mistake or screw up at all. In fact, the project served his purpose for the time, and even though he's moved on, there are others who like it enough to maintain or improve it.
A fork is (or can be) a healthy, natural thing that happens.
Do you think it would have been feasible to make wheels compiled against the python 3.x C APIs also compatible with Tauthon?
It’s a tragedy of the commons: A lot of companies making good money are still using Python 2, but none of them are willing to pay programmers to maintain an open source currently maintained Python 2 fork.
It was no secret that everything else was portable, that's essentially all what python 2.7 was - backporting 3.x features to Python 2. It was a waste of resources for developers to maintain 2 forks of Python so that effort was stopped in 2015. There was 5 extra years to move application to Python 3.
5 years is frigging long time in computer terms, and your project felt still like giving f-you to the core developers for trying to improve the language. Trying to call the project Python 2.8 was very aggressive and would create a lot of confusion if Guido would accept that.
I'm glad you gave python 3 a try, I feel like the people that were against it didn't write any new application in it, and their experience was porting Python 2 code to 3. It can be very frustrating when Python 3 complains about bugs that Python 2 just ignored, but if you write a python 3 application from scratch you don't even notice the Unicode, and that was the goal.
But I agree, the warning would be a good idea.
Which aspects?
I've worked in and around Python for more than a decade, and I've yet to meet a single developer who didn't initially dig their heels in over `print` being changed, and then later realize that they were flat wrong to have dug their heels in because the new way is measurably better.
I mean, I guess I hadn't until today.
It would be bonkers for print to be a unary operator because it's just a side effect, and so I feel that your point in this situation is correct where it would not be in your other examples.
python-modernize --write --no-six PROJECT_DIR
(https://python-modernize.readthedocs.io/en/latest/) futurize --write PROJECT_DIR
(https://python-future.org/quickstart.html)old.py:
bs = raw_input()
# Call a 3ps which originally accepts a bytestream, but now accepts a string
some_3rd_party_function(bs)
After python-modernize: from six.moves import input
bs = input()
# Your code using bs as a bytestring still works
# Runtime Error:
# This now accepts a string, and you need to modify your code to deal with random TypeError popping from everywhere
some_3rd_party_function(bs)In both cases, as shown in the linked documentation, these tools are designed to support staged migrations for exactly this reason.
...wrong, and probably also mythical. There's no excuse for having it as a language directive rather than a function call. I don't believe that any human who's spent even a few seconds thinking about it actually prefers the former over the latter.
It's supposed to be a convenient, quick language for scripting. There was no good reason to force some academic consistency on the print call. It's extra keystrokes for no reason. It makes me sick every time I have to use the parentheses.
Yes, you can do the same thing with a kwarg on the print function, but it goes at the end of the line and ends up substantially messier in practice. This is the #1 reason that I still write Python 2 every single day; it makes a huge difference for code generation tasks.
p = partial(print, file=fp)
p("foo")
p("bar")Maybe in Python 4 we can have Nim's approach[1] (allowing both) and another decade of arguing about it.
[1] https://nim-lang.org/docs/manual.html#procedures-command-inv...
I'm not a python dev btw. I've mostly used go/rust/swift where assignment expressions are a big thing so I am a bit biased
FTR, I could care less about the walrus operator.
It's been more than a decade now. I'm almost tempted to think that anyone who hasn't started porting their code simply isn't going to.
Their point has been made, but it's not like the Python community is going to unmake those changes. If you're starting a project, you should definitely use the current version of a language that is supported by a zillion developers, rather than picking a (practically experimental) fork of the old version.
That said, I understand there are certain applications that are not compatible with 3.x and the company does not have the resources to dedicate to rewriting it. So let’s suppose there is a valid reason someone is forced to use Python 2.7. In that case, the number one priority of this fork should be back porting the security fixes.
You can live without async, type-hinting, and f-strings. But please be responsible when it comes to security vulnerabilities.
This isn't a great metric for evaluating Python 2 forks, that's all.
[1] The three states being 2.x prior to recognizing bytes, 2.x sorta recognizing, and 3.x hard recognizing.
There hasn't been a main branch commit since python 2.7.17 was merged in ~6 months ago.
There also are platforms where new Python is not and not going to be available, e.g. DOS.
What piece of code is taking advantage that print is a pure function nowadays?
[EDIT] I am genuinely curious.
Edit: no, the author built it as a proof of concept.
Does this use SIMD instructions?
It's not that big of a project, even for a moderately large codebase. If you think you can't get it done in a reasonable period of time feel free to hire me as a contractor and I'll knock it out.
Do you mean calling unicode('some non unicode string')? That just uses system default encoding. ( sys.setdefaultencoding() ). Just find and replace them and slap in a .decode('UTF-8') or whatever your default encoding was in python2.
Grep for the strings encode, decode, unicode and just mechanically fix them one at a time by making the old implicit behavior explicit. How many times could you be doing that anyway? A few hundred? A thousand? You could even script this pretty reliably and just page through the diff you end up with to eyeball them one at a time.
I guess you might mean 'str' + u'unicodestr' or something, but again you can find these pretty easily by rooting out where the non-unicode strings are being produced and fixing the problem there. They are either literals or they are coming from IO or calls to str, right? Anyway, I've done this quite a few times and the main concern I've always had was trying to get the patch in place before people commit too much stuff for me to be able to merge the fixed up branch.
Of course, you could just do it little by little by taking out places that you are relying on systemdefaultencoding by monkeypatching the default decoding function in sys to log tracebacks whenever it is used, and then whacking them as they come up so you end up with properly handled and explicit unicode decoding before you move away from python2. I bet you could find and fix 95% of the cases in a day of effort.
And yeah, u"" literals aren't hard to find, but the problem is when you get data flowing from different sources, so both sides are variables. For example, one is read from a text file, and another one comes from parsed JSON - so the former is raw bytes, and the latter is Unicode - and you need to combine them together. Like you said, the proper way to do this is to ensure that as soon as data crosses the I/O boundary, it should be of the correct type (i.e. unicode rather than bytes) - which, ironically, is exactly what Python 3 encourages with its changes. But it can be hard to find all such places - you have to actually audit every use of I/O one by one, because in Python 2, the code by itself doesn't always reflect whether it's supposed to be dealing with text or binary data.
Strings, most things are generators now, exceptions are different, relative imports changed, division, etc. On top of that, moving to python 3 requires updating your dependencies.
https://www.python.org/dev/peps/pep-0373/
Edit: I can't find a way to delete this, but I don't think it's fair to be downvoted because I didn't get it was a sarcasm :(
(I got the sarcasm)
Modern software isn't fire and forget. Developers who have grown up around security exploits understand that. Higher management and FORTRAN77 programmers don't. To them it's developer busywork.
The real judgement being made is money and resources.
[citation needed]
I have encountered exactly zero bugs in Python 2 in production. Code doesn’t magically stop working
Your locally running, toy projects will continue to run under Python 2 indefinitely. Large [esp network-facing] projects will start to fall apart.
I predict that Python 2.7 will continue being supported by some capable organization for at least the next 10 years. As a lower bound, Red Hat has promised to continue supporting it for RHEL 8 customers until at least 2024.
Many deployments rely on more than a supported cpython binary. External packages and the rest of the server stack. The Python community has largely dropped Python 2 support. If you don't need that, great. If you do... Then what?
You could pay people to support your entire stack. You could upgrade. It's that choice that makes Python 2 dead.
You can still run Windows 2000. Doesn't mean it's not dead.
Unnecessary changes and the overall fragility of "modern" software takes money and other resources away from dealing with actual security issues.
There's no reason why things like a library of geospatial calculations (distance between two lat/lon pairs, etc) should need constant maintenance. They could be written once and then used for decades.
That's the biggest problem with the "old" ways of doing things, you simply don't know where you're vulnerable. Refusing to allocate resources to revisit old code, only to add new features is a recipe for disaster.
Upgrading to Python 3 does not fix this alone, but ignoring it, and the other dependencies that you might have, and that old Redhat 7.1 box running Linux 2.4 that's all probably fine, it's just card processing, right? Nothing else has changed.
The "modern" ethos is testing everything all the time. Yeah, it's a shed-load of extra stuff if it's new to your project, but it's only codifying things that you should be doing anyway, even if only seasonally.
Yes, there is occasional fragility, but that's more than offset —at least in my network-facing world— by the constant stream of security fixes across my stack.
Sticking your head in the sand won't protect you.
I have code bases that will be forever stuck on 2, because they're not active projects and are archival only.
No it wasn't; it was the better part of a decade. Look, we all know Python 3 sucks (even if some poeple, unfortunately including the former Python maintainers, refuse to admit it), but this sort of blatant falsehood helps noone.
Everyone had 10 years of warning.