Python is two languages now
threeofwands.com
threeofwands.com
That's funny, because last time I had to write json-related code in Python, I really missed Rust's serde, which makes deserialization + validation much easier to write in Rust than in Python.
YMMV, of course.
That's not a gain in any way.
Since I/O is usually untyped, you need to cram it into a typed schema somewhere. In strictly typed languages, that must occur upon ingestion. The author is saying you can delay it by having a layer on top of the I/O transform the data in untyped fashion before being typed later on.
I have seen large applications (think hundreds of millions of daily users) break because of such oopses.
Libraries like Jackson make it entirely painless.
I would assume JSON parsing is a solved problem in any typed language.
record Thing(String name, int size, LocalDate when) {}
The only way I could see this being easier is to forgo type safety completely.These things are just better when you aren't in a nominatively typed language with some pretty crusty assumptions.
By contrast, duck-typing means that the intricacies of which exact type you are deserializing into is a lot less important in python.
How do you explain the popularity of type hinting in Python then? Type hinting arose _because_ duck typing is horrible for large systems.
I’ve come to Python 20 years ago after a short stint with PHP and I don’t really care if the incoming stuff I have to handle is a list or a tuple (let’s say) as long as it quacks and walks like the proverbial duck.
More generally, you could say Python the original language is the victim of its own success (the same goes for JavaScript in relation to TypeSctipt).
The only not so easy decision is: is it better to store as Unix epoch or as ISO string. That is a tradeoff between compactness and easy searching of date/time ranges.
Am I missing something?
That's a strange take, not one I've heard anywhere else. Rust macros are code generators - they're run once at build time and generate regular Rust code that's strongly typed. Maybe it's because the author has a different interpretation of "infrastructure" than the commonly accepted one.
However, a more technically correct interpretation would likely divide based on hygienic vs unhygienic macros. The C Preprocessor is most definitely a different language to C – it has the same differences as Rust/Rust Macros do, but additionally is implemented by a different binary, is de-coupled from the C language to some extent, and most importantly, has different syntax and semantics making it actually a different _language_ in the strict definition of the word.
> While these annotations are available at runtime through the usual __annotations__ attribute, no type checking happens at runtime.
At first I thought it was disproving the commenter's "strange take" assertion but it actually proves him right. While it is true that python's type hints are a build time feature like in Rust, they do not generate code for the interpreter in the way Rust does.
Actually they do not generate anything besides metadata that can be interpreted (at build time) in any way undefined way. The undefined is explicit.
I personally liked the article and agree that eventually the undefined specification of the PEP will be removed and a new language a-la Typescript will emerge.
If anything this is just a sign of a more versatile language.
import numpy
import pandas
...
and those that don't. Everything the follows from there might as well be two different languages. I write some 'python' code that starts with "import django" and some code that starts with "import numpy", and they're completely different in both structure and semantics. Being familiar in one style won't really help you much in understanding the other style.I see a potential separation as more: Application vs Data Science. Someone purely writing Python inside Jupyter notebooks for data science purposes tends to be writing a very different flavor of Python than I do while writing a web application. As least in my experience there is less concern for formatting guides, idiomatic Python, PEPs, type hinting, or things along those lines. And that is totally fine, its not something they need to really be concerned within that context. It does result in very different outputs though.
Not really. The thing with typing is that it gets hyped, but in practice, it matters very little. Both science code and web application code don't really need typing or really type hints.
If you go back to your CS professors and ask them to explain what is the advantage to types and OOP, the overarching reason would be that they create contracts in your code that reduce the risk of errors. In practice, if you look at well set up deployment workflow and pipelines, you still have unit tests/integration test that aim to have 96%-99% code coverage with sufficient test coverage, and people don't realize that it makes all the typing system moot. You can write equivalent code in any language and make it fully correct solely by ensuring that unit tests and integration tests for your usecase pass.
There is value however in extremely strong typing, where you can essentially do static proofs on the codebase to ensure correctness without running it.
I guess typing is moot when all the code and the tests are written, but having types sure does help while you're writing that code and those tests.
It helps to use the right data type when 'connecting' two objects. But it constraints the usefulness of code that can rely on the behaviour of the objects instead of the type.
It clutters the code: you have to read (and write) way more text. Which makes errors more probable.
It forces to think about primitive data types, so steals time maybe better spend thinking about higher data structures (list, dictionary,...) or algorithm.
YMMV.
One of the other differences I have seen, but didn't specifically mention, is unit tests themselves. The web app probably has a ton of them. The Jupyter notebook probably does not. Which is fine, the language is being used for different purposes.
But an advantage of having those types in place is that you can do the static analysis to ensure everything is being used correctly within your codebase. And then you can have your unit tests concentrate on the actual usage rather than other stuff. There is no need to check for / test with a null value when you already know that function is never being passed a null value. And adding those hints has helped me many times with catching edge cases before they even happen where the existing unit tests were never going to hit that edge case.
I find this sort of optimism both cute and frustrating. The reality is that almost no company is willing to foot the bill for such a level of code coverage as it makes writing the application take easily 2x as long. Sure, it'll theoretically be more provably correct, but the cost difference isn't easy to swallow. The reality is more like you have good path coverage for a small subset of business critical code with the rest of the code getting tested at a more macro level with integration and functional testing. And that's not even touching on the theory vs reality of how much better the code will be with those unit tests. Writing _good_ tests is harder than most people think. Writing tests that have good path coverage and good input fuzzing is _really_ hard. Writing a test that checks the box for 'code coverage' can be done trivially in many cases, but provide no real value.
This is only true if the codebase uses some highly inefficient setup that makes writing/running unit tests a pain.
And a general corporate policy for many companies, including big ones like Amazon, is to aim for >90% coverage with strongly typed languages like Java, regardless.
Where "science code" is "this python script I'm working on for my PhD and no one else is ever going to use it", sure. Where "web application code" is "I need to get an MVP out for Tech Crunch Disrupt", sure. Where neither of those things are true - if your grad student project code becomes the basis for sucessive grad students' projects, or gets popular in industry, or if the web application code underlies a company for a few decades - for sanity's sake you want type hinting. Tests only go so far, and if I have to run a piece of code's test cases before I know what the code actually does, instead of just reading it, something smells rotten.
That is, did you remember to check for null? Did you remember to write a test for floating points in languages that just have generic number types?
I don't see the value in relying on unit tests to enforce types.
It might be possible, but it relies on the programmer to:
1. Remember all of the types (was this an int? Or a float? Or a int | null union?
2. Implement exhaustive tests for each type
3. Remember to update the tests when something changes.
4. When reading code, if the types emitted are not specified, having to examine the unit tests to determine the type limits.
And the advantage of this is that it saves me a few keystrokes so i can write:
X = 5
Instead of
int X = 5
1. Language with strong type features.
2. Testing packages that automatically include fuzz tests and null tests to ensure correct behavior?
Granted most of my experience in production level software involves backend services, which essentially means json transformations 95% of the time, Ive rarely ever seen an error that results in the service saying everything is ok but returning a wrong result. For example in the case of null fields, something downstream fails and errors out - this is easily caught by testing.
However, I think tests are high-maintenance. Often they consume the same amount of time to write as the actual business code.
I don't know how automatic null checking tests would work. Wouldn't you need to specify what happens if a null is input?
In contrast, I find the cost if typing very low for the value it delivers. Unless the code is very ephemeral (like a one off script) then I highly prefer typing.
What static typed language, apart from SQL, does automatic null checking?
You can do a lot of productive work with a language without knowing everything about it. It might be safe to say everyone working with Python is aware of indentation, though.
Sometimes the disconnect between what people in the field are doing and what bloggers say is really striking.
I'm sure such people exist. I'm just surprised to see them on Hacker News.
To be clear, I don’t fault these people at all. They probably have other priorities in their work.
The answer to the mystery is: I don’t write production code. I write code that supports my testing and data analysis consulting work. There are hundreds of thousands of hackers, data scientists, AI people, DevOps people, etc., who get a lot of value out of programming languages just by using a subset of what they offer. They don’t study them as Zorro studies fencing.
Lots of people in the testing field dabble with Python without becoming experts in it.
It just strikes me, from time to time, how people really into something can lose touch with the majority of humans who are doing similar work.
Infrastructure by definition is supposed to be stable with well established interfaces, this is a perfect match for types.
Business logic is constantly evolving, focusing too much on types will cause you to be disconnected from business value and you'll just waste time fighting typing system.
> TypeScript… JavaScript… Going to guess it's similar.
> I haven't touched Java for almost a decade
> I don't really know enough Rust to be able to comment with any confidence
You would first add types to things that are obvious as you implement features or fix bugs and leave the rest to Any and then add more complex types or even construct types for your application.
Typed Python is not better or worse than untyped Python, it has different tradeoffs.
If you're setting a goal to "transition" your application, it means you don't understand what those tradeoffs are. You have a non-nuanced thinking is "types equals good" or "new equals good".
> then add more complex types or even construct types for your application
In other words, your intention is to add complexity to your application, while preserving it's functionality. Neither engineering nor business as a whole is going to benefit from additional complexity.
The "edge case" where complexity can be beneficial is where the ratio of library maintainers to library users differs by orders of magnitude. E.g. if you have a library with 5 maintainers and 1,000,000 users, then types can be a net positive.
The problem I have currently with Python is that all these options are increasing the number of "pythonic" ways to do something, and making the language more complicated.
Python was exciting because it was quick to write, and elegantly simple. It didn't make exciting and risky choices on syntactic sugar -- it took a moderate approach.
But I get the impression that Python is trying to evolve into more and more use cases that it isn't well suited for. I think this runs the real risk of becoming too generalized -- doing many things, but not doing them well.
Node Version Manager - POSIX-compliant bash script to manage multiple active node.js versions
https://github.com/nvm-sh/nvm
Volta - The Hassle-Free JavaScript Tool Manager
https://volta.sh/
asdf - Manage multiple runtime versions with a single CLI tool
https://asdf-vm.com/
Eventually we might need a node package manager manager manager to make sure we are using the right one for a project lolI also don't get what you mean by "dumb byte strings", it's just that the default* string type, i.e. when you have a literal `"foo"` in source, it's a sequence of codepoints vs. a sequence of octets. Assuming file systems use UTF-8 was broken in very early 3.x (I think before 3.3?) but it's something I've never struggled with (pure ASCII filenames for me). Forcing the programmer to know if they're dealing with characters & code points vs. bytes/structs/network buffers is a good thing, IMHO.
What a weird take, unironically asserting that the only "exciting" code is the glue, and all of the actual things software does is just an excuse to build more frameworks.
The vast majority of truly impactful code is really boring at a technical level.
So, yes, that’s sometimes why people write frameworks: to get some excitement in their code.
calculating that 38302 + 8204 = 46506 is pretty boring
noticing that ∀x,y: x + y = y + x is more exciting, and it's a lot more exciting when you can prove it
group theory? continuous braingasm
changing the report code so that an invitee who rsvps to an event even though the host has marked them as 'not coming' still shows up as 'coming'? that's pretty boring; it solves today's problem
writing a practical extraction and report language so that you need only a third as much code to define that report, but also all the future reports you write? that's exciting as fuck, it makes the rest of your life better
generality in software is highly desirable, just as in older branches of math; it's the difference between subsistence and progress
With more complex types, I've found it isn't necessary to do anything more complex than specifying the element type of a list or dict most of the time. I have typehinted lambda function arguments before, but just using the plain Callable typehint is usually enough. When I forget if I need to use an Iterator or an Iterable (which is every time), then I just try one and run the type checker and change it to the other if i guessed incorrectly.
Type checking Python can become complex if you expect to be able to express everything in the typehints, but Python's type system isn't powerful enough for that. I feel its strength is that it lets you add as much detail to the types as is convenient.
The author means that in the sense of there being two distinct camps in the Python community, but using type hints is not exclusive in any way.
Gradually typed [1] is the recognized term.
But YMMV, I know very little about Python so I can be persuaded this is no big deal.
The author's take that annotated code and unannotated code are different languages is a bit silly. It's like saying code with comments is a different language than code without. While it's true that different camps favor annotations differently, but the same is true for code documentation, and we don't treat it like different languages, even if the comments are in a language you don't speak.