In some sense - there are no large projects. Only projects that have failed at being modular.
Maintainability involves keeping clear boundaries between functional units so that you can keep enough code in your head at any one time to reason about it.
If we assume that Python is capable of handling more code than you can keep in your head at any given time, then surely the problem is simply one of making sure your modules are nicely self-contained from each other.
Write function in module A that expects some complex data structure that comes in. In a static typed language, you put a type on the argument and thus guarantee, to some degree, the data in the object can be safely manipulated.
Over in module B we call that function. Because Python doesn't enforce types, I can pass in some type of data structure that might behave well in many situations, but in some that function is going to fail because I didn't properly construct the data structure.
How do other Python projects avoid this? Do they avoid large data structures of multiple fields?
Let's question some of the assumptions. Do you NEED the entire thing to be passed around? If your functions are small and only do one thing, can't they be fed the portion of the data they need and only manipulate that? If they are getting unnecessary data, soon someone will like the convenience and that piece of data will now become necessary.
I have seen Java projects with humongous "DAO" objects being passed across all apps. Even if Java had Haskell's type system, at some point it breaks down.
So, you do not enforce that on Python itself. Rather, you have to apply due diligence and ensure your data structures are clean, and functions do what they are supposed to do and no more.
> I can pass in some type of data structure that might behave well in many situations, but in some that function is going to fail because I didn't properly construct the data structure.
Then you construct the structure in a single place, and that should have some sanity checks and sane defaults.
Even if you are using a static typed language, there's no guarantee that what's inside the data structure even makes sense, only that the right type of structure is being passed on. That's helpful, but not as much as people would think.
And you write tests. Lots of tests. It is amazing the amount of Python code that gets written without proper testing. And Java, for that matter. Only Python programmers have less excuses, it's easier to replace what you need for testing.
You should always put validation in constructor so you are 100% sure that the object created is valid.
I guess I'm arguing both sides now. My point is that one needs to root out the weird cases somehow and that it's often hard to have "100%" validation of inputs. One's initial understanding of what is the valid and invalid state space is often incorrect.
Also, even without actual type checking (mypy), IDE support + type hinting (PEP 484 or reST et al) means you can find most issues before the interpreter/compiler is involved. PyCharm is fantastic at this.
https://github.com/jonathanslenders/pymux/blob/master/pymux/...
Here's one way. Those asserts will typically pick up on errors during spikes, development and tests and turn a head scratching problem into a straightforward issue.
A couple of times on a disastrously tech debt ridden project I used it to pick up obscure configuration errors in prod, too (shouldn't see it in prod on a relatively well written project tho).
Never been much of a fan of duck typing. It only really makes sense when you have objects that approximate built in types.
Take a random object with a .run() method, for example, and you don't have a fucking clue what it might be doing so there's no point interpreting that as a quack.
Divide code into modules.
If you are worried about types, use object oriented programming to define the classes you want and then use the assert() instruction to check variables to comply with said classes/types.
A good IDE like PyCharm will take advantage of such assert() statements.
* Realistic integration tests that cover as many user stories as possible.
* Raise exceptions ASAP for any kind of invalid data or state (could be checking types, but could also be checking for file/directory existence). Could also be pip installing and using 'schema'.
* Do not build on poor quality modules and strive to decouple and uninvent poorly reinvented wheels.
* Just in general: loosely couple everything.
I don't see this as a static typing issue at all. Static typing only helps with a small part of the problem and it usually does that at some expense (usually more verbose & less flexible code).
I'm a little skeptical that super-strong type systems like haskell necessarily help all that much either. They clearly help eliminate classes of bugs, but the overhead is very high (the amount of everyday software written in haskell is seemingly small relative to its apparent popularity).