Open-sourcing MonkeyType – Let your Python code type-hint itself
engineering.instagram.com
engineering.instagram.com
I must say this is the first time I've been disappointed with the quality of discussion on HN. For a community that promotes using the right tool for the job at the time, I would have thought people would be more open to the choices the early engineers made. I'm sure that Instagram are using a variety of tech across their stack.
For a long time there have been efforts to ensure at least a degree of quality and robustness through processes, practices and verification tools. One such tool is a type system which allows encoding requirements and expectations that will be automatically verified with the help of a checker. This is exactly what Instagram is attempting in an effort to increase quality and ease maintenance.
It seems many are wondering why they haven't done this from the beginning. A typical (and probably correct) answer is development speed and flexibility: early stage companies need to be nimble and compromise on quality if they want to survive. Fair enough, I understand that. We can applaud Instagram's success on the markets, but that doesn't have to mean that they're a role model of technical excellence for building a million line Python code base.
This tool is the proof that Python has significant problems at scale, which is something the Python community has denied for a long time. They're still doing it in this thread, but the lesson looks pretty clear to me: if you plan on building large scale, don't use Python. Or PHP while we're at it (see Facebook). Definitely not JavaScript (see FirefoxOS).
The founders of a start-up can continue to do whatever they want in the name of success. Instagram at their beginnings was basically a different company from today's Instagram, no one would have used them as a technical role-model. Now they're at best a role model for large Python code bases, but many here seem to be drawing the wrong conclusion, namely that it's a good idea to do large-scale Python in the first place.
This does not prove your point. Annotating a large dynamically-typed codebase with type information is a large amount of work, regardless of the language. This tool makes that easier.
It seems to me it's so difficult to manage such a code base, that they decided to do that "large amount of work" in addition to the large amount of work required to develop the necessary tools!
In which case it seems prudent to avoid the said amount of work by picking another programming language for one's large-scale code base.
> In which case it seems prudent to avoid the said amount of work by picking another programming language for one's large-scale code base
When I first saw Instagram I thought it was nonsense, I was wrong. When the founders started Instagram, I wonder if they had the amazing foresight to see what it would become. I would suggest getting to market was more valuable to them than worrying about what maintenance they would have to do once they had a billion dollars
While I'm a type-adherent in my day-to-day (F# represent, wut wuuut), I think this isn't reflective of the chicken and egg dilemma for startups... There are an immense number of things they "should be" doing at scale that they can't do early because they're relying on their early product to scale. Time to market, and windows of opportunity, are critical to startups.
Any immature technical decision at that part of the lifecycle needs to be made not with an attempt to make perfect forever from the start, but rather with an eye to transitioning to smarter solutions aggressively as you scale up. The company may be 6 pivots away from success, so better to validate solutions in the market than prognosticate.
I don't accept the premise that a lack of typing significantly improves development velocity, per se, but language decisions are about ecosystems, key components, and local talent. Where these companies get into coo-coo land is not integrating those immature components into better systems as they're getting bigger. Next thing you know someone is writing a whole compiler chain for PHP in an attempt to reinvent a sane programming language, or trying to get Python to be Java.
Smart, modern, functional languages provide development velocity and pleasure equal to Python with fundamental type safety and guarantees. But in a world where just putting a button on a webpage means multiple dialects in multiple languages I think we should be ok mixing and matching on the backend to scale smartly.
Putting aside concrete technical issues regarding the python runtime’s performance envelope (eg: startup time, FFI inter-op call time, etc) and memory footprint, there is no reason not to use Python. Again, concrete technical issues aside, Python will not be an issue until your business is big enough that having to push the performance of your product is a nice problem to have and you will probably have the money to spend solving it by either optimising your python or rewriting parts of your application, as all these large companies have done.
0 - https://www.linkedin.com/pulse/top-10-sites-built-django-fra... 1 - https://worldwebtech.weebly.com/blog/top-ten-most-popular-we...
Development speed, maintainability, error count are also important development issues. Unfortunately we don't have much data to judge, but the little that we have such as this article indicate that dynamic typing has a non-negligible negative impact on the above.
""having to push the performance of your product is a nice problem to have and you will probably have the money to spend solving it by either optimising your python or rewriting parts of your application, as all these large companies have done.""
Not considering the topic at all is negligent. I don't understand why you're so sure that the only possible outcomes are either not reaching scale or having the money to optimise or rewrite.
First of all, not even Facebook had the money to rewrite their PHP code base, so that's probably out of the question. And it's very well possible that one will reach scale and not have the money (or worse the time) to optimise, if optimising means inventing type checkers.
At least one should spend some time thinking about this topic and picking a language that is flexible and can scale at least somewhat. Having to stop writing production code in order to invent a Ruby static type checker or native compiler is decidedly not a good problem to have.
It depends on how you define technical excellence. I would suggest a low memory usage might get you a thumbs up from HN, but isn't real technical excellence the ability to solve users problems?
In my own experience I see it on a spectrum: there are design decisions you can avoid by using dynamic typing that give you speed in the small that start to show their absence in the large. Savings on inputs (quicker coding), tend to evaporate as complexity grows because you're using so much time analysing outputs (the system), to determine behaviour. Every bit of up-front work saved gets amortised across issue after issue, and runtime testing during development, and increasing stagnation in high-level architecture.
For me the ideal is the type system of Haskell with the linguistic power of Haskell and the type inference of Haskell... only on a mainstream platform I can convince management to use.
Rarely is a programming language chosen because it actually is the best tool for the job. Often, it’s whatever the people at the ground floor were most comfortable getting a prototype out the door with. Prototype working well? Fix this bug, add yonder feature. Before you know it you’ve built Facebook in PHP. “PHP is at the very least suitable for the development etc etc”, yes, that’s why Facebook invested all that effort in Hiphop VM.
There’s an irony in saying “there’s no reason not to use X” in reply to an article that is about a company spending a ton of effort working around a problem in X.
My point is: as a community, it’s our duty to learn from these mistakes. Let’s admit there is a problem, investigate, adapt, overcome. We can improve whatever will be the next Python so the next Instagram doesn’t have to go through this. But that won’t work if we keep saying “this is a nice problem to have, there’s nothing wrong.” It isn’t. There is. Look at the article.
Yeah, absolutely, don't do this! These are two examples of successful companies that did it and look at them now! </s>
Being able to move fast and produce a winning product on time is much more important for startups. What does it matter if you used <your_cool_scalable_thingie> for a project, when it never went past 10 users because you were concentrating on wrong side aspect of your business? PHP is fine. Python is great. Use the tools that fit your problem and you know how to use, not the latest toy.
It seems that this same pattern plays out with many tools, and not just languages. When you've built something and you now have a team, processes, etc. built up it becomes difficult to see the forest for the trees, or to make the hard decisions because it might involve replacing people.
1) That market success implies having quality software. Average seems to be enough in my experience.
2) That start-ups are a good example to follow if one wants to achieve good quality. In fact they should be ignored, because they will absolutely murder quality in order to stay alive. Sometimes the product doesn't even work and is held together with duct tape in order to get past that important demo... It's quite pointless to discuss quality and start-ups.
The lesson I mentioned should be heeded by mature companies that are able to do some project planning, complexity estimation, etc.
"Scale" can be seen under different angles. You can run thousands of boxes with relatively simple code if it's designed to scale horizontally. Twitter used to run on Ruby this way, until they accumulated money and expertise enough to rewrite the whole thing and save on operations and further development costs.
Your code can be millions LOCs and run on a few boxes; it's a very different sort of "scale".
In one company where I worked they had a crazy mix of codebases, from modern Scala down to ancient PHP code. But since it was architected reasonably, it was possible to replace the PHP code piecemeal, without stopping the system. Do the devs that started it 15 years ago with PHP deserve blame for choosing a poor language, or praise for coming up with a serviceable architecture?
You can see the original post as a step in the right direction: in a complex codebase, static typing has a large number of advantages. Barring a wholesale rewrite, how would you gradually transform your code to using it? Yes, by documenting the current state in a formal way, introducing a typecheck build step, then maybe splitting out certain components and rewriting them in other languages, etc. Look at TypeScript.
Unfortunately, there's no way around using quick-and-dirty prototypes at the very early stages, if nothing else, for lack of technical expertise among founders and their first employees. They have business ideas first and foremost. So a tool like that would help exactly the step you want done, switching to a nicer and more manageable stack as growth allows / demands it.
Would be interesting to see how MonkeyType and PyAnnotate compare.
I'd also like to know how much quicker or better it could have been completed if it had been done out in the open.
I assume they scratched their own itch so it was probably faster to do it on their own addressing their specific needs.
We knew we're going to open-source each others' implementations, esp. that Instagram's is focusing solely on Python 3 which isn't useful for Dropbox at the moment. It just took a while to get through the process of open sourcing what we had (cleaning up the early implementation with limited documentation, decoupling from internal data stores, etc.).
Would it be cheaper if this started out in the open? Probably, but I don't think by quite the margin as you expect.
PyAnnotate couples re-applying the types with the tool, and uses type comments for this.
MonkeyType generates .pyi files that you can either use directly or re-apply them to your code as proper type annotations.
Other than that, MonkeyType is used on a daily basis internally and solves a bunch of common annoyances of systems like these, like duplicates in unions, applying better types to stuff that was already hand-annotated with Any, etc.
I do find it intriguing though, that adding back types manually is so hard and slow. Is it slower when done retroactively? Or is it just as slow when done at the same time, but we don't realize its overhead?
http://preshing.com/20141202/cpp-has-become-more-pythonic/
I have been using C++ a lot lately but really wish there were more tools for reflection at compile time, e.g., ability to iterate over all the members of a class. Other than that, I'm really loving C++17's auto template parameters and type deduction capabilities, plus code that's 200x faster at runtime than most interpreted languages. I've found autocompletion in CLion to be slightly better than autocompletion PyCharm, but not quite as good as IPython or IntelliJ with Java.
sorted([(k.weight, k.name) for k in somelist], reverse=True)sorted( (k.weight, k.name).for_(k).in(somelist), reverse=true));
You would never be able to get that past code review, however.
The other thing to note is that
sorted([(k.weight, k.name) for k in somelist], reverse=True)
is essentially already typechecked: def biggest_ks(ks: K):
return sorted([(k.weight, k.name) for k in ks], reverse=True)
The above code now has all the same type guarantees as your c++, actually maybe more since the macros you use are going to be...uhhh, mysterious.My point was just that you can implement almost whatever you want (even without macros, they'll just expand the design space).
std::vector<std::pair<float, std::string>> in_order;
std::transform(data.begin(), data.end(), std::back_inserter(in_order),
[](auto& item) {return std::make_pair(item.weight, item.name);});
std::sort(in_order.begin(), in_order.end(), std::greater<>());
While I won't claim it to be as elegant as Python, it doesn't seem too ugly. Does anybody have an idea about how `auto` could be utilized to avoid the long vector<pair...> type declaration?For the adventurous: https://repl.it/repls/RepentantGentleKudu
Pulling things out of tuples isn't great, though.
Right now, I think that's approximately,
vector<tuple<int, string>> output;
transform(
somelist.begin(), somelist.end(),
back_inserter(output),
[](const auto &f) { return make_tuple(f.weight, f.name); }
);
sort(output.rbegin(), output.rend());
Ranges, I believe, would reduce this a lot, possibly even to a single line. If I am reading the docs on it correctly, something like, vector<auto>(somelist | view::transform([](const auto &f) { return make_tuple(f.weight, f.name); })) | action::sort;
I think.Two notes, however:
1. I feel like most of the desire for a static language is to know what type something is. Is C++ exactly as brief as Python? No, as I think you've demonstrated. But I think you're a lot more likely to know the type of something. Rarely do I think I find that Python has annotations, and annotations can be wrong.
2. C++ is, in general, I feel, much more explicit about where copies occur. I elided one of the copies in your example, opting instead for an in-place sort (but this is trivial to fix in the Python).
Really? O.o How could this possibly work?
Sure, but you don't have to choose between them, there are plenty of languages where you can have both Pythonic terseness and full type safety. E.g. Scala:
(for {k <- somelist} yield (k.weight, k.name)).sorted.reverse
Many other strongly typed (ML-like) functional languages are similar. somelist | transformed([](auto &&k) { return make_pair(k.weight, k.name); }) | sorted | reversed; someList.OrderByDescending(x => x.weight).ThenByDescending(x => x.Name);
It's quite similar to expressing the concept in English, certainly more than using list comprehension in Python.And how would you order it in Python by ascending on the first field and descending on the second using list comprehension?
In answer to your question though,
sorted(((k.weight, k.name) for k in some_list), key=lambda x: (-x[0], x[1]), reverse=True)
appears to work. This does use a non-obvious trick, but being more explicit is a smidge difficult, since the key function is called only n times, as opposed to O(nlogn) in the C# example.Alternatively, you can use
sorted(sorted(((k.weight, k.name) for k in some_list), key=lambda x: x[1], reverse=True), key=x[0])
Which is more like the original example, and if you're doing it in place, you get outs = [k.weight, k.name) for k in some_list]
outs.sort(key=lambda x: x[1], reverse=True)
outs.sort(key=lambda x: x[0])
Python's builtin sort is timsort, so despite sorting the list twice, this will still run in approximately NlogN comparisons, not 2NlogN.You could also manually define a custom comparator, ie
lambda s, o: (s[0] > o[0] * 10 + s[1] < o[1])
and pass it to `functools.cmp_to_key`. std::vector<P> people = { P{"jane", 47}, P{"mary", 71}, P{"john", 65} };
ranges::sort(people, [](auto& x, auto& y){ return x.weight > y.weight; });
Compile this gist with `c++ -std=c++1z -I range-v3/include`:https://gist.github.com/cieplak/dcd587c67d989768900e4110e776...
Picking C++ over Python is like picking woodworking over metalworking.
Python being slow is never an issue unless someone is insisting on using the wrong tool for the job.
So you think, for example, Numba (and everything that uses it) is misguided?
Numba could be seen as misguided from some points of view. E.g. when using Python for high performance scientific computing, you will typically be writing your computational kernels, I/O etc. in some compiled, superfast language (C/Fortran/CUDA/whatnot) and all the input handling/case setup/etc. in Python. If 1% of your compute time is spent Python and 99% is carefully optimized C, Numba is obviously pointless.
But that's for one application. Python is used for so many different things that you can't make blanket statements like this.
What language do you use where you can get these kinds of guarantees? As far as I know very few languages provide those kinds of dependent types statically.
Typechecking is very useful here: if you try to transform a point in space represented by a Vector3 by a general 4x4 matrix, it fails compilation because you have to convert the point to homogeneous coordinates first. Very useful information from the type system.
That is, generalized matrix types, not simply rotation matrix types or whatnot for special cases.
There are a few shining exceptions, but not many.
One thing that I think could really improve the documentation is a few examples! One of my favorite things about the Python docs and the community is the wealth of examples. From looking at the docs, I couldn't find the main thing I wanted to see - what would MoneyType's annotations look like if I used it?
Man, that's crazy. At the time they were acquired by Facebook, they had 13 employees.
Or that's total number of engineers and way fewer actually twiddle the Python...
Good question about abstract base classes! Paraphrasing a well known cliché: types in functions should be forgiving in arguments (what the function accepts) and strict in return values (what the function emits). In our case, the human reviewer needs to decide if the argument types collected by MonkeyType should be generalized. In fact, the collected types might not even work in all cases and the type checker might complain. It's because annotations describe "what should be" whereas MonkeyType finds "what is". This is why a system like MonkeyType shouldn't even attempt to use abstract base classes in place of concrete types that it collected.
Having so much dynamically typed code to maintain that you need to run production code using a separate tool just to figure out the types sounds just wrong. Why not use a statically typed language for such a large code-base? Is this done by purpose, or did they end up with a million lines of Python code and are looking for ways to make the maintenance easier?
And before I get down-voted to hell - I completely understand using Python for many things. It a good technical choice for many different problems, but navigating a million lines of Python seems just daunting to me (although maybe I'm just not experienced enough with Python).
maybe it's easier to do this half-measure than rewrite your entire code base
They went with a dynamically typed language, were successful, the language they chose added an optional type hinting system, and they wrote a tool that would automatically type hint their code in order to reap many of the benefits of a static type system.
I think that the amount of man hours that went into writing the tool is negligible, so it's a net win for Instagram.
(Though really you should just use Scala and get both Java-like safety and Python-like productivity)
I think the reverse is true. Static typing is liberating for humans because it tames complexity. Because I'm not a machine I cannot possibly keep track of fuzzy programs that arise from dynamic typing.
I'm one of the dumb ones and just let the computer do the checking for me.
It doesn't, though. Not with the currently existing type systems and implementations.
- Without type inference you end up righting multi-tier type declarations everywhere.
- With overly powerful type systems you need something close to a PhD in math to create proper types and then figure them out half a year later when you've already forgotten most of what you did
- Union and intersection types which are extremely valuable are missing from a lot of statically typed languages
And because I'm not a machine I often cannot figure out what a yet another two-hundred multiline error message wants of me. Often I'm happy to just throw an `if (x && x.field){}` and be done with it.
If I look at some of the python code I've written, I'll be perfectly honest I often cannot tell you what to make of it any more.
Quite often people construct complex type hierarchies just because they can (or don’t know better). And it’s a pain to wade through and coerce to what you want it to do.
I’m very much on the fence between static and dynamic typing, having used (and probably abused) both. I prefer a “pragmatically” typed language, but I haven’t come up with a proper definition for it yet :)
This tool? This tool takes an existing set of codebases and makes them safer. No major boiling of oceans required.
From a technical aspect, I do find these projects cool. I wonder if its more efficient for large companies to initially develop using dynamic languages then transition them with these optionally typed languages.
We just recently moved our entire codebase from JS to TypeScript which was pretty hard but super worth it.
1. Startup builds thing fast in dynamic language because they need to optimize for development speed and iteration, not maintainability or scalability.
2. Startup grows and continues to hire for expertise in the tech stack they are mostly already using.
3. Repeat for some years and some hundreds of engineers and you arrive at this exact scenario.
People are running a business, not writing an a treatise on code maintenance and hygiene.
The software developers among us should also learn their lesson: don't build large-scale software in dynamic programming languages unless you can afford to spend time later adding a static type system on top.
Just because there may be better tools now, should they scrap their working code that earned their fortune?
In terms of what Instagram should do now, I'd say they should do what Facebook did: introduce thrift or similar, gradually move business logic into backend services written in more suitable languages, leaving the Python to eventually become just a thin web frontend. Retrofitting types onto code involves a lot of the same effort as rewriting it into a better language, and the rewards for the latter are higher, IME.
We still have to work with the world as it exists. Not as it should be or will be. And even with the crop of modern languages, its often still faster to start with less optimal languages and fix shit in the 1/10 chance you're actually successful.
Haskell remains impractical for many use cases, it is not used much outside of academia, it's not documented to be used outside of academia, and it didn't even have a working package manager until a few years ago.
Not sure why the negative vibes toward this comment, I think it's fairly sensible (of course, this is a very subjective subject).
> 3. Repeat for some years and some hundreds of engineers and you arrive at this exact scenario.
Then as the product gets bigger, you'll hire python developers to keep up with the workload - and the best ones will be the ones who have committed their lives to Python. So you'll now end up with more and more python code.
Before you know it, you aren't writing small scripts anymore, but now you are writing quite large features, that take weeks and that require intimate knowledge of the code base so you don't keep backtracking and repeating yourself. But Python doesn't help you at this point, you traded static types for flexibility and now you have to pay the price.
At this point you're screwed, too many man hours spent on the codebase to redo it, so what to do? Well if you have the man power, build your own static type checker of course! I mean after all, if you have 100s of engineers who cares? You just throw more people at the problem until it goes away. Then wrap it up in a nice little package, and slap yourself on the back while you ride the instagram bubble.
Did it occur to you that instagram published a valuable and useful tool that now just exists, and this is now a non-issue for anyone else in their situation?
Like, why are you complaining about resume boasting?! It's like you want people to do useless work that has no positive effect on the open source community. Are you just upset people are using Python or what?
Definitely the latter. I've seen this discussion a few times before, and it's always the same. Your initial developers are not looking down the road to the million lines of code milestone, they're just trying to make a product that might actually make some money here and now.
I'm sure Instagram was exactly that. They needed to handle images and some guy knew how to do it in Python. They wrote Python code, and then people liked Instagram. They eventually became a billion dollar company with millions of lines of code and no where along the road was there time to say "hey we need to refactor this whole thing". Or if that was said, management laughed and said "we need this feature".
So here is where you end up. The developers need to clean things up but they don't have time to clean it up by using a language, realistically, they probably don't know as well as the Python they wrote the millions of lines of code in.
Re: your last comment, navigating a million lines of any codebase is daunting, and especially more so if you aren't a developer in that language. I'm not sure what exactly "Python" has to do with that, besides that you're not a Python dev.
Well as he said, a statically typed language is better in that kind of situation because it enables a better class of tooling and the typing system enforces certain style constraints, that enables better quality of code analysis en mass.
Python specifically is very lightweight in this regards with little in the way of naming constraints (vs for instance Ruby having different formatting rules for different types)
So yeah, it’s not “just like any other language” - horses for courses
Stricter typing goes a long way to achieve this, and gradual typing allows you to upgrade the code base at your own pace, which is great.
Consider this study[0] about TypeScript and Flow, which use the same approach for JavaScript, which found both able to detect ~15% of runtime bugs. So no wonder companies with large Python code bases would be the first to invest in this space.
Personally I feel this is a great addition to the language, and hope type checking becomes a first class citizen too, instead of being delegated to external tools like mypy[1] or pytype[2].
You can statically type Python 2 codebases, but the language does not offer native support for it. Thus, all needs to go to docstrings or comments.
First you can't read what the types passed into and out of functions are. You have to find their usages to work it out. Second, you can't reliable do things like "find usages" or "go to definition" because of the dynamic typing.
But more often, when I'm looking at c# or c++, it's not code I wrote, it's not code I intend to change, it's code that's interacting with my code (written in another language) that I'm trying to see why it's misbehaving, so I can get the owner to fix it. I could be reading the code on GitHub or some other web view, I might have checked it out, but I have no interest in setting up a (probably new) IDE to look at it as the author would; I dig into too many projects to learn that many tools -- and deal with the upgrade cycle for them.
Sure, it would be useful to hover and get more information, but I'm used to loosely typed languages, so it's not awful. It's just jarring to see that the type information is apparently not important enough to write down the name in c++ or c# anymore.
In my experience PyCharm can do both correctly for the vast majority of cases.
def some_func(foo):
foo.run()
...
Find usages in the run() method will return dozens of results, the IDE can't help you any more, to find what 'foo' is at runtime.A hero arises, offering a sacred herb to calm the torrent and light the golden path. The hero is elevated, yet they continue to pray
They have password managers you know and also spell checkrs too. :P
Kevin Systrom " thought of combining location check-ins and popular social games. He made the prototype of what later became Burbn and pitched it to Baseline Ventures and Andreessen Horowitz at a party. He came up with the idea while on a vacation in Mexico when his girlfriend was unwilling to post her photos because they did not look good enough when taken by the iPhone 4 camera." (Wikipedia)
He used Django because I guess that was an easy way for one guy to do it fairly quickly. The app was Burbn which then pivoted into Instagram.
By the way I kind of surveyed the "what framework should I use" stuff on HN over the last year and Django still seems the most popular, probably followed by Rails and Phoenix.
Clearly something is right with the situation when the incentives are aligned for a tech company to contribute back to the open source community in such fundamental ways. So why look for the mole and think "They should have done it differently", when doing it differently has a high likelihood to mean not being as successful as they are today, and not having the occasion to contribute back?
It's like telling a successful charity "You should just take everyone's money and spend it on lamborghinis instead of wasting time building wells in africa".
Python is an excellent enabler of this kind of dynamic system evolution.
Python allowed them to build a successful company. Now, when their stack is mature and maintenance is more important than rapid prototyping, Python allows them to add type hinting.
Because they are engineers, they built a tool (in Python) that allows them to do it in an automated manner.
And all of this is great!
They are evolving their code to fit their needs; it's nothing like making a octagonal wheel and wishing you'd have gone for a round one in the beginning.
The cost of doing it right from the start is negligible.
Big assumption.
Both because they might have known Python already and also because Python is quite a bit more newbie-friendly, concise and expressive than the mainstream statically typed languages.
But no, ignoring that it is not a big assumption really. The benefits of static typing comes pretty quickly, especially if there are more than one programmer.
Type hints allow external tools to check some things, but at this point you're basically imposing static types so why not use a language with the tooling and optimizations to take advantage of that?
Python is a good choice to prototype, write small (less than a few thousand lines of code) projects with non-trivial complexity, and somewhat larger projects with more boilerplate (e.g. Django webapps). Beyond that its utility diminishes until it starts to become a hindrance.
Python's type system is, imo, currently better than Java's, and the syntax is cleaner than java's or C++s. You get all the benefits of static typing without having to put `auto` and `List<>` everywhere. And at the same time, you get all of the advantages that python has over statically typed languages that aren't haskell (like comprehensions). And, when you need to, if you're doing something that's especially tricky or dynamic or whatnot, you can fall back to untypedness.
I think the closest parallel I can draw is to something like Rust. You get a huge set of guarantees for free, but can opt to do unsafe things when it's absolutely necessary, and better yet, you can start in unsafe land and then go back later and make sure your code is safe.
I'm curious what tooling you feel that say, Java, has over type-annotated python.
This is also possible in traditionally statically typed languages. Nothing stops you from doing unsafe casts or using reflection. Much like its exceedingly unlikely that you'll run across this in "normal" java or C++, its exceedingly unlikely for you to run into any issues with this in python. And, in fact, the typechecker has ways to handle unusual things like dynamically created attributes, for when that comes up.
And yes I mean this quite honestly. I've seen a lot of typechecked code, some of it quite ridiculously dynamic. Typecheckers perform absolutely fine.
>The "tooling" the other languages have includes a compiler that performs these checks in a way that Hints + Checker-of-choice is unlikely (or unable) to.
What way is that? Typechecking is static analysis. There's really no difference between how java or cpp does typechecking and how mypy does, other than that the python typechecker isn't installed by default.
>The answer is: not nearly what a compiler does.
This is not an answer.
Neither of those is "silent" or "unknowing".
> What way is that? Typechecking is static analysis. There's really no difference between how java or cpp does typechecking and how mypy does, other than that the python typechecker isn't installed by default.
The typechecker can't handle un-hinted code (or, rather, it chooses something very permissive, like 'Any' for all hints). It's incomplete at best.
> This is not an answer. It is. That you don't like or agree with it doesn't make it not an answer.
And in Java or c++, un-hinted code couldn't compile. The python type checker can do more than a java or c++ checker in this regard.
>Neither of those is "silent" or "unknowing".
They're exactly as silent or unknowing as you would get in typed python code. You appear to be comparing untyped python. That's an incorrect comparison. Offhand, I actually can't think of anything I could do in typed python that would get around the type checker, that wouldn't be considered reflection or a dynamic cast, and be very obviously so in python too. If you have an example of a silent or unknowing failure of well typed python code that passes on mypy, you should probably file a bug report ;)
>That you don't like or agree with it doesn't make it not an answer.
You're right. Its not an answer not because I disagree with it (I don't), but because it doesn't actually answer anything, which is why I don't disagree with it.
To summarize this:
Python typecheckers are capable of more type inference than Java, and require less syntax than c++ or Java to get well typed code. A typed python codebase can interact cleanly with an untyped python codebase, and within the typed parts of the code, you get equivalent safety guarantees to what the type systems of Java or C++ provide.
Your appear to be ascribing magical powers to compilers in other languages, when those compilers have exactly the same type information as mypy does.
In other words, going back to your first statement:
>Hints don't provide any guarantees.
Hints provide exactly the same guarantees as any other type system: "Assuming you write reasonable code that doesn't attempt to subvert the type system, the type system will catch any dumb mistakes you make."
That's the exact same guarantee you get in any statically typed language.
But you're right, I should have specified "traditional" static language.
There is no reason to to save the literally 0.2s (you can still spend that time reasoning about your code) it takes to write the type. It is better for yourself writing it and for readability to be explicit.
auto s = "Hello world";
for (auto c: s) {
cout << c;
}
Those autos don't need to exist, they're completely inferable, otherwise you wouldn't use auto. It's not like you can use auto in function declarations, nor should you, I agree.Tell that to any serious Numpy/Scipy/Pandas user.
The empirical evidence of reality is against you: there are successful large (in terms both of codebase and contributors/development team) projects in these awful terrible children's languages, and there are unmaintainable failed piles of crap in even the most grown-up of languages you'd care to name. The choice of language, and choice of type system, seem not to correlate with the success or failure in a meaningful way.
The interesting thing to note is that a language that's perfectly acceptable at the above at small or medium scale might turn into a hindrance at large-scale. An otherwise fast to develop in language like Python won't be so fast if every change has to be painstakingly reviewed and tested due to the complexity of interactions in the code base.
Using a type system to verify assumptions/requirements is not a recipe for success, but it can improve reliability.
You can write spaghetti code in any language, it turns out. Blaming the language for that is not really an indicator of understanding the problem.
Here we're talking about average or best-effort: large code bases are complex in spite of the best intentions of their maintainers, so using tools that can manage that conplexity in an easier way through e.g their type systems could lead to better results.
The preconcept that dynamic languages are more productive is just an illusion because you can easily take shortcuts that will hamper your progress in the future. A proper typed language with HM type inference has the ability to mostly avoid writing the types with the guarantee that the compiler will catch most of the pitfalls. And if you don't do any logic error pretty much every time your code just works. Saying that in a million line application Python is a better choice than F# or Haskell it's frankly ridiculous in my opinion.
Do you disagree?
People do use ML family languages, and they are better. There are plenty of non-technical reasons they aren't as widespread as dynamic languages or shitty static languages.
I'm a big fan of rapid prototyping with Python to map out problem domains, and once the domain has been mapped properly, rewriting in a statically typed language if necessary.
Python is much better for prototyping than the other langs I've used. Because the syntax is almost pseudocode, and the duck-typing makes a lot of design patterns and boilerplate obsolete, so I can dedicate my headspace to the problem at hand.
Right tool for the right job, as they say.
Honestly, your octagonal wheel metaphor works, too. Building the first car you spend a lot of time on octagonal (crude) wheels, but later spend a lot more money on round (precision) wheels. You could have gone bankrupt spending money originally on round wheels that were the wrong size.
[1] https://www.joelonsoftware.com/2000/04/06/things-you-should-...
You can churn out a greenfield project faster as you don't have to spend upfront time mapping out interfaces, DTO's etc. It's also easier to 'hack'.
Of course the above makes the code harder to maintain and reason with, unless written by very disciplined engineers.
So these languages are a natural fit for startups (and things like prototyping and scripting).
Instagram would have been in this category, and it's worked out well for them.
Being in a similar (though much smaller scale) situation ourselves, I suspect your latter suggestion is the answer; they wound up with a large amount of Python code due to expediency (Python is good at quickly getting things done), and are now finding the code base quite hard to maintain.
Something to remember is that as of 10 years ago, statically typed languages consisted of C (for some definition of "statically typed"), C++ (complicated and cumbersome, and didn't yet have widespread availability of "modern" C++ features), Java (cumbersome, heavyweight, lots of missing opportunities for abstraction back then), C# (heavily tied into the Microsoft ecosystem, not yet open source), and a bunch of weird academic languages that required tutorials about burritos to learn (ML, Haskell, etc).
While statically typed languages like C, C++, and Java were the mainstream languages of the 90s, C was too limited for a lot of people, C++ and Java too cumbersome to use, and so lightweight dynamically typed languages like Perl, Python, PHP, Ruby, and JavaScript picked up a lot of steam due to how much easier to pick up and more productive many programmers found themselves in those languages. But now we have large, fairly un-maintainable code bases in these languages, and people are realizing the value of static typing, in part due to the maintenance hassle and in part due to newer, more expressive and accessible statically typed languages being available (Elm, Rust, TypeScript, Scala, Go, Kotlin, as well as improvements to C++, Java, and C#).
But that leaves all of these old codebases, that are hard to maintain. Rather than doing a complete rewrite, adding static typing capabilities that can be applied to existing codebases is a way to make them more maintainable without spending all of the time of a complete rewrite and having everyone have to spend all of the time learning the new language while still maintaining the old codebase.
The worst thing -- to me -- is dynamic typing.
So why not fix it?
It's not that bad.
The first thing you learn when you're in a million lines codebase is that you will only work within your project of maybe a hundred files.
Once in a while, there is a guy who is asking for help on his project or there is an old weird bug to fix and you dive in other stuff. Otherwise, it's like it's not there.
Currently, I think the nearest competitor is node with typescript, and I'd rather stick with python. Please tell me if I'm wrong.
I also think that if they want to use types the correct approach is to apply a tool like this as a stop-gap but write new components going forward and bug fixes/feature re-writes in a language that supports types "properly" (i.e. in the way they seem to want, that is static types checked at compile time).
I think tools like this are great for companies in situations like this. I don't think they're good to use from the outset: the team should just use an actually statically typed language instead.
Though type inference has been around in so-called "academic" languages for decades, it hit a tipping point in the last 10 years, to the point that every major dynamic language has static type checkers or dialects (like Typescript) that support static checking. Meanwhile, even traditional statically typed languages are growing stronger type checkers.
Python's type annotations are in fact very similar to Flow and Hack in the sense that they provide gradual typing. The specification (see: PEP 484) describes that only annotated functions are type checked. Calls to non-annotated code are treated as accepting any type in arguments and returning the Any type (a special type which effectively silences the type checker).
This generates a chicken and egg problem: if you don't have enough functions annotated, the type checker won't be able to provide meaningful output to you. So convincing people to annotate their code is harder: they don't see the benefit right away. Worse yet, you already have tens of thousands of functions in your code that you know work in production but were written before type annotations were introduced. It's not really feasible to come back and fill this information manually.
MonkeyType is a tool that gathers types at runtime and enables putting them back in your code as annotations. The goal is for mypy to have more information to work with, making it way more useful.
[1] https://medium.com/fhinkel/runtime-type-information-for-java...
I too thought this but even just one check in a chain has been helpful to me.
That said, I am not saying this tool is bad. It could very well help a lot of codebases, but I would warn against using it as part of the operational workflow.
Do you mean cases where you accept "anything iterable", an issue with callbacks / template types, or something else?
Can you give examples? In practice, I haven't noticed this being an issue.
Can you give examples? In practice, I haven't noticed this being an issue. A lot of python code already had/has docstrings explaining the types, these annotations just formalize that a bit more.
Are most of their code not yet in production ? Or why do they produce more bugs now with static types, then before with dynamic types !? This sounds a lot like homeopathy, that can both detect and cure diseases with placebo.
It always amazes me that some of the most popular products around are built with the worst technology choices. And now they had to build their own static type checker, which slows down random samples of real users, just to shore up the language's weaknesses? Outstanding.
(Not totally a joke. My mother wrote in octal before she got an assembler.)
I'm not commenting on the specific issue of whether Python creates good trade-offs, just that I think tallying successful companies is a poor measure of a programming language.
(If this goes well maybe you'll get an email from Mike Krieger with the subject line "Here, you do this".)
I think it could be pure shit -- and indeed a very simple idea. They got big and famous, but it wasn't for dexterity in language choice or coding.
I don't know. I suspect being half-pregnant is a far bigger issue than the technical debt of dynamic typing, whether it stems from being slower to implement new features, spending time choosing the "correct" tech stacks, or spending more time/money hiring qualified engineers.
There's something to be said for the fact that many of the most successful startups in the world are saturated partly or fully with Python and PHP: getting features and new hires off the ground in as little time and money possible is worth its weight in gold.
I've always been very much of the opinion that the smartest way to go is making good use of FFI where it counts, and doing everything else in whatever gets code out the door most effectively. This TypeScript/Hack/MonkeyType trend just seems like a very welcome extension of that.