Selecting a programming language can be a form of premature optimization
snarky.ca
snarky.ca
This isn’t a knock on rust - I just think rust trades off programmer productivity for correctness and performance. And it shows, on both sides.
I recall a study being quoted in my university courses, which described the amount of code needed to get certain things done between different languages, which showed that Python, Ruby and others are on the less verbose side, the reasoning being that on average they'd also be more productive.
This seems to coincide with my personal experience, where JavaScript with React was much easier to work with in smaller projects, whereas using TypeScript with React lead to much slower development because of all the typing that needed to be handled, especially in cases of union types. Now, one can say that it's worth the effort to be more confident in your code doing what you want it to both now and after X months, much like you could sometimes prefer the type systems of Java or .NET over Python, Ruby or PHP, but in my eyes those tradeoffs always come with slower development velocity.
The sad thing, however, is that DuckDuckGo (and possibly other search engines) failed to return the original study or anything like it, only resulting in low quality blog content:
- https://duckduckgo.com/?t=ffab&q=programming+language+comparison+amount+of+code
- https://duckduckgo.com/?q=programming+language+comparison+verbosity
- https://duckduckgo.com/?q=programming+language+lines+of+code+java+ruby
- http://libgen.is/scimag/?q=programming+language+verbosity
- http://libgen.is/scimag/?q=programming+language+lines+of+code
Anyone have any better search queries for this? Any idea why the search result quality is generally so low? Any ideas which study it might have been referencing?The first result for "program language comparison", at least when I view https://scholar.google.com/scholar?q=programming+language+co... is: Prechelt, Lutz. "An empirical comparison of seven programming languages." Computer 33.10 (2000): 23-29.
There's all sorts of work which cite that paper, like "An empirical investigation of the effects of type systems and code completion on api usability using typescript and javascript in ms visual studio".
For fun, here are some of those citation which themselves use the phrase "An empirical study" or similar in the title:
"An empirical study on the impact of static typing on software maintainability"
"An empirical study of the influence of static type systems on the usability of undocumented software"
"Do developers benefit from generic types? An empirical comparison of generic and raw types in Java"
"An empirical study on C++ concurrency constructs"
"An empirical study on the factors affecting software development productivity"
"An empirical study to revisit productivity across different programming languages"
This seems like a pretty apples-to-oranges comparison. Rust is intended to do things you couldn't/wouldn't use JavaScript for.
If you're talking about a small-to-medium project where you don't care much about mishandling memory, of course JavaScript is going to be more productive. Rust is forcing you to tell the compiler a lot of things that JavaScript assumes you don't care about (and is, most of the time, correct).
I think it was 2 things: one, it is genuinely harder to write a valid C++ program. The language imposes non-trivial constraints on the programmer, for better and worse, and you have to navigate them for the thing to even compile. But two, C++ invites pedantry. I heard endless discussions about "ooh, template meta-programming", "struct or class", "but is this idempotent", "you should be using move semantics", "why on earth is this a unique_ptr".
I'm not saying C++ is worse, or that its trade-offs are not worth it, but for sure, from my experience, translating an algorithm into code is done far more quickly in Python than in C++. YMMV
Ah and also, this was definitely not a question of deficient C++ coders. The hiring standards for C++ were so high we were always short of devs. Meanwhile, Python prototypes were usually written by part-time ex-academia Python dabblers.
-PEP 20 -- The Zen of Python
—The next line
This ZoP line is by far the most misused cliché of all things Python. How is it determine that “prototyping in C++ and rewriting in C++” is not the obvious “one” way? Because you don’t like it? A Zen is intentionally self contradicting to allow introspection, not judge others :)
-- Django
-- Perl
-- assembler
-- fpga
Back at you.
-hdl
This seems like a monumental level of operational waste. I'd love to hear more about this particular setup
I still agree that Python can be much more efficient to write than C++. I just wish it wasn't so unbearably slow.
Even then though, I have the impression that Python just makes it easier to write it all more quickly. It just doesn't invite you (as much) down rabbit holes of doing things even more properly.
You would initially trade some of the C++ performance, but the code shouldn't be more complicated or difficult to write or compile than Python. Then, you'd only need to change the performance sensitive parts to use real C++ arrays or structs.
I wonder what the result would have been in eg. Go or Rust, where that level of conversation doesn't exist. True, others would have taken place instead, but at a higher level nearer the design, where they should have been happening anyway.
I've heard reports like this for many years.
My pet theory is that the discrepancy here is ~10% static types and ~90% memory management. Having a GC (or refcounting or whatever) fully baked into the language so that you don't have to think at all about how values are passed around and stored is a monumental productivity boost. Possibly the biggest win in software engineering productivity in the history of the field.
Interestingly, I've consistently had the opposite experience. In C++, when I ask, "how do I do X" there will be a few different options, but pretty much any of them will do. As long as you avoid things that are definitely UB, you're fine. In Python, I'll find endless religious arguments about which is the most "Pythonic" method. Whenever I try to just hack out Python code to just get things done, I always feel like I'm being judged, because usually the quick-and-dirty method is far from the most Pythonic.
Ruby on the other hand embraces many different syntaxes to do the same thing, but they’re all accepted as correct in the traditional view.
I consider myself as someone who knows "far too much about C++", but recently when I had to spider some webpages and read some values out of them to populate a database, I did that in Python because it's a handful of lines due to high quality libraries that all work nicely together. I honestly wouldn't know how to do that in a short amount of time in C++.
Python and other dynamic languages don't use static types and that may give them a (very) small advantage in productivity as well, but only for very small programs... as program size increases, in my experience, statically typed languages take the lead in productivity... as OP is about writing general software, not just small scripts, I really have a hard time agreeing that Python is either the best tool for the job, or the most productive tool at all.
#include <string>
#include <cctype>
#include <algorithm>
#include <iostream>
int main()
{
std::string s("hello");
std::transform(s.begin(), s.end(), s.begin(),
[](unsigned char c) { return std::toupper(c); });
std::cout << s;
}
Same thing in Python: print('hello'.upper())In fact it's really mostly name space management. Python doesn't require you to pull in "upper" because it starts with a relatively big name space. It doesn't require you declare the main program because you're relying on the fact that Python executes a file on loading it (and that reliance is arguably bad Python style). All of the "std::" stuff is a name space management choice.
... and the rest has nothing to do with static typing either. The way the code is indented on multiple lines is a stylistic convention. The transform primitive requiring you to specify bounds is a library choice, related to the language's apparent lack of a universal way to map over a container. I'm not sure why you need the lambda. The rewrite in place is memory model stuff inherited from C.
In Haskell, which is fully compiled and is about the most rigid static typed language you could ever imagine:
import Data.Char(toUpper)
main = putStrLn (toUpper <$> "hello")
If you'd asked for something that didn't require oddball functionality, then that could probably also have been a one-liner. std::string ToUpper(const std::string& str) {
std::string as_lower(str);
absl::c_transform(str, as_lower, std::toupper);
return as_lower;
}
to enable: int main() {
std::cout << ToUpper("hello");
return 0;
}
but this would be no good! You have chosen an API that forces your users to be gratuitously inefficient. Maybe that is OK for your personal project, but an API like that has no business in the standard library.As a counter point to this - Python has a large number of C-based libraries with excellent bindings. To take Numpy has an example, you can the static-types and speed of Numpy to do all your maths, whilst Python acts as the manager. Using Numpy "feels" like normal Python, you don't really feel like you're working at native-level, but you get all the advantages of native speeds and memory usage.
And to your other point, I've worked on large C++-based projects which were a complete mess, especially when you had to make any changes to existing code. I'm not sure the language has as much an affect here as just using the correct design principles.
Recent-ish commentary (July 2021) about it at https://renato.athaydes.com/posts/revisiting-prechelt-paper-... , HN commentary at https://news.ycombinator.com/item?id=28108806 .
Google Scholar gives about 476 paper which cite Prechelt's work. I have not followed other work in that field.
Just saw that, nice. I like Norvig because he really cuts to the chase.
It's not just that they are more popular, Java reigned supreme for years when Ruby and PHP and later Node.js outpaced it's communities by miles. And you can't say Java developers are not open source oriented, because they 100% are.
And the argument that programmers are simply most productive in the language they know best is trivially untrue. When I first learned of Ruby, I could dream in C#, I was absolutely fluent, knew the standard library by heart. I switched to Ruby and only looked back whenever I needed tight performance. And this is not just an anecdote, the whole Ruby on Rails movement was basically Java developers fleeing to greener pastures.
I don't know a single highly experienced multi-lingual developer who does not reach for a dynamic language when they need to deliver something quick and easy.
The only exception I know off is using Go for small networked services, which is quite comfortable and intuitive for a static language. Also outside that niche Go quickly loses its productivity edge.
If you're picking from a popular language, the language itself rarely matters for this at all. Every popular language does roughly the same things. Writing speed (length of keywords, for example) is a non-issue. The ecosystem is by far the most important thing.
It doesn't matter how efficient I am in Go if I'm missing a crucial library that I'll now have to write myself.
JavaScript is a classic example where, even if you are pretty fast at writing your code, you will be dramatically slowed down by: 1) needing to add a new package every 5 min to do something trivial, 2) looking through five half-dead libraries to find one that seems maintained and usable, and 3) finding out that some of the libraries you chose are buggy.
So for a quick prototype, I'd weight a language's qualities as follows:
- ecosystem: 90%
- static analysis/tooling: 5%
- stdlib: 3%
- syntax: 2%
And for a long-term, complex project, I'd weight them closer to:
- ecosystem: 50%
- static analysis/tooling: 46%
- stdlib: 3%
- syntax: 1%
This argument is not addressing the claim for general purpose software, only for "quick and easy" software.
To people in this thread: please stop to think before responding to the "wrong point". I think we can all agree that dynamic, scripting languages are more adequate for "quick and easy" software, but this is not what the point of the discussion is. The point of the discussion is, or should be in my opinion, whether it's true that for general purpose software, written by teams, that is not trivial to write by oneself in 5 minutes, that scripting languages are more productive than statically typed, compiled languages.
I think you misunderstand the scripting language revolution of the mid 2000's. It's not that we suddenly realized scripting languages were the best for quick and easy projects. We realized scripting languages were suitable for a whole lot more than some data processing. We could much more effectively build huge scalable software platforms.
It got so bad that the sentiment flipped, and people like Joel Spolsky had to go out of their way writing blog posts that you could in fact build successful modern web platforms in C#. And then he went building the world's most popular project management tool in Node.js anyway.
Maybe this shines light on your social circles more than actual language impact on proficiency?
The "our language is so much more productive" myth exists in many languages, including the one I currently use. But when you dig deeper, there are almost always other social or technical factors that explain the productivity gap. That kind of gap exists even between people using the same tool.
The reality of it is that assessing what makes a developer productive is incredibly hard, and people doing so to claim their language of choice is better rely on anecdata and couldn't explain what "methodology" means to save their life.
Hi, highly experienced, multi-lingual developer here (Python, Ruby, PHP, JavaScript, Java, Go, C#, F#, OCaml, Elixir/Erlang, ReasonML, SQL, Smalltalk, Objective-C, Swift, leaving off quite a few prior Web 2.0 ones for brevity) nice to meet you. As you can see, I've done just about all of it: every paradigm, every syntax. It all honestly blurs together after a certain point, and it just becomes easier and easier to pick up a new language the more you learn.
Anyways, these days, I never reach for a dynamic language when I need to deliver something quick and easy. The difference is just too marginal. Not worth the downsides (and the downsides are immense!).
Because, as my experience has taught me, invariably one of those quick and easy ones will turn into something that becomes business critical, lives on for years after you are gone, and will be much more difficult for other, typically more junior, developers to update or enhance your code. Static typing is not only marginally less productive these days with all the great tools and IDE's out there (not to mention Go which has one of the least obtrusive static type checkers I've seen, or, even better if the org allows you, a language with a HM type system), but for a marginal improvement in productivity, you pay a heavy long term cost. Not worth it.
Dynamic languages are great for rapid prototypes. After that, convert it to a static language. Your junior devs that join after you, who aren't familiar with the entire ecosystem your work has will thank you.
I don't think you can have a hard and fast rule like this. It will depend on the situation. In many instances, a "prototype" built in Python or similar is perfectly fine.
In addition, one can make a horrendous mess in statically typed languages as well. One of the absolute worst projects I ever worked on was written in a popular statically typed language. This is of course an anecdote and not to say statically typed languages are worse, but they're not automatically better either.
Just a note too that I'm a big fan of, e.g., Rust & TypeScript, and I often use type hints in Python, so I'm not anti-static typing.
There is a floor for how bad you can write static typed code.
There is no floor to how bad of JavaScript or Python or Ruby you can write. The madness can descend to the inner most circles of coding hell. Only Perl exceeds it.
Explained:
Go's type system is generally less rigid. It strikes a good balance of strict enough. A lot of Go's converts aren't from "systems" languages like it targeted originally, but rather former Python/Ruby/PHP/JavaScript backend devs. I love the performance and low level levers I can pull with Go (although to be fair Java is quite fast enough). But finally, Go is easy to learn (26 language keywords?), the standard library is mostly great, and the worst developers I've seen still write mostly maintainable code that builds fast, which is what I optimize the most for these days.
Java, for all its warts and legacy cruft, these days you can write fairly good java, utilizing modern libraries. I love most things from Codahale, who in turn I think pushed orgs like Spring to write better libraries, so now everything's pretty good. Plus all the legacy stuff comes in handy when you have to deal with arcane government or financial systems, something I have to interface with frequently.
But if I had my choice, I'd use something where you can express functional programming concepts intuitively, without fighting the language or having to do it at a heavy performance cost, like Rust or better yet, just a full fledged FP language like OCaml (whom I understand heavily inspired Rust)
When I want to develop further, I put a debugger statement right where I want to pick up, where all the data is available, and develop from execution.
Some static languages have REPLs but few are as good as dynamic languages for mutating existing code in the middle of a debugging session.
If you don't write much code which involves exploration of data or unknown APIs, then this mode of development may not be as useful to you, but it's a significant productivity advantage for me and the reason I reach for a dynamic language. Ruby is my go-to, with binding.pry as my debugger REPL.
Popularity and expansive use may have nothing to do with quality, and a lot to do with financial influence on peoples agency.
Business wants people templating out directories of performant code, not generating syntactic art for the ages.
“Strings and string formatting” is a very broad domain.
For most specific tasks, there is one low-impedance approach and it takes very little (but not zero) reflection and/or experience to find it.
Most new Python features directly address specific tasks for which there are currently multiple relatively high-impedance approaches taken because there is not one obviously correct way.
Python's “one obviously correct way” is not “one possible way” (the latter approach is closer to Go.)
"People who know one language think it's the greatest in the world. People who know more than one think they all suck."
The Blub paradox: http://www.paulgraham.com/avg.html
(BTW, the TL;DR is people that know more than one language admit they all fall short one way or another, but there are some languages that are more productive than others).
This is true, but most software is built by multiple people (for corporations) and will eventually be maintained by multiple other people. Whether an individual is "most productive" in a language is rarely important.
I’m much more productive in Python. Not as in “I feel” more productive, but as in measurably more productive.
Where Python falls short of something like .Net and C# is that I probably wouldn’t have been more productive in Python if I didn’t have 7 years experience with a rather strict environment, and as far as using Python on major projects, well let’s just say that there is a reason people use TypeScript instead of JavaScript and Python doesn’t fall into that category.
But for most programming and for most minor systems or services, Python is just wildly good.
The second point matches my experience with dynamic languages as well. To use them well really takes a certain level of discipline, but it pays off. It's why I'm a bit skeptical of Python as a beginners language, which it sometimes has been touted as.
About large projects, presently I am working with Python on a somewhat large project (currently 53k sloc, I guess large is relative). Its been great, though we are pretty strict about almost absolutely everything being type-hinted.
Maybe not the best example, but I used to work with a guy who had pretty large scripts that he was writing for one of our projects. It turns out it was mostly copy-paste, so it sort of got the job done -- except he fixed bugs in one place but not in the other. The whole thing was a huge mess to understand and maintain. And he was a senior engineer at the time (now principal, to my great amazement).
Why are they writing the same code again instead of writing different code? Why aren't they learning about new APIs for topics they aren't so experienced in? What makes them "senior developers" and not "senior transcriptionists"?
Who is tasked with digging into 10 year old code, written by an ex-employee, to find a subtle bug? Who is identifying and fixing performance issues?
How are these "senior developers" able to write 1KLoC/day plus documentation and developer tests? I find test code and documentation each take as much time as writing the code in the first place.
Why aren't some of these senior developers producing negative LoC? https://www.folklore.org/StoryView.py?story=Negative_2000_Li...
For example, 1M+ LoC of code, including build system, CI, test cases, built-in documentation, in 6 months is about 7kLoC per day, divided by 5 developers it's about 1400LoC per day per developer.
I completed 30+ projects already, I have 20+ years of experience. Few years ago I was able to close up to 20 tickets per two week sprint, when I worked with same level senior developers and dedicated PM, product owner, and QA team, at large outsource company at fixed-price projects.
It's easy when tickets are properly sized and described by PM, when PO is responsive and easy to reach, and when QA covers your back for complex test cases. 6 month project + 1-2 months to recover after, and then another such project.
Those numbers are ridiculous, and I say this having 25+ years of professional experience.
The typical industry numbers are under 100 LoC/day.
Eg, slide 20 of https://www.slideshare.net/ddskier/calculating-the-cost-of-m... ("A world-class developer (e.g. Facebook or Google senior engineer) will write 50 LOC per day")
"Improving Speed and Productivity of Software Development: A Global Survey of Software Developers" at https://uweb.engr.arizona.edu/~ece473/readings/9-Improving%2... (Fig 6) has Lines-of-Code per Total Man Months at about 1,750, so about 81 lines per day (assuming 21.62 work days per month).
"A Practical Approach to Software Metrics" at https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=819938&... says "Many industry rules of thumb describe programmer productivity in terms of lines of code, for example 350 [noncomment source statements] per engineering-month of effort." That's 16 lines per days.
Going the other way, in the COCOMO estimator at https://strs.grc.nasa.gov/repository/forms/cocomo-calculatio... your 1 MLoC for an "organic" project estimates 3390 person months. Not the 30 months you mentioned.
That said, it's really easy to distort LoC measurement. And I mean beyond the "put in 1,000 lines of the form 'a=1'" cheat.
For example, LoC traditionally refers to source code, not test, CI, etc. that you use. Why did you use that non-standard definition?
Test code, for example, may contain a lot of data records and autogenerated code. This doesn't require the same work as a source line of code.
Consider SQLite. https://sqlite.org/testing.html says the library is 143.4 KSLOC, while the test suite is 91911.0 KSLOC - nearly 100 million lines of test code! These tests were not all written by hand. At that ratio of test code to source code you would have about 1,400 lines of source code.
And it's possible to mis-measure source code. A few years ago I added about 400,000 lines of code to my project in a few weeks. These were auto-generated when I replaced a very confusing set of C preprocessor directives with a homebrew template system to compile specialized versions of functions across my parameter space.
I then added some Cython projects, which generates C files about 50x larger than the original pyx files.
Counting auto-generated code makes LoC a worthless measure.
You also wrote you have a QA team. I assume their test cases are in your repo, and counted in your 1M+ LoC count, but you didn't include their time.
The post-release defect rate per KLoC is about 7.47 (average) and 4.3 (median). See An Overview of Software Defect Density: A Scoping Study at https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=6462687... .
If you write 1KLoC/day then you're likely generating 2 post-release defect bugs per day. That's in addition to bugs your QA team catches.
Your "up to 20 tickets per two week sprint" simply isn't enough to catch up with the number of bugs you likely generate.
Your numbers are so far outside of industry standards and academic findings that, with 20+ years of experience, you must know they are exceptional and difficult for anyone else to accept on your simple say-so.
> Are your fixed-price estimates based on LoC?!
I don't now. I was not a part of sales.
> Else, why aren't you embedding that knowledge in a corporate library ("contractors tools") which you license to all your future customers?
Because each customer has different requirements, different goals, different language and framework. Nobody wanted to pay for a library.
> If you write 1KLoC/day then you're likely generating 2 post-release defect bugs per day. That's in addition to bugs your QA team catches.
In practice, I had about 1 serious bug slipped per 1-2 weeks. Professional QA team known a lot of corner cases to test already. Professional developers know them too, even when writing in a different language using different framework. Moreover, we know how to prepare good CI, efficient automated tests, how to write efficient documentation, who is responsible for what, etc.
> Your numbers are so far outside of industry standards and academic findings that, with 20+ years of experience, you must know they are exceptional and difficult for anyone else to accept on your simple say-so.
How much I will be paid if you will accept my story? :-)
> Because each customer has different requirements, different goals, different language and framework
Sure, but that doesn't match your earlier statement that the senior developers 'wrote almost same code dozen times already'.
That is, I don't see how "almost the same code" is able to handle different requirements, etc.
How much would you pay me to accept your story? ;)
> there is a reason people use TypeScript instead of JavaScript and Python doesn’t fall into that category.
I thought the dynamic vs static typing debate was pointless until I joined a large Python project. Now I pretty much consider anyone that can keep up with one superhuman.
Does that mean Python has all the performance you would ever need? No - we know it doesn't, but we also know that you rarely need that kind of performance and when you do it's usually localized to a very narrow portion of your code. Take that portion of code and implement in C and optimize to your heart's content. Even with that you'll still get your project done much faster.
Python is the tool of choice for those developers who just want to get stuff done.
I’m primarily a python user, but I spent time writing code in Ocaml to understand the potential benefits, but… (I would love to have my mind changed)… it feels like many of the features people touted about Ocaml have already made their way into python…?
1. Immutability / referential transparency seem more like nice-to-haves for a codebase, rather than the real reason people tout FP…?
2. Sum/product types and pattern matching are being added soon to python
3. Mypy is starting to gradually enable python typing, although I assume it’s still a work in progress
4. Python allows us to use map/reduce, and to pass functions around as arguments…?
I want to understand the potential benefits of FP more, but my experience with Ocaml hasn’t shown me great improvements yet. Open to having my mind changed.
In python, most of these "features" are done half right and duct-taped on at the last second as an afterthought.
Comparing Python to Java, C#, Haskell, Erlang, Go, maybe even C or Rust, would show a much smaller performance difference.
Have you tried RAII?
Would you like to know more? https://stackoverflow.com/questions/2321511/what-is-meant-by...
Rust is a little better since at least the ownership concept is known to the compiler and automatically enforced.
(Tracing/Copying/Compacting) GC is much easier since this entire concept goes away. You always pass references to data, and the physical storage is "owned" by the GC itself. It's also much faster for certain workflow patterns, though it always consumes more memory than deterministic destruction schemes.
Not if you don't go looking for it: http://www.norvig.com/java-lisp.html
ETA: And don't get me started on Design Patterns: https://norvig.com/design-patterns/design-patterns.pdf
We generally do everything in Java. Python is avoided because of all the usual complaints - dynamically typed, not fast enough, GIL, etc.
So, as a PoC, I decided to try rewriting one of our services in Python. And, compared to the Java one, it is:
Cheaper to write and maintain. About 1/10 as many SLOC. About 1/20 as many person-hours.
Has better static type checks. For example, Java's type checker cannot statically verify for null safety. The Python one I chose does do that. Note that, since I did choose to put type hints on everything, the productivity boost in question cannot be attributed to dynamic typing.
(That said, not all 3rd-party libraries have type hints, so, if you want to type check everything, you may have some yak shaving to do.)
Uses less RAM. About 1/2 as much.
Is faster. Admittedly I'm leaning on numpy, Cython, and friends for this. It may well be much slower for projects where that is not possible. But still, I think that the point about premature operation stands in this case.
(Disclaimer: I also have a colleague at $FAMOUS_COMPANY who tells me they are moving off of Python because they found the opposite of the above in many cases. Though I personally suspect that a mitigating factor is that their profit margins and scale are large enough that all the coefficients in their cost/benefit formula are wildly different from the norm.)
Check out NullAway.
> Uses less RAM. About 1/2 as much.
Were you on a modern JVM?
If you find yourself commonly having to cast to/from Object in Java, that is almost certainly something of your own doing.
I don't write Java, though, just guessing.
If you add a case, for example, you're 100% on your own to make sure that all code that interacts with the type is updated to handle the new case. Slip-ups will produce run-time errors rather than static ones.
It's also rather tricky (and awkward) to set things up such that consumers are forced to check the case value before attempting to coerce it.
All of this can be worked around with custom linter rules, but then that becomes its own maintenance burden.
My preference is to try and invert things such that you can rely on dynamic dispatch and "tell, don't ask." That eliminates the need to coerce things at run-time in the first place.
My gut would be a visitor for the fully general case. Though, I would also accept that you are using the wrong language for the abstractions your are writing? My guess is it will depend on why you are needing/wanting one of those types.
For a lot of plumbing code, these aren't as necessary as stipulated. Yes, there could be some duplication of code. But, especially in a microservice landscape, most of that duplication will be each service doing their version of whatever you are doing.
With newer tools and most important of all: your rewrite being a second iteration with the benefit of hindsight, you may notice that the language choice makes less of an impact than expected. IME language differences tend to shine the most when writing something the first time around and its expression capabilities shape the design.
My sense was that the big benefit was actually down to the libraries and syntactic sugar. Dataclasses, comprehensions, and generator functions all have a big impact, but so do things like Python just generally having more ergonomic libraries for talking to databases, producting/consuming REST APIs, and reading/writing data files.
To put it another way, perhaps what you tested here was having a great array library vs not having a great array library! It would be interesting to test Python vs Java on a more generic task, where there isn't a huge advantage from a particular library or set of libraries.
It would also be great if someone would write a NumPy-style library for Java. My team would use that a lot!
Why ? Because it's easier to write great frameworks in Python.
It compounds.
I don't think the lack of any of these things is because they're impossible, or even particularly hard. It's because the community did or did not have the need and energy to build them.
Remember that Java had Java EE forced on it quite early. That defined web development for a long time. Eventually Spring overthrew it, and became the new tyrant. Neither of those are quite like Django, or Rails, but they occupied the space that would have had to have been empty for a Django or Rails to emerge. I think this is rather unfortunate; it would be really useful to have a vibrant Django- or Rails-like framework in Java.
There's no Typer because there isn't much of a culture of building command-line tools in Java. There are some good command-line parsing libraries, but nothing with all the bells and whistles and extensions that Typer has, as far as i know.
Of course there is. In fact, SQLA is easier to use than Hibernate, and the ecosystem around is as rich. As for Spring, crossbar fits the bill.
> I don't think the lack of any of these things is because they're impossible, or even particularly hard. It's because the community did or did not have the need and energy to build them.
I've never say they are impossible. I've said they require less energy to build them in python :)
Also Hibernate is terrible and should be erased from the Earth.
I don't think this is entirely down to the language itself. A lot of it is cultural. The Java community tends to favor a more "enterprisey" style of programming and API design that, at least to my tastes, tends to be rather overwrought. That could change, in principle. I'm just not optimistic about it happening any time soon.
I wish there was some way i could take a sabbatical from my job and come and try to rewrite your service in what i consider effective modern Java. I am pretty optimistic it would be competitive with Python in effort and maintainability, but i have no way to prove it!
Things like this are the greatest disappointment to me after fifty years of writing software. It should not be necessary that Java equivalent exists. It really ought to be possible to connect to libraries from any language by now.
Numpy is just wrappers for BLAS and LAPACK.
Certainly not. I encourage you to have a look at the API of BLAS/LAPACK to perform basic operationa. And that just scratches the surface.
It is. Maybe it's a very fancy wrapper.
CORBA got a lot of bad press, but towards the end, it was actually pretty decent. Certainly the best way to do the thing it was trying to do!
Not to your main point - but this line threw me off. I really thought you were discussing this as a "person of color", and I was so confused as the relationship of persons of color and Python. ^_^
In fact, in 2018, Java 8 was still massively used:
https://jaxenter.com/java-8-still-strong-java-10-142642.html
Then you have libs and frameworks, and they break compat too. The JS ecosystem basically broke everyone code for 10 years every 6 month.
It's not a Python thing.
This is not by way of complaining about Java 9. The changes are good and needed to happen. And the reasons why most people are still on Java 8 are also legitimate. Everyone's handling the situation in a fairly mature way, from what I've seen. The level of histrionics about Python 3, though... I've got a lot less sympathy for that.
Julia one of the target is solving the "two-language problem"
"Julia seemed to have solved the “two-language problem”—a conundrum often facing Python programmers, as well as users of other expressive, interpreted languages. You write a program to solve a problem in Python, enjoying its pleasant syntax and interactivity. The program works on a test version of your problem, but when you try to scale it up to something more realistic, it’s too slow. This is not your fault. Python is inherently slow—something that doesn’t matter for some types of applications but does matter for your big simulation. After applying various techniques to speed it up but only realizing modest gains, you finally resort to rewriting the most time-consuming parts of the calculation in C (most commonly). Now it’s fast enough, but now you also need to maintain code in both languages, hence the two-language problem."
https://arstechnica.com/science/2020/10/the-unreasonable-eff...
> Clearly, multiple dispatch, or some other way around the expression problem, is necessary for the kind of fluent composability that I’ve described above—but it is not sufficient. Julia has enjoyed an explosive degree of uptake in the scientific community because it combines this feature with several others that make it very attractive to numericists.
That's incredibly handwavy. So what's the special sauce?
There is no such thing as a free lunch when it comes to dynamic vs. static. It also seems like Julia is trading off expressiveness and easy of use in favor of efficiency, based on comments from people that have used Julia. It's one thing to be faster than any inherently slow language (Ruby, Python, Smalltalk, etc.), but keeping that flexibility and being as fast as C/C++ is a rather bold claim. Most languages hit some middle ground between the two, such as Java. But no one is under the delusion that trade-offs weren't made to get there.
1. You can't add fields to a type (struct) after definition. This means that Julia's structs have no overhead and are essentially equivalent to structs in C (although they are parametric)
2. No local eval. Eval in Julia only happens in the global scope and results of eval are only visible the next time you visit the global scope. This may sound kind of unintuitive, but in practice people don't generally use this for good reasons. This allows Julia to never need to de-optimize code. Once a method is compiled that code remains valid.
3. Macros. Julia has really good macros and other code manipulation (since it is basically a Lisp). This makes it possible to generate very complicated but fast code that you would never write yourself. The tradeoff here is that it makes the language more complex, but that's a pretty good tradeoff. (especially compared to the C/Fortran land of using a preprocessor that works on text).
4. Just-In-Time (just ahead of time). Julia at it's core runs as if it were highly templated C++ code. If everything got compiled ahead of time, Julia would be generating terabytes of compiled code and never finish compiling. Instead, Julia makes the tradeoff of only compiling for the argument types that are actually used in the program, which means that it only compiles a reasonable amount of code. The tradeoff here is that compiling small binaries with Julia is very difficult (not possible to do automatically yet).
The TLDR is that most expressive languages started by giving away as much expressiveness as possible, and then looked at how they could be sped up. Julia started by being a modern fast language and looked to see how much expressiveness could be added without slowing the language down.
A huge part of the design considerations for Julia essentially boiled down to "what sorts of dynamism and language semantics can we disallow while keeping the the good parts of dynamism"
The two biggest things that had to go in order to make Julia fast was
1) the ability to change the memory layout of a struct in a running session
2) the ability to eval in the local scope (our eval always occurs in the global scope)
These two things are huge performance problems. We might oneday solve 1) with Revise.jl (though it'll mean recompiling all your code if you do change the layout) but 2) is basically just a very bad idea and likely to never happen. Instead of a locally scoped eval, we have macros, multiple dispatch, parametric types, and generated functions. These give an incredibly powerful suite of metaprogramming tools that are beyond anything available in Python.
One of the key things in CL is that it has its metaobject protocol which forces a lot of decisions on what gets executed to runtime. There are ways to speed it up, but if you have something like:
(defgeneric foo (x))
(defmethod foo ((x number)) (print 'number))
(defmethod foo ((x integer)) (print 'integer))
(defmethod foo :before ((x number)) (print 'also-a-number))
(foo 3)
(foo 3.0)
Then CL won't call foo specialized on number when given an integer, but will call foo :before specialized on number. It determines this at runtime by searching for all applicable methods based on the type (at least as a first pass, you can cache this to speed it up but then you also have to have cache invalidation if a definition is changed).Julia doesn't have that aspect of CL's MOP. So this helps to simplify the search for applicable methods and dispatch. Even if it did all its dispatch at runtime, it would still be simpler. The other thing Julia does is aggressive JIT compilation. So if you wrote something like (with the Julia equivalent of foo from above):
function bar(x,y)
foo(x)
foo(y)
end
And, only considering floats and integers, later called it with each pair of float and integer then Julia would compile specialized versions for those 4 combinations. Now when you call bar it still has to properly dispatch it, but once inside bar the search for the correct foo can be bypassed because the types will be known. CL, again thanks to the MOP, doesn't make that as easy to achieve.- multiple dispatch
- parametricity
- lightweight subtyping
- staged programming
In interesting ways.
felleisen's class talks (partially) about it here: https://felleisen.org/matthias/4400-s20/lecture15.html
I guess the best way to put it is that Julia encourages a style where 90%+ of code can go through paths that are static.
Personally, I think Julia starts off as easy as python, but to get C++ or Fortran speed, you can't just code naively. Things go into a steep learning curve at that point, but perhaps there isn't yet as much know how about how to code "professional Julia" yet. There needs to be a book like Fluent Python or Effective C++ for Julia, or perhaps a condensed version of the Julia manual (see the 1 page zig manual for inspiration).
The other problem I have with Julia right now is lack of static type checkers. "modern python" (e.g, python in production in the last 5 years) tends to leverage the large ecosystem of things that hook into mypy (I'm taking about tools like pydantic) to reduce the inherent brittleness of the language. Ruby, php, and every other dynamic language has also seen that trend.
Right now, I've barely seen that with Julia, and it needs this badly for higher uptake in industry. It's why for example, perhaps you see a lot of Julia packages written for people's phds right now.
http://ucidatascienceinitiative.github.io/IntroToJulia/Html/...
As for expressiveness, this then leads to different programming styles which I explained in a blog post:
https://www.stochasticlifestyle.com/type-dispatch-design-pos...
Prototyping in Python is a viable strategy. Personally I tend to use Perl for that, because I am comfortable with it and I find it better adapted to quick prototyping than Python. I then rewrite in another language, usually C++, with more care about data structures and optimization. The author strategy to do everything in Python than optimize the slow parts can make sense, though I tend to prefer the opposite: white your framework/engine in a static language (like C++) and embed an interpreter (Python, LUA, etc...), as it is common in game dev.
But all that are language decisions! Whatever you do, there are going to be consequences. Choosing Python with C/C++/Rust optimizations is not bad, but it is a rather strong opinion, not the obvious default choice the author makes it.
You also end up spending a lot of time and resources in the later stages of the project trying to work around the problems: hire more people to work on the code base, spend a lot of time investigating optimization strategies, hire a devops team to turn the program into a distributed system with maybe hundreds of instances running on CloudProvider (TM)
... and give tens to hundreds of thousands a month to cloud providers when you could spend 1/100th if it were written in Go.
def div(x: Int, y: Int) -> Optional[Int]:
return x // y if y != 0 else None
The above should be what production level python should look like nowadays.https://www.infoworld.com/article/3575079/4-python-type-chec...
That is not to say the type checking situation for python is perfect, far from it, but the situation is good enough that arguments against python that utilize the dynamic nature of the language are mostly no longer relevant.
None/Null in many languages isn't type checked even if the language itself has static type checking. It can be passed to other functions mysteriously and remain a hidden error. Javascript and C++ are guilty of this. He's complaining of a safety issue most likely.
However note the type signature I specified Optional[Int] encodes the None type into the static type checker. In modern programming this implies Null/None safety at the type level.
I'm not up to date with the current state of python type checkers so those type checkers may let some things pass depending on some options you set. See my reply to the other guy for a more detailed expose.
This seems like the typical response of someone who posts bait ready to flex their "advanced programmer" knowledge. There's really no need to call out people's level of experience, ever. Personally I'm past the point where I think I've learnt everything so I just ask questions when I don't know.
The problem with your earlier post was that you weren't asking a question. You were stating something as if it were a fact.
The reason is because an exception is a runtime check. It caught an error that is hidden until runtime. The better way is to encode this logic into a static type checking system so this error isn't even permitted to run.
With the type Optional[Int], Any other function that utilizes the output type of this function MUST also be an Optional[Int] and not just an Int, meaning that the function signature must handle the None case and this is caught statically before run time. Insofar as the function type signature... python type checkers should handle this error.
However tbh I'm not sure how extensive type checkers for python are and whether or not they can track this type of error for the logic itself. For example while the below will catch a type error if you feed it a bad parameter, the signature itself could be wrong and not caught by a type checker:
x1: Int = 2 # results in type error (Good)
x2: Optional[Int] = None # results in runtime error. This is bad in terms of safety
def addOne(x: Optional[Int]) -> Optional[Int]:
return x + 1
As far as I know the above is permitted by python type checkers when truly correct code is below: def addOne(x: Optional[Int]) -> Optional[Int]:
return x + 1 if x is not None else None
However with python 3.10 pattern matching the level of static type checking can increase to the point where this type of safety is 100% possible: https://www.python.org/dev/peps/pep-0636/. x1: Int = 2 # results in type error (Good)
x2: Optional[Int] = None # No error (also good)
def addOne(x: Optional[Int]) -> Optional[Int]:
match x:
case None: # if this case were deleted should result in static type error (also good)
return None
case _:
return x + 1
Not sure if python type checkers can handle the above code, but exhaustive type checking is 100% viable in the future if the syntax utilizes the pattern matching shown above. Haskell and Rust already have this level of type safety. See: https://rustc-dev-guide.rust-lang.org/pat-exhaustive-checkin...What this means in short is that you are wrong. Whatever your notions of "modern" python should be having division by zero return a runtime error instead of a None or an Error type is definitively worse. We are at a point in our technology that this type of error should be caught automatically BEFORE a program even runs. All type errors should be encoded into types and caught via static type checking and NOT via exceptions.
div(x: Int, y: NonZero[Int]) -> Int
Exceptions at least have the benefit of retaining the location the context for the error occurred which gets lost without extra bookeeping when using optionals. NonZero = NewType(‘NonZero’, int)
if x != 0
x = NonZero(x)
just to call your function, especially since you still have to handle the case where someone does NonZero(x) blindly without checking. Should the constructor throw?The NonZero type should be responsible for checking the wrapped value is non-zero, you will probably want safe and unsafe constructor functions
Int -> Optional[NonZero[Int]]
Int -> NonZero[Int]
where the unsafe version throws.If this is overkill for your application then I'd prefer throwing an exception in div rather than encoding the failure in the return type.
This still moves the check out of runtime and into a static check. An exception means the error is handled runtime and as I said before is a definitively worse metric.
Let me put it more simply. An exception means that you can compile an error or mistake and it might accidentally make it to production. This happens because exceptions are only thrown on compiled programs during runtime.
An optional means that run time errors of this nature are impossible to happen. A program that throws an exception on division by zero is impossible to even EXIST. This occurs because if the optional is used, the static type checker will prevent such a program from even compiling. That is why it is definitively better.
Your encoding has really changed the meaning of the div function - it now always accepts any two arbitrary ints and always allows the possibility of returning None. So you've weakened both the precondition and the postcondition, which is now easier for the caller but imposes a cost on every location where the return value is used. As the optionality encoded in the value propagates further from the call to div, the context for the source of the missing value is lost. This is 'safer' in the sense of avoiding crashes at runtime but doesn't actually help resolve the actual issue of avoiding a zero divisor when calling div.
No it did not. That is simply your interpretation of it. I can rename "None" to "undefined" or to "Exception" and the etymological meaning is now more inline with your thinking.
The concept of exception or undefined can be encoded into a type or the set of all ints. What you are doing is saying is roughly equivalent to, "No I want to use the mathematical definition of Z, and I don't want to encode the reality of what's actually going on into the type."
You are getting hung up on a wording issue because you are stuck on the mathematical concept of undefined and division of Ints. Think of the problem from a higher mathematical perspective of sets and define a set (aka type) that fits the use case. That's just one perspective on the flaw in your thinking. There is another isomorphic perspective as well:
There is nothing weak going on here. There is no difference in saying that division by zero maps to "undefined" vs. Not defining it period. The concepts are isomorphic. When you construct the mapping in the english language when you talk to people. You literally say "division by zero is undefined" or "division by one is the same number" all these sentences are equivalent to saying "division by x/y/zero/ maps to undefined/z/somenumber."
Don't get too lost in mathematical conventions. Understand the higher meaning behind the concepts and you'll be much more flexible.
>This is 'safer' in the sense of avoiding crashes at runtime but doesn't actually help resolve the actual issue of avoiding a zero divisor when calling div.
You're not seeing the bigger picture. The only way to avoid a zero divisor is to use your made up type NonZero[Int]. There's no other way. Not even your exception handling can prevent a zero divisor.
The best way to do this is to FORCE the user to handle a zero divisor pre-compile time. This does not happen with an exception. Your program just crashes.
The pattern matching I illustrated above is a mechanism for forcing the user to handle any zero divisor related logic or the program won't Compile period. It is definitively better.
Think of it this way. Exceptions are cheating. It's some sort of escape hatch built into the language. It's better to have no escape hatches. Build all possible outcomes into the type.
The type for your static version of div is:
def div(x: Int, y: Int) -> Optional[Int]:
The precondition for the dynamic behaviour of div is that the divisor is non-zero. If you want to think of it in set terms this means y must be a member of Z - #{0}. But your encoding allows any member of Z. This is a larger set and therefore a weaker precondition.Likewise the postcondition of the dynamic version is that an Int is returned, but yours only guarantees an Optional[Int]. Again in set terms this is something like Z + #{None} which is a larger set than Z and therefore a weakening of the post condtion.
Your version encodes the dynamic behaviour of throwing an exception in the return type, but I'm saying that
def div(x: Int, y: NonZero[Int]) -> Int
is a more precise static definition of the behaviour of div, and if you're going to add types to the dynamic version you should prefer this approach. Our two definitions are not equivalent since your function can easily be implemented with mine: def yourDiv(x, y) = map (fun nz: myDiv x nz) (fromInt y)
but you can't (safely) implement my version using yours since it returns an Optional[Int] and an Int is required. This shouldn't be surprising since algebraically (Int, Int) -> Option Int is a larger type than (Int, NonZero[Int]) -> Int.> You're not seeing the bigger picture. The only way to avoid a zero divisor is to use your made up type NonZero[Int]
Obviously I know this because I suggested using NonZero[Int] in the first place. Your version accepts an Int divisor and just encodes the partiality in the return type. But this doesn't force the user (i.e. caller) of div to handle a non-zero divisor at all, it forces the user of the return value to handle a potential missing value without any context for why it was missing. Propagaging this missing value isn't 'handling' the error at all, since that can only be resolved by ensuring the divisor is non-zero before calling div.
Like I said think in terms of isomorphisms. Exceptions also force the caller to handle the code. There is no difference here.
None is isomorphic to a Null and an exception is isomorphic to a result type. If you want context. Do this:
type Result[Int] = Int | Exception[String]
def div(x: int, y: int) -> Result[Int]:
return x // y if y != 0 else Exception("Division by zero on line ${line}")
The structure of the code handling this type of div is identical to code handling an actual exception. The main difference is type safety. Exceptions compile with a possible runtime error, Result types do not.Additionally your NonZero[Int] is not fully type safe. You have to think about long reaching consequences and precursors. Here you aren't thinking about the precursor. What generates this NonZero[Int]? Usually at some point you have some type casting function of the form:
Int -> NonZero[Int]
What then happens when I pass a zero? All you did is propagate the issue to somewhere else.Math doesn't have the syntax to fully compose two functions with different codomains and domains. So strict math formalism is irrelevant here. You can switch some attributes to isomorphic concepts (like how mapping division by zero to an element in a set called "undefined" is equivalent to the operation actually being undefined) but in the end you have to invent something because math simply doesn't have any elegant formal syntax to cover this codomain and domain mismatch. Therefore strict adherence to the concept that a Int can never be inserted into a div function is unnecessarily pedantic. Literally every other mathematical operation covers all of Z in both domain and codmain except for division.
You would never write an exception handler to handle such a failure from div. The divisor being non-zero is a precondition of calling div in the first place, which is something the caller is responsible for upholding. You shouldn't ever need to write an exception handler to catch precondition violations. Do you also write handlers to 'handle' null dereferences? Representing the partiality in the return type is just pushing the responsibility to some code that can't reasonably do anything.
> What then happens when I pass a zero?
I've already explained this, you obtain a NonZero[Int] from a function
fromInt : Int -> Optional[NonZero[Int]]
and you can optionally add an unsafe version with type Int -> NonZero[Int]
> All you did is propagate the issue to somewhere elseYes, the check has to be done somewhere since that is the point of encoding the property in the types. But encoding it in the argument type ensures the check is done before div is called which is where it needs to be done.
This is just your arbitrary preference. There is nothing wrong with going from either perspective. But your exception is completely worse from every quantifiable metric except for your opinionated qualitative metric.
fromInt : Int -> Optional[NonZero[Int]]
This is functions suffers from the same problem you describe. You're just trying to justify a convention of doing this check before rather than later. Also Your unsafe version is again worse because it will trigger an exception on zero, so I don't see how it helps your argument.>But encoding it in the argument type ensures the check is done before div is called which is where it needs to be done.
This is the core of your argument and it is highly flawed. There is no "need" for it to be done this way. It is simply your preferred convention.
Your argument loses on both fronts. Exceptions are definitively worse and Encoding non zero type safety into the parameter is not necessarily proven to be better.
It's not arbitrary since it's possible to write your function using mine but not vice versa. If you disagree then please implementing the following function without casting:
def convertDiv(f: (Int, Int) -> Optional[Int]): (Int, NonZero[Int]) -> Int
> You're just trying to justify a convention of doing this check before rather than laterThe convention that callers are responsible for upholding the preconditions of the functions they call is well established: https://en.wikipedia.org/wiki/Design_by_contract. You obviously can't fix precondition violations by checking the result after the fact.
> Also Your unsafe version is again worse because it will trigger an exception on zero
That is the point of the unsafe version, yes. Sometimes you will statically know the argument is non-zero e.g. NonZero(3). If you want to avoid an exception then use the safe version.
First, Why does this even matter? It doesn't. Being able to write something in terms of the other doesn't mean anything.
Second you can't implement the converse without casting EITHER. The Optional[Int] doesn't exist so how do you create it?? You CAST. It's a zero cost implicit type in python and in C++.
>The convention that callers are responsible for upholding the preconditions of the functions they call is well established: https://en.wikipedia.org/wiki/Design_by_contract.
Should I use the fact that Optional is more well established then NonZero to win this argument? Yeah if you want to talk about "Well Established" then Optional is more well established then NonZero or this Design by contract convention that is so unestablished I barely even heard of it.
Additionally even reading about this convention I see no requirement that division by zero must never return an undefined or that zero should never be the divisor. The description reads that these pre/post conditions just need to exist, but they're your choice what you need them to be. These conditions are encoded in the type.
>If you want to avoid an exception then use the safe version.
The safe version suffers from your same problem just moved. Nothing is magically solved by this move other than it fulfilling your arbitrary opinion and convention.
The reason you can't write my version using yours is that the types are less precise and you can't recover the imprecision in the output type after the fact. The only safe way to obtain an Int from an Optional[Int] is by providing a default value which doesn't exist in this case.
> The Optional[Int] doesn't exist so how do you create it?? You CAST
By casting I mean an unchecked narrowing conversion e.g. of the type Optional[Int] -> Int. There's no casting in my version.
> if you want to talk about "Well Established" then Optional is more well established
This is a false dichotomy, contracts are still used in static languages where you can't or don't want to try represent properties at the type level. You could for example define a function
lookup: Map -> Key -> Optional[Value]
and still add preconditions that the map and key were non-null. The failure to uphold these represent a different kind of 'failure' than the key not being found so it wouldn't make sense to lift them into the return type.> The safe version suffers from your same problem just moved
It didn't 'just' move, it moved to the point in the program you actually need to deal with the possibility of a zero divisor i.e. before calling div. Where does the divisor come from in the first place? You seem to be assuming there is necessarily some call to NonZero.fromInt at each call site to div but this is wrong. The non-zeroness of the divisor could be established at some prior point in the program and used in multiple places. In contrast your version has to deal with the possibility of returning None everwhere even if you've already established the property of the divisor beforehand.
Irrelevant to my statement. I said why does it even matter not why can't you do it. The answers are it doesn't matter at all AND you can't do it for EITHER case.
>The only safe way to obtain an Int from an Optional[Int] is by providing a default value which doesn't exist in this case.
No the safe way is through exhaustive type checking via pattern matching. If you're not sure what this is, look it up. Suffice to say it's static safety on all sum types including Optionals prior to execution.
>By casting I mean an unchecked narrowing conversion e.g. of the type Optional[Int] -> Int. There's no casting in my version.
There is 100% casting in your version. 100% percent. There is no narrow conversion here you're just making that shit up. The inverse of what you wrote is THIS:
def convertDiv(f: (Int, NonZero[Int]) -> Int ): (Int, Int) -> Optional[Int]:
There is no way to create an Optional[Int] WITHOUT a typecast. I'm sorry, but your statement is definitively wrong no need to build some scaffold of strange logic around it and "narrow" the definition of a cast. I get your point though (even though I disagree). However, this does not change the fact that your example is completely wrong from a logical standpoint and completely off base.>and still add preconditions that the map and key were non-null. The failure to uphold these represent a different kind of 'failure' than the key not being found so it wouldn't make sense to lift them into the return type.
Uh no. You can do Exactly what you did with NonZero[Int] with Key in your example. Imagine a map with RGB colors as keys.
type KEY = Red | Green | Blue
type VALUE = ...
lookup: Map[KEY, VALUE] -> KEY -> VALUE
Like I said it's just your preference here. There is a false dichotomy when it comes to things being more correct when "Well Established" and that false dichotomy isn't coming from me. It's coming from you.>It didn't 'just' move, it moved to the point in the program you actually need to deal with the possibility of a zero divisor i.e. before calling div. Where does the divisor come from in the first place? You seem to be assuming there is necessarily some call to NonZero.fromInt at each call site to div but this is wrong.
Ok let me reframe this. I completely AM not Assuming NonZero.fromInt at the call point AT all. Once you realize that your assumption is wrong, maybe you should consider the fact that you're NOT understanding me.
>The non-zeroness of the divisor could be established at some prior point in the program and used in multiple places.
The above is 100% what I am assuming. This prior point involves the creation of the type NonZero[Int] which involves: NonZero.fromInt. Every other mathematical operation (+,-,x^y,/,) returns an Int not a NonZero[Int] so this cast must occur. And that is my point. Think about it.
> In contrast your version has to deal with the possibility of returning None everwhere even if you've already established the property of the divisor beforehand.
This is where you're getting hung up. Let's clarify something your NonZero.fromInt is of the type:
Int -> Optional[NonZero[Int]]
With that out of the way let us continue:Yeah so my division returns an Optional which could be a None. I can either handle the None immediately or let it propagate all the way to IO and handle it just before it hits this wall. This is a bad thing I get it.
But your NonZero.fromInt Also returns an Optional which could be None. I can either handle the None immediately or let it propagate all the way to IO and handle it just before it hits this wall. This is a bad thing I get it.
Notice how the above two sentences are the same? That is what I mean when I say you're just moving the problem to another place but the problem essentially remains the SAME THING.
As I stated before and I'll repeat it again. The only reason why you prefer NonZero[Int] over Optional[Int] is the same reason why someone would prefer blue over red. There is no logic, rhyme or reason behind it. It is just your style and your personal taste.
I've explained why it matters - the types are more precise in my version and if you start from that you can always throw away the extra precision if desired to get to your version. You can't go in the other direction, so starting from your version makes it impossible to safely recover an Int from the returned type of Optional[Int], even if you've already established the precondtion beforehand.
> There is 100% casting in your version
Creating an Optional[Int] from an Int is a conversion, not a cast. I thought it was obvious from the context but for the avoidance of any doubt, by 'casting' I mean an unsafe narrowing conversion. Optional[Int] is a larger type than Int, so it's trivial to create one from an Int:
def pure(x: Int): Optional[Int] = Just(x)
you clearly can't safely go in the other direction, whether using pattern matching or otherwise. If you disagree, just complete the following definition: def fromOptional(o: Optional[Int]): Int =
match o with
| Some(i) => i
| Nothing => ...
eventually you need to provide a default value for the case of no value.> Imagine a map with RGB colors as keys.
Your example doesn't make sense, what would you expect (lookup Map.empty Red) to return? The optional return value is used to represent the key being missing in the map. Nonetheless the point I was making is that you wouldn't return Nothing from such a function in the event of a precondition failure e.g.
def lookup(m: Map[K, V], v: K): Optional[V] =
if m is None return Nothing
...
you would instead throw an exception if the input map is null and force the caller to handle it. The majority of static type systems are not powerful enough to encode arbitrary properties about values, so you have to decide which ones to check dynamically and which statically. Checking preconditions dynamically is reasonable if encodng them in the type system is too cumbersome.> This prior point involves the creation of the type NonZero[Int] which involves: NonZero.fromInt
No, this is not necessarily the only way to create instances of NonZero. You could have a PosNat subtype with members one: PosNat and succ: PosNat -> PosNat. You could have a non-empty list type with a length member.
> Every other mathematical operation (+,-,x^y,/,) returns an Int not a NonZero[Int]
They don't return Optional[Int] either so I don't see how this is relevant. There's no reason the input has to come from some application of a different operator, it could come from configuration, user input, a property from some other type etc. The question is whether and how to model the constraints in the type. The constraint exists in the argument so it makes sense to constrain the input type, not widen the output.
> Notice how the above two sentences are the same?
Yes, if all you want to do is avoid establishing the property you care about and silently propagate some information-free 'failure' value to the top level, then you can do it either way. But the entire point of encoding properties in the types is to force you to establish them. These statements highlight the difference:
1. I've established the divisor is non-zero, called myDiv, received an Int and continue
2. I've established the divisor is non-zero, discarded that information to call yourDiv, recieved an Optional[Int] which cannot be empty, but which must be propagated. You could immediately unwrap the value but now you're just re-creating the dynamic behaviour of a function (Int, Int) -> Int which you've already rejected.
Ok man, take a look at this: https://en.wikipedia.org/wiki/Type_conversion
There is no garbage redefining of a well known word to fit your convenient argument. We are using English and well known terms. This is not obvious to anyone.
Let me illustrate what happened. YOu used a well known term INCORRECTLY, then redefined later claiming that your redefinition was obvious. That is a comedy excuse. No.
You were using a term incorrectly. You man handled the term and you are wrong in that regard.
> eventually you need to provide a default value for the case of no value.
So? It's still type safe. My point is you can't get rid of this problem at all, not even with NonZero int. But via an optional type you guarantee that it's handled. The Rest of your composition needs to live in the Context of that optional. This problem isn't solved by moving it to a post condition.
>Your example doesn't make sense, what would you expect (lookup Map.empty Red) to return?
Ok this I admit I'm wrong. I wasn't thinking detailed enough in that respect so my example was wrong . But this doesn't prove your case at all. There is nothing wrong with using a Maybe monad here. There is zero reason why you need to throw an exception.
>No, this is not necessarily the only way to create instances of NonZero. You could have a PosNat subtype with members one: PosNat and succ: PosNat -> PosNat. You could have a non-empty list type with a length member.
This level of pedantism is too extreme. Nobody uses that type of arithmetic. But if you want to go there fine. You now have a hole in subtraction as 3 - 3 would be a singularity. Your program can no longer use subtraction or addition.
Division and all other mathematical operations have incompatible types. Type casting is 100% necessary to use all operations. The only way I see your Peanno NonZero safetly working is to restrict division to be the exact First set of operations. Then a safe cast can be done without an optional from NonZero to Int for subsequent operations. This is the expansive type casting that is safe and what you're referring to earlier. Either way such an ordering restriction is highly unreasonable but it does offer total safety at the type level.
>They don't return Optional[Int] either so I don't see how this is relevant. There's no reason the input has to come from some application of a different operator, it could come from configuration, user input, a property from some other type etc. The question is whether and how to model the constraints in the type. The constraint exists in the argument so it makes sense to constrain the input type, not widen the output.
If it comes from user input are you saying there's a terminal where the user can never input a zero? If it comes from IO anything goes, you have to be prepared for Anything. Usually the input is String and String -> Int has a shitload of holes.
>Yes, if all you want to do is avoid establishing the property you care about and silently propagate some information-free 'failure' value to the top level, then you can do it either way. But the entire point of encoding properties in the types is to force you to establish them. These statements highlight the difference:
>1. I've established the divisor is non-zero, called myDiv, received an Int and continue
Number 1 cannot be established without creating a singularity or creating strange restrictions as I explained above. If all your program ever does is do division without accepting terms from IO then fine, you win, but let's be real. This is not a realistic expectation and that is my point.
> But via an optional type you guarantee that it's handled
It's not 'handled', which is the point. There's no default value so you have no option but to propagate the missing value. Callers that establish the precondition still have to deal with an Optional value even though it can't be empty. Callers which establish the precondtion through the types receive a more precise type and don't have this problem. Callers that establish the precondition in the more simply-typed version (Int, Int) -> Int also receive an Int. Only your version imposes the imprecision on the return type.
> There is nothing wrong with using a Maybe monad here. There is zero reason why you need to throw an exception.
The Optional[Value] returned from a map lookup is used to incidicate the possibility of a missing key, which is an expected outcome of the operation itself. Passing a null map represents a structural error in the construction of the program at the point the call is made. If you represent both in the same type you can't distinguish these two cases.
> You now have a hole in subtraction as 3 - 3 would be a singularity.
There's no hole here, subtraction on NonZero just has to return Int instead.
> If it comes from IO anything goes, you have to be prepared for Anything
Yes, if it comes from user input you have to establish the property dynamically. Encoding the property in the argument type forces you to actually do it (whether statically or dynamically), and help you push the constraint to the highest level. Putting the optionality in the return type doesn't force you do do this, and imposes a cost on every call that actually does.
I would call it fixation if it was some obscure reference. But it's not. You're calling a cat a dog. Literally everyone knows what type casting is. I'm saying hey buddy I know what you're talking about, but a cat IS NOT a dog.
I've always accepted the intent of what you're saying. You're just not understanding my intent. You literally didn't even address division coming before all other mathematical operations. I don't think you understood.
>It's not 'handled', which is the point.
It is handled. Otherwise it won't compile.
>There's no default value so you have no option but to propagate the missing value.
Understood. Let me put it this way an "=" represents a section of a program with no Optionals a "-" represents a section of a program with Optionals and an F represents a section of a program with a function that does division. A program is can be represented by a serial list of these characters. To you a program is bad if it looks like this:
-------------------F====================
After the division call the output is polluted with optionals because F returns an optional. You think this is bad.
You think a program should look like this:
-------------------F---------------------
And you think this is achievable with division if you use a NonZero Int type.
What I am saying is that it is not generally possible. To create a NonZero type your program will be polluted before the function call simply because Int->Optional[NonZero[Int]]:
===================F----------------------
>There's no hole here, subtraction on NonZero just has to return Int instead.
First of all this a type cast. subtraction and addition are defined on sets, not between two different sets.
Second off this defeats your whole point. You wanted to create a NonZero int for safety subtraction just made it unsafe again.
Third off does it return an int for 3 -3 and a nonzero for 4 - 2? So your subtraction returns a variant type. That basically opens up the same singularity as your null once you try to evaluate this sum type via pattern matching.
Wasn't your point this:
f(x: NonZero, y: NonZero) -> int
If I do a subtraction operation BEFORE this function you can't even use the resulting zero value. So basically I have to turn subtraction into this which is identical from a cardinality stand point to returning a variant of a Int or NonZero: minus(x: NonZero, y: NonZero) -> Optional[NonZero]
> Encoding the property in the argument type forces you to actually do it (whether statically or dynamically), and help you push the constraint to the highest level.As I said above nothing is pushed to a higher level you're just moving the singularity around. Making it happen at the type cast rather then the division.
>and imposes a cost on every call that actually does.
You either pay the cost before division or after. Pick your poison, that is what I am saying.
The only way your methodology works is if Division is the first thing that happens. Then your program looks like this:
F--------------------------------------
That's the only time a nonzero type makes better sense then a function that returns an optional.
> The reason is because an exception is a runtime check. It caught an error that is hidden until runtime. The better way is to encode this logic into a static type checking system so this error isn't even permitted to run
But in
def div(x: Int, y: Int) -> Optional[Int]:
return x // y if y != 0 else None
don't you have a runtime check? It has to do a runtime check of the value of y in order to decide whether or not to return None?I don't see how this is encoding the logic of dealing with divide by zero into the type system or permitting the error to run. I'd expect that to deal with it in the type system to not prevent it at runtime would require y to be a type that does not allow 0 (does Python's type annotations allow constrained integers?).
Handling it by returning None means the caller has to deal with None. If it deals with it by just return None too, then its caller needs to deal with it do. If whatever code finally really deals with it instead of just passing it upwards is farther up than div's caller or maybe div's caller's caller it seems like what you've ended up with is something that is less clear than exceptions and requires more runtime work.
For those cases where the return None approach is better than using an exception, can you get rid of the runtime check for y being non-zero with something like
def div(x: Int, y: Int) -> Optional[Int]:
try:
return x // y
except ZeroDivisionError:
return None
or do exceptions in Python occur sufficient overhead when not taken to make explicitly checking y win?As I mentioned this safety does not usually extend to the function definition. I stated that it can though (and it is in haskell and rust), and this is usually done through a feature called pattern matching. Pattern matching is available in python 3.10 but I'm not sure whether exhaustive type checking on pattern matching by external type checkers is available yet.
This is just defensive programming.
The None state here is already an invalid state. If you want to use the type system to ensure correctness you would change the second argument to be a type representing multiplicatively invertible integers.
For haskell it will give you a warning. For Rust it will actually not compile. The key here is that you MUST use pattern matching.
Having to wrap every division call in a maybe monad or rust result enum is bad from a usability standpoint and potentially unacceptable from a performance standpoint.
Exhaustive pattern matching is great. But not a replacement for other types of error handling.
There are languages that employ this philosophy everywhere. Such languages are actually incapable of having a runtime error outside of edge cases like ffi.
Just so you know, you don't have to wrap every type in an enum. Types and enums share a bijective relationship. A type IS an enum and vice versa. You CAN in fact define an Int in terms of an enum (and an enum in terms of an int, in C). The mechanisms behind an Int and a bool are roughly the same, it's just the cardinality and mappings that are different. Pattern matching implementations usually check types based off of cardinality and do not differentiate between enums and other built in types like floats or ints.
I think you would be surprised by the amount of type-checked bugs that can be eliminated in dynamic languages with standard software engineering practices. have a look at how common lisp, for example, handles type annotations and type warnings: https://lispcookbook.github.io/cl-cookbook/type.html
instead of blaming dynamic languages, i think project managers need to take on better practices
> but the larger the code base, the more you need strong guarantees over flexibility
i think i disagree with this. the larger the code base the more flexible the program needs to be towards its input, otherwise your program is too sensitive and it will crash. the output however i think should be strict
To second this, I just read someone's anecdote about how unit tests/TDD is eliminating the need for type checking: https://www.artima.com/weblogs/viewpost.jsp?thread=4639
> i think i disagree with this. the larger the code base the more flexible the program needs to be towards its input
And again, I'll second this with another anecdotal posting, but with the emphasis that input is not just run-time, but also specifications for programs: http://ivy.io/common-lisp/2015/03/03/guerilla-lisp-opus.html
And the interoperability is not as awesome as it's painted in the article, it's always more pain to have more languages in a project that need to talk together
When Knuth wrote that quote, he meant something very specific with the word optimization: low level micro optimizations. Spending a lot of time trying to get every little ounce of performance from every little CPU instruction.
High level reasonable design decisions are not "optimizations" in this sense at all. You should think of them as non-pessimization.
Choosing Python or another slow language is premature pessimization. You just make your program slower for no reason.
So choosing a language based on how programs written in it perform is simply non-pessimization.
If you are interested in an expansion on this idea, checkout Casey Muratori's lecture about philosophies of optimization: https://www.youtube.com/watch?v=pgoetgxecw8
The article makes claims about software being "fast enough" and not needing further optimization. But what you call "fast enough" is probably 10000 times slower than what it could be.
I'm not joking.
People who program in slow languages have a really skewed perception about performance.
If a page takes 5 seconds to load, they don't see that as a problem.
If a server takes 500ms to respond to a request, they don't see that as a problem.
This is simply unacceptable.
Using python for shell scripting requires extensive knowledge about the language and the standard library and all the gotchas that you can fall into.
This sort of time investment does not pay off if you just want to write small scripts that perform simple tasks.
If the task is simple, then a bash script would be better.
If the task is not so simple, then a language you are really familar with is better.
If you accept my premise about how programming in slow languages is generally a bad idea, you would not spend that much time learning any such language with all its ins and outs and gotchas.
For background, I personally have spent a lot of time learning python, and only later in life have a realizd what a mistake that was. Now I only use Go for server side programming. I'm also using it for command-line programs. Such programs usually start small but grow over time. Choosing Python for them would have been a costly mistake.
Bash script don't work on the Windows command line, and bash relies on UNIX tools as its "standard library", which also are not available on the Windows command line. That's the whole point really, Python has very good cross-platform support (and in this case cross-platform doesn't just mean "Linux and macOS", but also Windows).
> If the task is not so simple, then a language you are really familar with is better.
IMHO every programmer should be familiar with at least a handful programming languages.
Also doesn't windows now ship with a linux compatibility layer anyway (WSL)?
EDIT:
I just realized you are the guy who wrote sokol!
If your main work station is windows and you don't work on linux, then replace bash from my previous comment with bat(ch).
Asking Windows users to run Bash scripts is essentially the same as asking them to build a C/C++ project that uses Autotools and Makefiles. It may work with lots of tinkering, but it's a massive world of pain.
> If your main work station is windows and you don't work on linux, then replace bash from my previous comment with bat(ch).
I usually switch between macOS, Linux and Windows at least several times a week, sometimes several times a day, maintaining different scripts to do the same thing wouldn't make sense.
Windows was built by people who started by copying from Unix. This is part of the reason Bash works fine on Windows (ie What is your problem with it "mapping" to things?). You can pipe commands, you can stream text, you can edit binaries. There are things that slow down (or can't be done, modify permissions easily, etc) because of Windows tradeoffs, not because it's fundamentally different. An OS does OS things and a shell talks to the OS.
If you OS vendor refuses to supply you with appropriate tools to do your job, then just chose another vendor. Vote with your foots.
Bash has one strong point: 4 to 5 lines admin scripts for which concision and expressiveness matter more than anything else. That's it.
There is nothing else it can be better at than ruby, python, perl, php, or even go or kotlin.
The argument you need to learn x stands for bash as well, btw.
Sure, but that isn't to say his advice can't be applied in other contexts. Applying optimizations before knowing where a bottleneck is just a bad idea, regardless of the situation. To constrain the advice to the original context is sort of odd, but maybe there's more to it than that?
I always thought the risk of premature optimization was similar to the lean principle of deferring commitment: you should to wait until the last possible moment to a decision that has high switching costs. In both cases you want to have as much information as possible to taking action.
As an aside, this bugged me for a long time: how _much_ information do you need? Recently I co-presented with the CEO of a successful manufacturing company. At his firm they use the military's rule of thumb of taking action when you have 70% of the information. Deciding sooner than that and you risk making a bad decision; Waiting for more information after 70% and you likely are missing an opportunity.
But, general high level design decisions that don't produce slow code by default do not suffer from this. They don't make the code any complicated or hard to understand. Quite the opposite, they usually simplify things. Because most of the time, the way to get good performance is to do the minimal thing that needs to be done to perform the task at hand. The code you write will be very straight forward. There's no switching cost.
On the other hand, choosing a slow language has a high switching cost. You can't just take a bad design and try to optimize it. It usually doesn't work. You have to rework the whole design.
EDIT:
I have one other important point to add:
> Applying optimizations before knowing where a bottleneck is just a bad idea, regardless of the situation.
Ignoring any and all performance concerns using this kind of reasoning will result in a codebase that is permeated with slowness. The program is slow but you can't identify any single bottleneck. You end up with a monster that simply can't be fixed.
Worse, you may not even realize how slow it actually is. If you have profiled it and found no particular bottle neck, you may think this simply the limit of computer performance and the task at hand cannot be performed any faster than this.
This actually leads me to see the concept of "micro abstractions", and I'm going to see if I see those in hard to deal with programs, now.
Premature abstraction really is very very bad (I won't go so far as to say it's the root of all evil).
The higher level the abstraction, the worse it is.
Unless it's born from experience writing the same kind of program many times. In which case it's not "premature".
"micro abstractions" are actually fine. You notice a recurring pattern in the code so you "abstract" it away into a function.
I have in mind the projects that went from a handful of data structures and functions to dozens of each. Typically in the vein of abstracting away whatever immediate problem was faced in the code.
Have a workflow? Make a workflow abstraction that can run any work.
Have code that needs to retry? Make a retry facility that can retry for any reason. Than make a grammar to specify what backoff should be used.
While doing those, realize you want a rules engine to determine what work and what backoff to use.
Then, put in a sat solver. To determine if a plan can be done.
...
None of these are really bad things, in and off themselves. Odds are high they would all be poorly done, or themselves a giant undertaking.
Last responsible moment, not last possible, moment. That's a big difference. You shouldn't start on the work at the last minute, that's too late. But you also don't (necessarily) need to start on something 2 years before it's due (unless it's big and/or there are anticipated problems that will need to be addressed that the time would permit; that's what makes it the last responsible moment).
There's a great deal of scientific literature about this topic, BTW. Especially in machine learning.
Donald Knuth writes entire volumes of books in assembly language, because the details matter. If you want to have that point, then you probably shouldn't be quoting Donald Knuth.
Donald Knuth is the guy who points out that "for(int i=size-1; i>=0; i--)" is faster than "for(int i=0; i<size; i++)". When he says "don't sweat the small stuff", its because his books are filled with _incredibly_, detailed stuff. Far more detail than any other programming book / algorithms book.
IIRC, the "premature optimization is the source of evil" statement was about how GOTO statements and how efficient they were. The __context__ is that GOTO is indeed a very fast construct (especially in the context of replacing recursion).
> Programmers waste enormous amounts of time thinking about, or worrying about, the speed of noncritical parts of their programs, and these attempts at efficiency actually have a strong negative impact when debugging and maintenance are considered. We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%.
From: "Structured Programming with Goto Statements", emphasis mine.
Archive link: https://web.archive.org/web/20131023061601/http://cs.sjsu.ed... (I haven't read it in a while, and it's 41 pages, so I probably won't get to it today.)
I translated it to C so that people got the gist of it.
You can design slow crap in any language, but I'd rather use something that has the potential to be robust and fast for my problem domain so I don't have to rewrite the entire app in the future to language "X", when I could have easily started with "X".
Make no mistake - if you do this, you made an upfront error in selecting your language for the business problem. It may not be an unrecoverable error, but it is a failure in judgement regardless.
It's now right up there with "X considered harmful", which 80% of the time translates to "I don't like X".
Specific to Python from the article, I'll add that I personally am done with dynamically typed languages for anything where more than one person needs to work on it or it needs to be long-lived. Static typing is simply too useful. Dynamic typing is (IMHO) a false economy. I've often said "Python means writing unit tests for spelling mistakes". Refactoring a dynamically typed code base of any reasonable size is a comparative nightmare.
Shell good for simple scripts. — Shell scripts are pain to support!
Perl is astonishingly good for text manipulation. — Perl is write-only language!
Python is easy to learn and start coding. — Python2 to Python3 migration is decades long hell!
JavaScript is the language of the World Wide Web. — JavaScript is nightmare to refactor!
And so on.
If language easily accepts anything that developer pops out of his head, then program will be hard to support and refactor.
If language limits developer to certain good, safe, performant solutions, then it hard to write programs in it, because you will need to learn good patterns first, which may take years.
People really underestimate the time it takes for devs to figure this out. My company has been on TypeScript for over three years now and it's a giant mess. Everything is arbitrarily typed with no design philosophy or organization behind anything. Which leads to these massive interfaces with every single field being optional. Which means you're still doing run-time checking all over the place.
I've always maintained that if your team sucks at organization then no amount of tooling will save them. You're not so much working as a team but as bitter enemies. Possibly the turnover is too high in the company. You have mercenaries rather than ideologues. Not that I'd want to work with a team of pure zealots. But there should be enough people holding the architecture in place. It helps morale as well. You don't want to work on a project that no one gives a shit about.
On first load, not only the page is ready, but the video has started to auto play (including buffering), under 3 seconds. Less than 500ms on a reload.
The whole project has been running for 10 years on 7 servers. Hosting costs under $1K/mo.
And even better, the code base is a mere 11K LoC, so a single dev can manage it. And make a living out of it.
That's because:
- Python is almost never the bottleneck in a website
- Python wraps compiled languages for slow stuff. That's the whole point.
- your archi matters much more than the language you chose. Because good db providers, proper caching and optimized data fetching will be the pareto moves, unless you are google size.
- you will not code most things anyway. You are not going to code the video encoding, you will use ffmpeg. You will not recode a database, you will use an existing one. You will not code a server, you will use nginx or something else. What you will spend time on is creating logic for your clients.
So I think your are extrapolating this rare use case where your POV makes sense, and ignore the 99% of the case were, indeed, a slow language doesn't matter.
In fact, if electron success tells anything, it's that while I do value speed a lot, the market has voted.
Sounds much more likely that you wrote in python the code that bundle together the libraries that peform the actual network streaming, machine learning, encoding, search, etc, and that those libraries were not written in python.
Where in the stack do you see that Python performances matter?
Only one software out of a 1000 will actually need it. The others will mostly do plumbing or automation.
And scripting languages excel at both.
If you don't fall in that category, then yes, do use a fast language. If you are coding a chat server, use Erlang. If you want to do a rendering engine, use Rust. But if you most likely want to code your company CRUD software and glue it to some external services and providers, use Python.
As we just saw, most of the code forming your stack is not written in python because its performance would not allow it.
> Only one software out of a 1000 will actually need it.
Not what I see from your example.
It seems to me that you are overinflating the volume of the python code that you wrote and are blind to the huge volume of performance critical code you just take for granted.
You probably mean: most of the code that most of the programmer write is disposable code assembling reusable libraries. This code performance does not matter therefore python is fine.
I agree on this. The problem I see with this reasoning though is that since python is so slow and error prone compared to compiled alternatives that it is unfit when scaled up from small disposable programs to largish permanent software, in a world where all programs trend to grow and stick in production for longer than expected.
I don't know what is the website you're running or how many people use it or how much money you make. I know you can make a lot of money doing stuff in python or electron. Does that mean the market has voted that performance doesn't matter? I don't think so. Jira makes a lot of money but almost everyone hates it. I know a lot of other software that is very slow but still makes tons of money.
> you will not code most things anyway
You say that like it's a good thing! It's a bad thing.
You may say it's because you don't need to. But the flip side is that you can't, even if you wanted to.
You can't program a database or a video encoding or a web server in this language you are using to write your program.
So what is the point of this language?
You can't even use it to cache things in memory. That's why everyone uses redis or something similar.
In this case, since the entirety of the project is to be written in a specific language, it's usually worthy of placing that in the more macro-category - where that choice affects all other decisions.
Don't spend 4 hours of your time on a `for` loop but half an hour of your time deciding on the language and architecture.
I don't mean that as a criticism, either. I keep meaning to strive to having the level of cognizance to memory usage that he just drips into his software.
As a fun project, I started writing some code to do calculations on matrices. Wrote it in Python and used Numpy. Then I wrote it in Julia. Almost immediately it was a 100x improvement. I spent time doing "premature optimization" and netted a 10x improvement. So overall it went from minutes to milliseconds.
I guess if you're writing a website or something then I wouldn't worry, but for other things, doing optimization from the get-go can save you a ton of time and effort later on. And really, such optimizations can affect the architecture in ways you wouldn't think of without it. Reworking your architecture down the road is a huge time sink. I like to make sure mine is close to where I want it to be in the beginning.
Despite popular sentiment on HN and Reddit, a website is not a piece of software that can tolerate slow code. Unless you are ok with your website dying when 100 people visit it at the same time, or otherwise investing in a complicated "cloud" infrastructure to make it possible to have more than 100 concurrent visitors.
Viewing language choice as a high level design decision for a project of nontrivial scale is a symptom of bad high-level design decisions.
Requirements that are good enough to warrant a 5 second page load time being "unacceptable" must have, at some point, been prototyped with real use cases, real data, and in real environments.
Getting to that point - where the business case is connected with such a strong through-line to the user experience can literally take years. Early in a project lifecycle it's more important to design software that is easy to change than software that is fast.
Language choice does not matter, performance does not matter; the design philosophy of the developer(s) when building something new is everything. That philosophy should be "I'll likely have to change this" not "I need to make sure this runs quickly". Any high level design decisions should be made with the former in mind, not the latter.
Once you've earned that product wisdom, sure - carve out the slow thing and write it in CUDA or something. If things are easy to change, that's not a scary proposition; it's just part of the natural project lifecycle and Knuth lives to die another day.
"We either have perfect information or we make software that is very slow".
Optimize iteration frequency of features not for loops.
This kind of code does not tend to be rigid as you might imagine. Quite the opposite. Since it doesn't introduce a ton of premature abstractions and stays straight to the point, it tends to be flexible and easy to change/adapt.
You are still thinking in terms of a false dichotomy:
- Flexible but slow
- Fast but rigid
But it's a false dichotomy.
You don't need optimizations to get high performance software. Computers are very fast. You just need non-pessimized code.
It's funny. People use "computer are very fast" to dismiss all performance concerns. But this is not how we should think about it.
"Computer are very fast" means we don't need to optimize (as in micro-optimize on the instruction level) code to get high performance. We just need to get the clowns out of the car.
Sometimes yes, you can eat your cake and have it too w/r/t performance and flexibility; but the choice of programming language cannot be categorically relegated to "always pick the fastest one". There are no silver bullets, there will always be tradeoffs.
Flexibility matters, ecosystem matters, developer experience matters, third party libraries matter, good requirements (or lack thereof) matter, budget matters, timeline matters, security matters, and yes, performance also matters - it's just not the only consideration.
High level design decisions, language included, are not "pessimized by default" because they deem some of those other considerations more important than speed.
Visual Basic, PHP and Java come to mind. Probably C# as well.
Here's the conclusion, which summarizes the essay:
> All of this is to say while some companies do extract massive value from squeezing every CPU cycle out of their code, those companies also typically build data centres. So if you don't need a dedicated building to host your machines, please consider doing the math to see if it's truly worth making your developers less productive in the name of computational efficiency when you don't even know if that perceived efficiency is even necessary.
As an example, the kernel of my software is only a few hundred lines long. It's in C with AVX2 intrinsics to eek out the last bit of performance. The Python overhead is <1%, which means that even swapping out Python with a 1,000x faster language won't be noticeable. This is the '"glue" code' mentioned in the section on 'Use language bindings'.
Had I started with "It needs to be fast therefore I can't use Python", then it probably would have been a lot more work to do. (See, "numerous anecdotes of where someone implemented something in Python in 1/3 the time a competing team did creating the same thing in e.g. C++ or Java".)
I would say the article is very much doing that, it's just stopping short from spelling it out clearly.
It starts from the flawed premises that Python is much more effective to come up with actual production code than other languages [1].
The author then moves on to telling you that you should try at least a couple of times to code something in python before obvious performance shortcomings are glaring, then keep python as a binding language.
At no point does the author acknowledge that there may be other reasons to pass over python than performance.
From the way he approaches the subject, the author seems to live in the paradise where you only have to sling half-baked prototypes to prod, and then move on to another project while poor souls handle maintenance.
The take on why performance matters [2] also completely overlooks the fact that optimizing for infra costs is likely the least common reason to write for performance. Sometimes, if you don't reach a given threshold, your code simply doesn't work. Sometimes, your performance is your users' time (and sometimes, these users are even developers). Such companies don't necessarily build datacenters.
[1]: "This is what leads to numerous anecdotes of where someone implemented something in Python in 1/3 the time a competing team did creating the same thing in e.g. C++ or Java." -- of course, it's just anecdotes and not a claim, so the author won't defend it, but you're still expected to believe it.
[2]: "All of this is to say while some companies do extract massive value from squeezing every CPU cycle out of their code, those companies also typically build data centres."
That said, the few and limited empirical results which do exist, like Prechelt's empirical work from 20 years ago, support the author's premise.
As another example, Boehm's work across a range of projects (<100KLoc) shows that code size has the biggest impact on development effort, and slightly non-linear scaling in KLoC (KLoC^~1.1 in COCOMO); and various publications show Python generally requires fewer lines of code.
(As one example picked semi-arbitrarily from a Google Scholar search, http://jucs.org/jucs_19_3/a_comparison_of_five/jucs_19_03_04... on a small project has Python at the smallest effective LOC, smallest effective byte size, and nearly the smallest (after Java) eByte/eLOC size.)
Modern C++ is a much more compact (and complicated!) language than 20 years ago, so of course all of these previous results must be reviewed and reconsidered.
As a field, we have not provided that funding for solid empirical research!
> At no point does the author acknowledge that there may be other reasons to pass over python than performance.
Shrug. Okay. And from that you conclude the author is saying that Python should be used in all cases? "I'm going to talk about why you should eat more vegetables" doesn't mean "you should only ever eat vegetables." Similarly, "I'm going to talk about why you shouldn't reject Python from the get-go for a project you think is performance intensive" doesn't mean "you should only ever use Python."
The author even talks about Rust in an earlier blog entry, from August of this year, at https://snarky.ca/introducing-the-python-launcher-for-unix/ saying "And so over 3 years ago I set out to re-implement the Python Launcher for Unix in Rust. On July 24, 2021, I launched 1.0.0 of the Python Launcher for Unix".
So the evidence is that Cannon does not believe that Python is the best language for all projects.
> Sometimes, your performance is your users' time
Cannon's essay is about the case where people select a programming language based strongly on expected computational cost, and forget other factors in the cost model. Here's the relevant quote:
] I think the jump to selecting a programming language based on potential performance needs often comes from a place where people think their computation costs are more important to optimize for than their developer time. I don't think that always holds, though, as software developers are expensive.
In your model, people are choosing a language based in part of the cost of users' time, which means they they are not primarily focused on computation costs, and so are outside the scope of this essay. That doesn't contradict the essay, only show that it's incomplete. Which we know already from the text itself.
> Sometimes, if you don't reach a given threshold, your code simply doesn't work.
And sometimes you can reach a given threshold with an off-the-shelf use of Pandas, rather than build your own analysis routine. If you start off thinking "Python must be slow" and don't even know what your thresholds or performance characteristics are, then selecting (say) assembly over Python is probably premature optimization.
Cannon didn't even argue that if you aren't building data centers then you must use Python. Your [2] trims the full paragraph I quoted earlier. You left out: a) the "please consider", which makes this a request, and not a rejection of any alternatives; b) the request to also consider developer costs, because essay argues people should consider more than computational costs; and c) the "you don't even know if that perceived efficiency is even necessary", which ties the essay back to Knuth's "premature optimization is the root of all evil" quote.
Also the author: Pick python no matter the problem, the team, the libraries, the deployment targets, etc.
I read it as the much more milder "don't reject Python because of vague concerns about run-time performance".
I can't find anywhere which suggests the author things people should use Python to, for example, code up their web app front-ends or to implement a 'hard' real-time operating system.
Looked to me like a "Python isn't than bad for perf" article.
If Rust or something really saves money because it means we don't need 50 extra web workers for the same TPS then I don't think it's "premature optimization" - unless the dev salary is too much!
As it stands we're all just saying stuff with no data or evidence to back it up.
Sounds exactly like business management culture. Makes sense though, developer productivity is first and foremost a business management issue so discussions about it ought to follow the same path as business management. And business management love their anecdotes, great person citations and jargon.
And of course the further you get from silicon valley the more the hardware/service costs matter. A developer in Sofia easily earns an order of magnitude less than one in San Francisco, which shifts the whole optimization tradeoff a lot.
My team does a ton of large-scale processing (most of which is ETL), and we have TONS of servers at our disposal. We effectively throw hardware at our compute, and even then, jobs can take 5-10 hours to run. No big deal - just run the job before you call it a day and it's done when you start tomorrow morning. Or when you have daily jobs, you schedule it and it's just done whenever it's done.
I can't tell you the number of times I've looked at a job someone wrote, made a 5 LOC change in 2 minutes, and cut processing from 8 hours to 5 hours. I literally had one that went from 5 hours to 5 minutes.
These extremely simple investments save a lot of money, but it's a problem that's spread across hundreds of different jobs, and it takes so much effort to convince people that we should invest in it. A daily job that costs $1500/mo to run (based on amortized hardware costs) needs only be optimized once, and those savings are yours forever. But I've had the conversation that is effectively: my time is more costly than a measly $1500 job, so it's okay if the job is expensive.
- Nobody
Our codebase is well over 1M lines of C++. We have about 100k lines of python. Running pylint takes the same order of magnitude of time (half as much IIRC) as running a full optimizing build + linking of the C++ code base. We run pylint over all cores, while we run the C++ build only on a subset of cores on the build machine.
I would call that dog slow, unless you think python is not the appropriate language write pylint. And no, it is not fine.
Or even doing it at all.
I like the busywork. Less spooky than 'cultic ritual'.
Moreover, the availability of "average" (and "cheap") programmers matters a lot in the long run: if only genius can maintain your system, then you'll have problem in the long run because you'll either need to keep them at all price or need a lot of time to replace them. So, in the long run, you should better use a wide audience language with a lot of available programmer (even if they are "average") than a specific language requiring good programmers. However, for an MVP, you can recruit a genius programmer using the fastest tool for the job.
Obviously, some domain are more oriented toward some language... and for ML for example, python is quite a good choice because of the libs (as Java could be for - lets say web servers)
So it matters a lot what your system will be used for and how much time you will require it to run before needing to rewrite it from scratch
If one of these steps requires a change of programming language, I'm never too lazy to recode :)
And by "make it fast, make it work", I mean benchmark first. Kind of like test first, you write the benchmark before the implementation and constantly look at how fast things runs, and ensues every single bit you add is fast.
Want to build an I/O utility writing to a DB? Sure C can do it, but Python is better suited. Want to write a toy compiler? You don't want to waste your time trying to wrangle on CPython extensions. C works out of the box.
> _'if you select a programming language based on your preconceived notions of how a language performs, you will never know if the language that might be a better, more productive fit'_
Part of the CS education is not about recognizing homeruns but understanding trade-offs. The experience gain is about learning how tools work & which tools to choose to work in tandem. Modern systems use a variety of languages - JS in the webpage, SQL DBMS for queries, C++ to run the performance bits, Python for ML, introperations - maybe even Rust in the security bits of late. In that sense, the title was unfortunately misleading to me, since author tried to demonstrate a lot of usecases with Python.
Python is great - but there has to be a reason why other languages co-exist. Not just for bankers, military or some enthusiastic hobbyist.
But we often overlook the fact that it's possible to scale those two, incrementally. With today's convenience in interop, we can always start with a prototype friendly technology and later switch tech for real bottlenecks.
Sure, features are important and if critical ones are missing it might be a show stopper. So there is an initial thresshold that all candidates must pass.
But problem solving is not a one-off exercise. It tends to be both dynamic (=facing unpredictable challenges) and recurring over long time horizons. Which means having a healthy, engaged, resourced community that will invest in adapting / solving future requirements is essential.
So the "optimization" problem includes quite a bit more than the presently known developer team, its software stack its hardware and current problem definition / user requirement.
I think you see this dynamic in several cases (including python) where you might not think that it makes rational sense.
You could just as easily say "Descriptive variable names are a form of premature optimisation. Code should be written with single-letter names, and descriptive names only added when absolutely necessary."
https://meadhbh.hamrick.rocks/v2/technical/six_hot_languages...
I mean, sure I can prototype in python, then put a lot of additional effort into bringing performance of the bits I absolutely need fast close to native code performance.
Or I could code everything in C or Rust and get everything in native code from the get-go.
To be honest, today there are many viable solutions with different languages. I would not recommend to start with a plan where you have to reimplement the system in another language at some point. Experience tells me that almost none of such projects survive.
Just because you can program a new feature quickly doesn't mean it is engineered well for the long-term & larger scales.
>Prototype in Python.
Stopped reading here.
Wait, what? The only thing I know about python is that it's indentation-sensitive, no idea about syntax or libraries. Suggesting me to use python is premature optimisation.