I think it was 2 things: one, it is genuinely harder to write a valid C++ program. The language imposes non-trivial constraints on the programmer, for better and worse, and you have to navigate them for the thing to even compile. But two, C++ invites pedantry. I heard endless discussions about "ooh, template meta-programming", "struct or class", "but is this idempotent", "you should be using move semantics", "why on earth is this a unique_ptr".
I'm not saying C++ is worse, or that its trade-offs are not worth it, but for sure, from my experience, translating an algorithm into code is done far more quickly in Python than in C++. YMMV
Ah and also, this was definitely not a question of deficient C++ coders. The hiring standards for C++ were so high we were always short of devs. Meanwhile, Python prototypes were usually written by part-time ex-academia Python dabblers.
I still agree that Python can be much more efficient to write than C++. I just wish it wasn't so unbearably slow.
Even then though, I have the impression that Python just makes it easier to write it all more quickly. It just doesn't invite you (as much) down rabbit holes of doing things even more properly.
Interestingly, I've consistently had the opposite experience. In C++, when I ask, "how do I do X" there will be a few different options, but pretty much any of them will do. As long as you avoid things that are definitely UB, you're fine. In Python, I'll find endless religious arguments about which is the most "Pythonic" method. Whenever I try to just hack out Python code to just get things done, I always feel like I'm being judged, because usually the quick-and-dirty method is far from the most Pythonic.
Ruby on the other hand embraces many different syntaxes to do the same thing, but they’re all accepted as correct in the traditional view.
I wonder what the result would have been in eg. Go or Rust, where that level of conversation doesn't exist. True, others would have taken place instead, but at a higher level nearer the design, where they should have been happening anyway.
I've heard reports like this for many years.
My pet theory is that the discrepancy here is ~10% static types and ~90% memory management. Having a GC (or refcounting or whatever) fully baked into the language so that you don't have to think at all about how values are passed around and stored is a monumental productivity boost. Possibly the biggest win in software engineering productivity in the history of the field.
This seems like a monumental level of operational waste. I'd love to hear more about this particular setup
-PEP 20 -- The Zen of Python
—The next line
This ZoP line is by far the most misused cliché of all things Python. How is it determine that “prototyping in C++ and rewriting in C++” is not the obvious “one” way? Because you don’t like it? A Zen is intentionally self contradicting to allow introspection, not judge others :)
-- Django
-- Perl
-- assembler
-- fpga
Back at you.
-hdl
You would initially trade some of the C++ performance, but the code shouldn't be more complicated or difficult to write or compile than Python. Then, you'd only need to change the performance sensitive parts to use real C++ arrays or structs.
Does that mean Python has all the performance you would ever need? No - we know it doesn't, but we also know that you rarely need that kind of performance and when you do it's usually localized to a very narrow portion of your code. Take that portion of code and implement in C and optimize to your heart's content. Even with that you'll still get your project done much faster.
Python is the tool of choice for those developers who just want to get stuff done.
I’m primarily a python user, but I spent time writing code in Ocaml to understand the potential benefits, but… (I would love to have my mind changed)… it feels like many of the features people touted about Ocaml have already made their way into python…?
1. Immutability / referential transparency seem more like nice-to-haves for a codebase, rather than the real reason people tout FP…?
2. Sum/product types and pattern matching are being added soon to python
3. Mypy is starting to gradually enable python typing, although I assume it’s still a work in progress
4. Python allows us to use map/reduce, and to pass functions around as arguments…?
I want to understand the potential benefits of FP more, but my experience with Ocaml hasn’t shown me great improvements yet. Open to having my mind changed.
In python, most of these "features" are done half right and duct-taped on at the last second as an afterthought.
I’m much more productive in Python. Not as in “I feel” more productive, but as in measurably more productive.
Where Python falls short of something like .Net and C# is that I probably wouldn’t have been more productive in Python if I didn’t have 7 years experience with a rather strict environment, and as far as using Python on major projects, well let’s just say that there is a reason people use TypeScript instead of JavaScript and Python doesn’t fall into that category.
But for most programming and for most minor systems or services, Python is just wildly good.
> there is a reason people use TypeScript instead of JavaScript and Python doesn’t fall into that category.
I thought the dynamic vs static typing debate was pointless until I joined a large Python project. Now I pretty much consider anyone that can keep up with one superhuman.
The second point matches my experience with dynamic languages as well. To use them well really takes a certain level of discipline, but it pays off. It's why I'm a bit skeptical of Python as a beginners language, which it sometimes has been touted as.
About large projects, presently I am working with Python on a somewhat large project (currently 53k sloc, I guess large is relative). Its been great, though we are pretty strict about almost absolutely everything being type-hinted.
Maybe not the best example, but I used to work with a guy who had pretty large scripts that he was writing for one of our projects. It turns out it was mostly copy-paste, so it sort of got the job done -- except he fixed bugs in one place but not in the other. The whole thing was a huge mess to understand and maintain. And he was a senior engineer at the time (now principal, to my great amazement).
Why are they writing the same code again instead of writing different code? Why aren't they learning about new APIs for topics they aren't so experienced in? What makes them "senior developers" and not "senior transcriptionists"?
Who is tasked with digging into 10 year old code, written by an ex-employee, to find a subtle bug? Who is identifying and fixing performance issues?
How are these "senior developers" able to write 1KLoC/day plus documentation and developer tests? I find test code and documentation each take as much time as writing the code in the first place.
Why aren't some of these senior developers producing negative LoC? https://www.folklore.org/StoryView.py?story=Negative_2000_Li...
For example, 1M+ LoC of code, including build system, CI, test cases, built-in documentation, in 6 months is about 7kLoC per day, divided by 5 developers it's about 1400LoC per day per developer.
I completed 30+ projects already, I have 20+ years of experience. Few years ago I was able to close up to 20 tickets per two week sprint, when I worked with same level senior developers and dedicated PM, product owner, and QA team, at large outsource company at fixed-price projects.
It's easy when tickets are properly sized and described by PM, when PO is responsive and easy to reach, and when QA covers your back for complex test cases. 6 month project + 1-2 months to recover after, and then another such project.
Those numbers are ridiculous, and I say this having 25+ years of professional experience.
The typical industry numbers are under 100 LoC/day.
Eg, slide 20 of https://www.slideshare.net/ddskier/calculating-the-cost-of-m... ("A world-class developer (e.g. Facebook or Google senior engineer) will write 50 LOC per day")
"Improving Speed and Productivity of Software Development: A Global Survey of Software Developers" at https://uweb.engr.arizona.edu/~ece473/readings/9-Improving%2... (Fig 6) has Lines-of-Code per Total Man Months at about 1,750, so about 81 lines per day (assuming 21.62 work days per month).
"A Practical Approach to Software Metrics" at https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=819938&... says "Many industry rules of thumb describe programmer productivity in terms of lines of code, for example 350 [noncomment source statements] per engineering-month of effort." That's 16 lines per days.
Going the other way, in the COCOMO estimator at https://strs.grc.nasa.gov/repository/forms/cocomo-calculatio... your 1 MLoC for an "organic" project estimates 3390 person months. Not the 30 months you mentioned.
That said, it's really easy to distort LoC measurement. And I mean beyond the "put in 1,000 lines of the form 'a=1'" cheat.
For example, LoC traditionally refers to source code, not test, CI, etc. that you use. Why did you use that non-standard definition?
Test code, for example, may contain a lot of data records and autogenerated code. This doesn't require the same work as a source line of code.
Consider SQLite. https://sqlite.org/testing.html says the library is 143.4 KSLOC, while the test suite is 91911.0 KSLOC - nearly 100 million lines of test code! These tests were not all written by hand. At that ratio of test code to source code you would have about 1,400 lines of source code.
And it's possible to mis-measure source code. A few years ago I added about 400,000 lines of code to my project in a few weeks. These were auto-generated when I replaced a very confusing set of C preprocessor directives with a homebrew template system to compile specialized versions of functions across my parameter space.
I then added some Cython projects, which generates C files about 50x larger than the original pyx files.
Counting auto-generated code makes LoC a worthless measure.
You also wrote you have a QA team. I assume their test cases are in your repo, and counted in your 1M+ LoC count, but you didn't include their time.
The post-release defect rate per KLoC is about 7.47 (average) and 4.3 (median). See An Overview of Software Defect Density: A Scoping Study at https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=6462687... .
If you write 1KLoC/day then you're likely generating 2 post-release defect bugs per day. That's in addition to bugs your QA team catches.
Your "up to 20 tickets per two week sprint" simply isn't enough to catch up with the number of bugs you likely generate.
Your numbers are so far outside of industry standards and academic findings that, with 20+ years of experience, you must know they are exceptional and difficult for anyone else to accept on your simple say-so.
> Are your fixed-price estimates based on LoC?!
I don't now. I was not a part of sales.
> Else, why aren't you embedding that knowledge in a corporate library ("contractors tools") which you license to all your future customers?
Because each customer has different requirements, different goals, different language and framework. Nobody wanted to pay for a library.
> If you write 1KLoC/day then you're likely generating 2 post-release defect bugs per day. That's in addition to bugs your QA team catches.
In practice, I had about 1 serious bug slipped per 1-2 weeks. Professional QA team known a lot of corner cases to test already. Professional developers know them too, even when writing in a different language using different framework. Moreover, we know how to prepare good CI, efficient automated tests, how to write efficient documentation, who is responsible for what, etc.
> Your numbers are so far outside of industry standards and academic findings that, with 20+ years of experience, you must know they are exceptional and difficult for anyone else to accept on your simple say-so.
How much I will be paid if you will accept my story? :-)
> Because each customer has different requirements, different goals, different language and framework
Sure, but that doesn't match your earlier statement that the senior developers 'wrote almost same code dozen times already'.
That is, I don't see how "almost the same code" is able to handle different requirements, etc.
How much would you pay me to accept your story? ;)
"People who know one language think it's the greatest in the world. People who know more than one think they all suck."
The Blub paradox: http://www.paulgraham.com/avg.html
(BTW, the TL;DR is people that know more than one language admit they all fall short one way or another, but there are some languages that are more productive than others).
“Strings and string formatting” is a very broad domain.
For most specific tasks, there is one low-impedance approach and it takes very little (but not zero) reflection and/or experience to find it.
Most new Python features directly address specific tasks for which there are currently multiple relatively high-impedance approaches taken because there is not one obviously correct way.
Python's “one obviously correct way” is not “one possible way” (the latter approach is closer to Go.)
It's not just that they are more popular, Java reigned supreme for years when Ruby and PHP and later Node.js outpaced it's communities by miles. And you can't say Java developers are not open source oriented, because they 100% are.
And the argument that programmers are simply most productive in the language they know best is trivially untrue. When I first learned of Ruby, I could dream in C#, I was absolutely fluent, knew the standard library by heart. I switched to Ruby and only looked back whenever I needed tight performance. And this is not just an anecdote, the whole Ruby on Rails movement was basically Java developers fleeing to greener pastures.
I don't know a single highly experienced multi-lingual developer who does not reach for a dynamic language when they need to deliver something quick and easy.
The only exception I know off is using Go for small networked services, which is quite comfortable and intuitive for a static language. Also outside that niche Go quickly loses its productivity edge.
Hi, highly experienced, multi-lingual developer here (Python, Ruby, PHP, JavaScript, Java, Go, C#, F#, OCaml, Elixir/Erlang, ReasonML, SQL, Smalltalk, Objective-C, Swift, leaving off quite a few prior Web 2.0 ones for brevity) nice to meet you. As you can see, I've done just about all of it: every paradigm, every syntax. It all honestly blurs together after a certain point, and it just becomes easier and easier to pick up a new language the more you learn.
Anyways, these days, I never reach for a dynamic language when I need to deliver something quick and easy. The difference is just too marginal. Not worth the downsides (and the downsides are immense!).
Because, as my experience has taught me, invariably one of those quick and easy ones will turn into something that becomes business critical, lives on for years after you are gone, and will be much more difficult for other, typically more junior, developers to update or enhance your code. Static typing is not only marginally less productive these days with all the great tools and IDE's out there (not to mention Go which has one of the least obtrusive static type checkers I've seen, or, even better if the org allows you, a language with a HM type system), but for a marginal improvement in productivity, you pay a heavy long term cost. Not worth it.
Dynamic languages are great for rapid prototypes. After that, convert it to a static language. Your junior devs that join after you, who aren't familiar with the entire ecosystem your work has will thank you.
When I want to develop further, I put a debugger statement right where I want to pick up, where all the data is available, and develop from execution.
Some static languages have REPLs but few are as good as dynamic languages for mutating existing code in the middle of a debugging session.
If you don't write much code which involves exploration of data or unknown APIs, then this mode of development may not be as useful to you, but it's a significant productivity advantage for me and the reason I reach for a dynamic language. Ruby is my go-to, with binding.pry as my debugger REPL.
I don't think you can have a hard and fast rule like this. It will depend on the situation. In many instances, a "prototype" built in Python or similar is perfectly fine.
In addition, one can make a horrendous mess in statically typed languages as well. One of the absolute worst projects I ever worked on was written in a popular statically typed language. This is of course an anecdote and not to say statically typed languages are worse, but they're not automatically better either.
Just a note too that I'm a big fan of, e.g., Rust & TypeScript, and I often use type hints in Python, so I'm not anti-static typing.
There is a floor for how bad you can write static typed code.
There is no floor to how bad of JavaScript or Python or Ruby you can write. The madness can descend to the inner most circles of coding hell. Only Perl exceeds it.
Explained:
Go's type system is generally less rigid. It strikes a good balance of strict enough. A lot of Go's converts aren't from "systems" languages like it targeted originally, but rather former Python/Ruby/PHP/JavaScript backend devs. I love the performance and low level levers I can pull with Go (although to be fair Java is quite fast enough). But finally, Go is easy to learn (26 language keywords?), the standard library is mostly great, and the worst developers I've seen still write mostly maintainable code that builds fast, which is what I optimize the most for these days.
Java, for all its warts and legacy cruft, these days you can write fairly good java, utilizing modern libraries. I love most things from Codahale, who in turn I think pushed orgs like Spring to write better libraries, so now everything's pretty good. Plus all the legacy stuff comes in handy when you have to deal with arcane government or financial systems, something I have to interface with frequently.
But if I had my choice, I'd use something where you can express functional programming concepts intuitively, without fighting the language or having to do it at a heavy performance cost, like Rust or better yet, just a full fledged FP language like OCaml (whom I understand heavily inspired Rust)
Maybe this shines light on your social circles more than actual language impact on proficiency?
The "our language is so much more productive" myth exists in many languages, including the one I currently use. But when you dig deeper, there are almost always other social or technical factors that explain the productivity gap. That kind of gap exists even between people using the same tool.
The reality of it is that assessing what makes a developer productive is incredibly hard, and people doing so to claim their language of choice is better rely on anecdata and couldn't explain what "methodology" means to save their life.
This argument is not addressing the claim for general purpose software, only for "quick and easy" software.
To people in this thread: please stop to think before responding to the "wrong point". I think we can all agree that dynamic, scripting languages are more adequate for "quick and easy" software, but this is not what the point of the discussion is. The point of the discussion is, or should be in my opinion, whether it's true that for general purpose software, written by teams, that is not trivial to write by oneself in 5 minutes, that scripting languages are more productive than statically typed, compiled languages.
I think you misunderstand the scripting language revolution of the mid 2000's. It's not that we suddenly realized scripting languages were the best for quick and easy projects. We realized scripting languages were suitable for a whole lot more than some data processing. We could much more effectively build huge scalable software platforms.
It got so bad that the sentiment flipped, and people like Joel Spolsky had to go out of their way writing blog posts that you could in fact build successful modern web platforms in C#. And then he went building the world's most popular project management tool in Node.js anyway.
If you're picking from a popular language, the language itself rarely matters for this at all. Every popular language does roughly the same things. Writing speed (length of keywords, for example) is a non-issue. The ecosystem is by far the most important thing.
It doesn't matter how efficient I am in Go if I'm missing a crucial library that I'll now have to write myself.
JavaScript is a classic example where, even if you are pretty fast at writing your code, you will be dramatically slowed down by: 1) needing to add a new package every 5 min to do something trivial, 2) looking through five half-dead libraries to find one that seems maintained and usable, and 3) finding out that some of the libraries you chose are buggy.
So for a quick prototype, I'd weight a language's qualities as follows:
- ecosystem: 90%
- static analysis/tooling: 5%
- stdlib: 3%
- syntax: 2%
And for a long-term, complex project, I'd weight them closer to:
- ecosystem: 50%
- static analysis/tooling: 46%
- stdlib: 3%
- syntax: 1%
Popularity and expansive use may have nothing to do with quality, and a lot to do with financial influence on peoples agency.
Business wants people templating out directories of performant code, not generating syntactic art for the ages.
Recent-ish commentary (July 2021) about it at https://renato.athaydes.com/posts/revisiting-prechelt-paper-... , HN commentary at https://news.ycombinator.com/item?id=28108806 .
Google Scholar gives about 476 paper which cite Prechelt's work. I have not followed other work in that field.
Just saw that, nice. I like Norvig because he really cuts to the chase.
This isn’t a knock on rust - I just think rust trades off programmer productivity for correctness and performance. And it shows, on both sides.
This seems like a pretty apples-to-oranges comparison. Rust is intended to do things you couldn't/wouldn't use JavaScript for.
If you're talking about a small-to-medium project where you don't care much about mishandling memory, of course JavaScript is going to be more productive. Rust is forcing you to tell the compiler a lot of things that JavaScript assumes you don't care about (and is, most of the time, correct).
I recall a study being quoted in my university courses, which described the amount of code needed to get certain things done between different languages, which showed that Python, Ruby and others are on the less verbose side, the reasoning being that on average they'd also be more productive.
This seems to coincide with my personal experience, where JavaScript with React was much easier to work with in smaller projects, whereas using TypeScript with React lead to much slower development because of all the typing that needed to be handled, especially in cases of union types. Now, one can say that it's worth the effort to be more confident in your code doing what you want it to both now and after X months, much like you could sometimes prefer the type systems of Java or .NET over Python, Ruby or PHP, but in my eyes those tradeoffs always come with slower development velocity.
The sad thing, however, is that DuckDuckGo (and possibly other search engines) failed to return the original study or anything like it, only resulting in low quality blog content:
- https://duckduckgo.com/?t=ffab&q=programming+language+comparison+amount+of+code
- https://duckduckgo.com/?q=programming+language+comparison+verbosity
- https://duckduckgo.com/?q=programming+language+lines+of+code+java+ruby
- http://libgen.is/scimag/?q=programming+language+verbosity
- http://libgen.is/scimag/?q=programming+language+lines+of+code
Anyone have any better search queries for this? Any idea why the search result quality is generally so low? Any ideas which study it might have been referencing?The first result for "program language comparison", at least when I view https://scholar.google.com/scholar?q=programming+language+co... is: Prechelt, Lutz. "An empirical comparison of seven programming languages." Computer 33.10 (2000): 23-29.
There's all sorts of work which cite that paper, like "An empirical investigation of the effects of type systems and code completion on api usability using typescript and javascript in ms visual studio".
For fun, here are some of those citation which themselves use the phrase "An empirical study" or similar in the title:
"An empirical study on the impact of static typing on software maintainability"
"An empirical study of the influence of static type systems on the usability of undocumented software"
"Do developers benefit from generic types? An empirical comparison of generic and raw types in Java"
"An empirical study on C++ concurrency constructs"
"An empirical study on the factors affecting software development productivity"
"An empirical study to revisit productivity across different programming languages"
I consider myself as someone who knows "far too much about C++", but recently when I had to spider some webpages and read some values out of them to populate a database, I did that in Python because it's a handful of lines due to high quality libraries that all work nicely together. I honestly wouldn't know how to do that in a short amount of time in C++.
Python and other dynamic languages don't use static types and that may give them a (very) small advantage in productivity as well, but only for very small programs... as program size increases, in my experience, statically typed languages take the lead in productivity... as OP is about writing general software, not just small scripts, I really have a hard time agreeing that Python is either the best tool for the job, or the most productive tool at all.
#include <string>
#include <cctype>
#include <algorithm>
#include <iostream>
int main()
{
std::string s("hello");
std::transform(s.begin(), s.end(), s.begin(),
[](unsigned char c) { return std::toupper(c); });
std::cout << s;
}
Same thing in Python: print('hello'.upper())In fact it's really mostly name space management. Python doesn't require you to pull in "upper" because it starts with a relatively big name space. It doesn't require you declare the main program because you're relying on the fact that Python executes a file on loading it (and that reliance is arguably bad Python style). All of the "std::" stuff is a name space management choice.
... and the rest has nothing to do with static typing either. The way the code is indented on multiple lines is a stylistic convention. The transform primitive requiring you to specify bounds is a library choice, related to the language's apparent lack of a universal way to map over a container. I'm not sure why you need the lambda. The rewrite in place is memory model stuff inherited from C.
In Haskell, which is fully compiled and is about the most rigid static typed language you could ever imagine:
import Data.Char(toUpper)
main = putStrLn (toUpper <$> "hello")
If you'd asked for something that didn't require oddball functionality, then that could probably also have been a one-liner. std::string ToUpper(const std::string& str) {
std::string as_lower(str);
absl::c_transform(str, as_lower, std::toupper);
return as_lower;
}
to enable: int main() {
std::cout << ToUpper("hello");
return 0;
}
but this would be no good! You have chosen an API that forces your users to be gratuitously inefficient. Maybe that is OK for your personal project, but an API like that has no business in the standard library.As a counter point to this - Python has a large number of C-based libraries with excellent bindings. To take Numpy has an example, you can the static-types and speed of Numpy to do all your maths, whilst Python acts as the manager. Using Numpy "feels" like normal Python, you don't really feel like you're working at native-level, but you get all the advantages of native speeds and memory usage.
And to your other point, I've worked on large C++-based projects which were a complete mess, especially when you had to make any changes to existing code. I'm not sure the language has as much an affect here as just using the correct design principles.
Not if you don't go looking for it: http://www.norvig.com/java-lisp.html
ETA: And don't get me started on Design Patterns: https://norvig.com/design-patterns/design-patterns.pdf
This is true, but most software is built by multiple people (for corporations) and will eventually be maintained by multiple other people. Whether an individual is "most productive" in a language is rarely important.
Comparing Python to Java, C#, Haskell, Erlang, Go, maybe even C or Rust, would show a much smaller performance difference.
Have you tried RAII?
Would you like to know more? https://stackoverflow.com/questions/2321511/what-is-meant-by...
Rust is a little better since at least the ownership concept is known to the compiler and automatically enforced.
(Tracing/Copying/Compacting) GC is much easier since this entire concept goes away. You always pass references to data, and the physical storage is "owned" by the GC itself. It's also much faster for certain workflow patterns, though it always consumes more memory than deterministic destruction schemes.