How many lines of C it takes to execute a + b in Python
codeconfessions.substack.com
codeconfessions.substack.com
They claimed that the hash function was used constantly —e.g. 11 times in print("hello world")—because it's used to look up object properties.
Apparently the default implementation is not optimized for performance but for security, just in case the software is exposed to the web. None of my Python programs are, so assuming all this is true, I'd much prefer to have a "I'm offline, please run twice as fast!" flag (or env variable).
https://docs.python.org/3/using/cmdline.html#envvar-PYTHONHA...
The original issue: https://bugs.python.org/issue13703
The hashseed is a per-process value, it has basically no impact on performances.
> The default hashing algorithm is currently SipHash 1-3, though this is subject to change at any point in the future. While its performance is very competitive for medium sized keys, other hashing algorithms will outperform it for small keys such as integers as well as large keys such as long strings, though those algorithms will typically not protect against attacks such as HashDoS.
https://doc.rust-lang.org/std/collections/struct.HashMap.htm...
However, Rust also lets you pick or implement your own hash algorithm if you want to optimise for your usecase.
However, Python does not use it for integers;
>>> hash(10)
10
>>> hash(100)
100
>>> hash(2**61-2) == 2**61-2
True
>>> hash(2**61-1)
0For the commonly used hash tables with prime size that use modulo to turn the hash code into a slot index, an identity hash for integers is usually fine (unless many integers are multiples of the prime size).
But other hash tables use power-of-two size to replace the modulo operation with a faster bit-and operation. Now an identity hash for integers is much more problematic, e.g. if all integers are multiples of 1000, only 1/8th of the table slots can be used.
The latter kind of hash tables would like all bits in the hash value to be well-distributed; and this is typically not true of the underlying integers. So an additional mixing operation needs to be used. Whether that mixing happens in the hash function or in the hash table depends on the implementation (for some, it's even configurable, e.g. is_avalanching marker in ankerl::unordered_dense).
>>> {i for i in range(10000)}
Takes 0.005s
>>> {i * sys.hash_info.modulus for i in range(10000)}
Takes 0.76s
I find that hard to believe that's a performance bottleneck. String hashes are all cached, and names like "print" are interned.
For a 2x overall gain I would expect to see the hash function pop up easily in my profiling, but I haven't seen it in my own profiling which was looking for simple things like that.
When siphash was evaluated, quoting https://peps.python.org/pep-0456/#performance , "In general the PEP 456 code with SipHash24 is about as fast as the old code with FNV" and "The summarized total runtime of the benchmark is within 1% of the runtime of an unmodified Python 3.4 binary".
Since then they switched from siphash24 to the faster siphash13. https://github.com/python/cpython/pull/28752
I’d use “thoroughly disproven” rather than “disputed”.
I’m sure the hash function can be changed, but as various comments noted:
- the benchmark was nonsensical
- cpython caches string hashes, and “symbols” are interned, so outside of dynamic attribute access from dynamically constructed strings each hash for attribute purposes or namespace lookup is computed once per process
- and finally (though probably not the biggest issue) xxhash is known for being mostly useful on larger sizes (>128 bytes, although you can find better hashes (city IIRC) up to 512 or so)
Much like Rust, CPython uses siphash as its default, it’s a pretty good all rounder though not the fastest. It actually used to use FNV before HashDOS.
CPython does suffer from the inability of users to configure hash functions since it’s an object property rather than a container property.
In Rust, the Hash trait is only used to feed data to a hasher such that the type can decide what should be hashed. The creation of the hasher is done by the collection, and thus provides better opportunities / flexibility in customising the hash function.
The problem is that because Python is such a ubiquitous language, CPython gets more attention than it deserves. People see it as an archetypical implementation of a scripting language. We get blogposts like this examining its inner workings, discussions about how its performance could be improved, comparisons of its speed vs. compiled languages, and tutorials on how to optimize code to run faster in it. I feel like all of this effort would be better spent on discussions about runtimes that actually try to be fast.
PyPy is in many aspects to be rated as a research project that tried a novel approach to reduce the workload compared to the manhours poured into V8,etc. LuaJIT managed with less with a focused language and a really capable lead. (Also I wouldn't be surprised if the PyPy team has also had to make compromises to get some kind of compatibility)
However it is true that this is much, much less extensive than it is in Python. As of 3.12, section 3.3 (“special method names”) of the data model documentation lists 107 entries (although some of them only apply to class protocols, and a handful are duplicates for async versions / context of sone operations).
I haven't heard anyone make this claim in a while. The inability to speed up Python beyond a certain point despite a lot of clever approaches taken was probably a good chunk of the reason, the remainder being the wall that JS has hit despite the huge effort poured into it where it is still quite distinctly slower than C.
If I were designing a language to be slow, but not like stupidly slow just to qualify as an esolang, but where the slowness still contributed to things I could call "features" with a straight face, it would be hard to beat Python. I suppose I could try to mix in more of a TCL-style "everything is a string" and "accidentally" build some features on it that bust the string caching, but that's about all I can think of at this point.
"We're going to redesign it the right way after we get this version out the door!"
I’m still finding broken print as a statement instead a function issues in codebases, somehow.
What changes would have to be made to speed it up? Obviously changing its core design now would break things, but my question is, can we can imagine an alternate universe Python that's as close as possible to our Python, except really fast? What would be different?
You can change things internally (e.g. optimizing opcode parsing), fixed object layouts, restricting mutability, converting everything to predictable array accesses, but you'll likely just end up with something like Lua or Wren rather than Python, and people like Python specifically because of the ecosystem that's built up around that dynamicism over the years.
How about objective c vs swift from a dynamic at least type vs static one. Can swift be glue.
The huge issue is that a big selling point for Python was the easy C-api integration providing lots of useful functionality via libraries now works as a chain that limits how many changes can be made (see any GIL-removal discussion).
The most sane way forward would be to mandate a conversion to a future-proof C-api (PyPy has already designed an initial one iirc that's tested and also has CPython support) that packages would convert to over time.
CPython will probably never go away due to many private users of the old api, but beginning the work towards implementation independancy in the package ecosystem at large could allow _language compatible_ runtimes with V8/JSCore/LuaJIT-like performance for most new projects.
It all depends on the entire community though and that in turn depends on the goodwill of the CPython team to support this.
The problem is that _ENV explicitly exposes the lexical environment as regular objects prohibiting optimizations, even JavaScript has _removed_ a similar feature (2: the with statement) when running in "strict" mode to simplify optimizations.
LuaJIT _could_ implement the _ENV blocks but it'd seep into large parts of the codebase as ugly special cases that'd slow down all code in related contexts (thus possibly breaking much performance for code in seemingly unrelated places to where _ENV exists).
To compare from an implementation optimization perspective, exposing _ENV is actually __worse__ than what CPython has with the GIL for example.
Luckily "with"-statements in JS was seldomly used so implementers ignored it's existence as long as it's not used(but still have to consider it in implementations, thus adding more workload), but it's an wart that will kill many optimizations if used.
For most practical purposes most people are fine without "with" or _ENV and the languages are fast enough.
https://luajit.org/extensions.html
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
If I write a function like this in Rust:
pub fn add(x: i32, y: i32) -> i32 {
x + y
}
this will compile to this assembly (on x86_64): add:
leal (%rdi,%rsi), %eax
retq
two instructions. This is because in Rust, free functions exist, have a name, and they are called by name. There's an additional twist here though too, let's check it out in debug mode, with optimizations off: add:
subq $24, %rsp
movl %edi, 16(%rsp)
movl %esi, 20(%rsp)
addl %esi, %edi
movl %edi, 12(%rsp)
seto %al
testb $1, %al
jne .LBB0_2
movl 12(%rsp), %eax
addq $24, %rsp
retq
.LBB0_2:
leaq str.0(%rip), %rdi
leaq .L__unnamed_1(%rip), %rdx
movq core::panicking::panic@GOTPCREL(%rip), %rax
movl $28, %esi
callq *%rax
ud2
There's a few things going on here, but the core of it is that in Rust, in debug mode, overflow of addition is checked, but in release mode, wrapping is okay, and so the compiler can eliminate the error path. This is an example of language semantics dictating particular implementation: if I require overflow checks, I am going to get more code, because I have to perform the check. If I do not require the checks, I get less code, because I do not perform the checks. (Where this gets more interesting is in larger examples where the checks get elided because the compiler can prove they aren't necessary, but this is already a tangent of a tangent.)In Ruby, there are no free functions. If I write a similar add function:
def add(x, y)
x + y
end
This function is not a free function: it is a new private method on the Object class. When you invoke a function in Ruby, it's not like Rust, where you simply find the function with the name you're invoking, and then call it. You instead perform "method lookup," which has some details I will elide, but for the purposes of this discussion, the idea is that you first look at the receiver to see if it has the add method defined, and then if it does not, you look at the receivers' parent class, and if it's not there, you keep going until you hit the top of the hierarchy. Once the method definition is found, you then invoke it.Now, it's not as if Rust doesn't also have method lookup (though the algorithm is entirely different), but Rust's design means that method lookup (in the vast majority of cases) is a compile-time thing: the lookup happens while you're building the software, and then at runtime, it simply calls the function that you found.
So why can't Ruby run method lookup at compile time? Well, for one, I left out an important second step: Ruby provides a method called method_missing, as a metaprogramming tool. What this means is, if we look the whole way up the object hierarchy and do not find a method named add, we will then re-traverse the entire ancestor tree again, instead invoking each class's method_missing method on the way. method missing takes the name of the method that was trying to be called, the arguments to it, and any block passed to it, and you can then do stuff to figure out if you want to handle this. This means that, even if no add function is defined, it still may be possible for the call to succeed, thanks to a method_missing handler.
Okay well why can't we do that at compile time? Well, Ruby also lets you redefine functions at runtime at basically any time. The define_method method can be called and generate a method on anything, anywhere you want, for whatever reason. You could do this based on user input, even! And yes, that would be a terrible idea, and you probably shouldn't do it, but the implementation of the language requires at least some sort of runtime computation to pull this off in the general case.
Now, I also want to point out that in my understanding, there's caching on method lookup, so that can help reduce the cost in many scenarios. But the point stands that the language has features that Rust does not, and those features mean that certain things must be more expensive than languages that do not have those features.
> can we can imagine an alternate universe Python that's as close as possible to our Python, except really fast? What would be different?
We could, but you lose compatibility with most Python code, and so you're effectively creating a new language. People do try this though, Mojo being an example of this very recently. I am excited to see how it goes.
I'm not an expert on Python but I don't see how Python is significantly more dynamic than e.g. JavaScript. I think PyPy and JS performance is comparable (or at least within the same order of magnitude), so I think it largely comes down to implementation, i.e. prioritizing performance.
I think if it had been Python (or Ruby for that matter) in the browser instead of JS, it would run about as fast as JS does today.
JS has much less in the way of magic methods that can affect "normal" object behaviour, and it doesn't have metaclasses in the way that Python does at all. Most of this customization goes unused most of the time, but the runtime still has to handle it in case it's being used this time.
Unbounded stack size is similarly difficult for WASM because like before, you have to be very careful about using the WASM stack.
Even with C++, you basically need to drop down to intrinsics or assembly to make full use of SIMD.
That is why the "as fast as C with the sufficiently smart compiler" never truly came to pass in a general sense, even if many languages that were slow to start have gotten way faster with better implementations.
And I'd love to peer in to some alternate universe that has an optimal Python interpreter (can't say "the" optimal because it's really a complicated frontier rather than a single point) that runs on our hardware and see what it looks like. Maybe even on multiple points on that frontier.
What really is the limit? I struggle to imagine what a C-speed Python interpreter could even look like, but is there some conceivable program that runs it at, say, half the speed? What even is the limit? What techniques would such a program use that would surprise us and be new and perhaps useful other places?
Or are actually pretty close to that frontier now?
To be honest, given the way optimizations tend to work, the answer probably is that we are relatively close today. They tend to have diminishing returns.
But I don't know. Is there some execution model nobody's thought of, or that has been thought of but simply hasn't had the effort invested to make it pay off, that would make huge gains? I can't prove there isn't. (In fact it's not far off junior-level computer science to prove that you can't prove it.)
I also suspect that we are relatively close, and there are diminishing returns. I suspect that you would have to start the language design with this goal in mind, and then balance out performance concerns with certain features.
I think "slow language" never meant that way. It was more like a counterpoint to the claim that there are inherent classes of languages in terms of performance, so that some language is (say) 100x or 1000x slower than others in any circumstances. This is not true, even for Python. Most languages with enough optimization works can be made performant enough that it's no slower than 10x C. But once you've got to that point, it can take disproportionally more works to optimize further depending on specific designs.
> If I were designing a language to be slow, but not like stupidly slow just to qualify as an esolang, but where the slowness still contributed to things I could call "features" with a straight face, it would be hard to beat Python.
Ironically, Python's relative slowness came from its uniform design which is generally a good thing. This is distinct from a TCL-style "everything is a string" you've said, because the uniform design had a good intent by its own.
If you have used Python long enough you may know that Python originally had two types of classes---there was a transition period where you had to write `class Foo(object):` to get the newer version. Python wanted to remove a blur between builtin objects and user objects and eventually did so. But no one at that time knew that the blur is actually a good thing for optimization. Python tried to be a good language and is suffering today as a result.
Yes, because that's also the era where this claim was flying around.
I'd say that each individual person may have their own read on what the claim meant, but certainly the way it was deployed at anyone who vaguely complained that Python was kind of slow shows that plenty of people in practice read it as I've described... that if we just wait long enough and put enough work into it, there would be no performance difference between Python and C. In 2023 we can look back with the perspective that this seems to be obviously false, so they couldn't possible have meant that, but they didn't have that perspective, and so yes they can have meant that. "Sufficiently smart compiler" was just starting to be a derogatory term. I also remember and lightly participated in the c2.com discussions on that, which may also contribute to my point that, yes, there definitely were people who truly believed "sufficiently smart compilers" could exist and were just a matter of time.
As for proportions, it's impossible to tell. Internet discussions (yea verily including this very one) in general are difficult to ascertain that from because almost by definition only outliers are participating in the discussion at all. Obviously by bulk of programmers, most programmers had simply never considered the question at all.
Maybe if we take the slowest C program, CPython no slower than 50x ?
https://benchmarksgame-team.pages.debian.net/benchmarksgame/...
fwiw PyPy doesn't seem to have been released in "the late 1990s".
Yet, not only are they in the genesis of JIT compiler research, their JITs are quite good, and their results went directly into JavaScript JITs research.
EDIT:
Also to note, those languages powered whole graphical single workstations, with microcoded CPUs + JIT.
"Efficient implementation of the smalltalk-80 system"
https://dl.acm.org/doi/10.1145/800017.800542
https://computerhistory.org/blog/introducing-the-smalltalk-z...
"Self-Confidence: How SELF Became a High-Performance Language"
We are not comparing Smalltalk JITs to Javascript JITs.
The whole point of this conversation is CPython refusing to add one, and the lame excuses regarding its dynamic capabilities, when more dynamic languages have had a JIT for decades.
Second, my question was about the possibility that Smalltalk and others were unbearably slow without JIT, so JIT was not a matter of choice for them. I'm not aware of Smalltalk implementations that don't have JIT, so it would be easier to compare Smalltalk with another well-optimized JIT implementation instead---in this case JS.
My thesis work was on AOT JS compilation, in it I refer to a bunch of experimental Python runtimes, Self,etc and the main issues in all these papers were basically the of same kind.
Heck, even PyPy exists and iirc when it comes to the core language is almost entirely compatible (except for code that relies on ref-counting semantics but that code should apparently big fixed anyhow).
https://www.pypy.org/compat.html
The real summary is: CPython is a turd in many ways, the old C api holds it back and the community hasn't put the effort into using cross-implementation compatible C-bindings instead making PyPy or others a second class citizen.
First off is the value model, the Python runtime handles ALL values as objects and that's fine for an initial naive runtime. All fast/modern language runtimes however use value models/encodings that fits "fast" values directly into machine register at the lowest level.
V8 has(had?) "small-ints" and objects (doubles,strings,etc) by setting the lowest bit in a register for pointers and otherwise dealing with them as numbers. So a+b when JIT'ed has a check (or stored knowledge from a previous opertion to elide the check) that both a and b are integers, if that is true then the actual addition is one single machine addition. if that ISN'T true then a more complex machinery is invoked that could methods like double-dispatch to see if more complex processing (like a "magic" method) is needed. This is how as JS engine handles that + behaves differently between numbers, strings, BigInt's, Date object's,etc.
(Other JS engines and LuaJIT use something called NaN/NuN tagging that also allows for quick passing of numbers w/o allocations and only a few small extra checks)
Re-implementing Python, you'd probably choose a small-int optimization (to better support Pythons seamless bigints) for values, put a runtime specific magic to the number add and make some kind of hook that detects writes to it from user code. Patching that from user code would trigger de-optimizations but for most applicataions it could continue running with optimized paths.
And even with larger objects (like heap-allocated BigInt's) a JS runtime can use inline caching to direct the runtime to fast direct dispatches, and then teams like the V8 team can detect commonly used objects that are often used and create fast-paths. A list addition for example will use common "slow" paths for dispatch but that's ok since it's an inherently slow operation that often involves allocations of some sort so the _relative_ overhead is fairly small in the big picture.
All this naturally assumes that you have the machinery in place, once in place though you can make simple code (numeric additions) fast while retaining magic for more complex objects (bigint, list,string,etc).
Tl;Dr; once you have that kind of optimizing in place, expensive processing can be allowed in special cases in slow paths thanks to type-guards, but 95% of the code will run the fast paths and having that handling in places with speed will give you most wins.
Modern tracing JIT engines indeed work by (heavy) specialization, often using multiple underlying representations for single runtime type. I think V8 has at least four Array representations? After many enough specializations it is possible to get a comparable performance even for Python. The question is how many, however.
For a long time, most dynamically typed languages and implementations didn't even try to do JIT because of its high upfront cost. The cost is much lower today---yet still not insignificant enough to say it's no-brainer to do so---, but that fact was not that obvious 20 years ago. Ruby was also one of them, and YJIT was only possible thanks to Shopify's initial works. Given an assumption that JIT is not feasible, both CPython developers and users did a lot of things that further complicate eventual JIT implementations. C API is one, which is indeed one of the major concern for CPython, but a highly customized user class is another. Herein lies the problem:
> Magic methods are not that "hard" to optimize (as long as you don't overload the add,etc operators of f.ex. the Number class in JS).
Indeed, it is very unusual to subclass `Number` in JS, however it is less unusual to subclass `int` in Python, because it is allowed and Python made it convenient. I still think a majority of `int` will use the built-in class and not subclasses, but if it's the only concern, Psyco [1] should have been much popular when it came out because it should have handled such cases perfectly. In reality Psyco was not enough, hence PyPy.
[1] https://psyco.sourceforge.net/introduction.html
At this point I want to clarify that magic methods in Python are much more than mere operator overloading. For example, properties in JS are more or less direct (`Object.defineProperty` and nowadays a native syntax), but in Python they are implemented via descriptors, which are a nested object with yet another dunder methods. For example this implements the `Foo.bar` property:
class Foo:
class Bar:
def __get__(self, obj, objtype=None): return 42
bar = Bar()
In reality everyone will use `bar = property(lambda self: 42)` or equivalent instead, but that's how it works underneath. And the nested object can do absolutely anything. You can specialize for well-known descriptor types like `property`, but that wouldn't be enough for complex Python codebases. This is why...> This is how as JS engine handles that + behaves differently between numbers, strings, BigInt's, Date object's,etc.
...is not the only thing JS engines do. They also have hidden classes (aka shapes) that are recognized and created in runtime, and I think it was one of innovations pioneered by V8---outside of the PL academia of course. Hidden classes in Python would be more complex than those in JS for this added flexibility and resulting uses. And JS hidden classes are not even that simple to implement.
After decades of JIT not in sight, and a non-trivial amount of work to get a working JIT even after that, it is not unreasonable that CPython didn't try to build JIT for a long time and the current JIT work is still quite conservative (it uses a copy-and-patch compilation to reduce the upfront cost). CPython did do lots of optimizations possible in interpreters though, many things mentioned above are internally cached for performance. One can correctly argue that such optimizations were not steady enough---for example, adaptive opcodes in 3.11 are something Java HotSpot used to do more than 10 years ago.
It's pretty well established that "Mike Pall" is the pen name for an AI sent from the future for unknown reasons. It disappeared from our light cone due to a rift in causality, presumably because it succeeded in whatever changes it wanted to make in the future.
Compared to JS pre-and-after modern engines it's a tiny improvement.
Of course, "good" and "bad" are relative. If you don't care about performance then there's nothing wrong with CPython.
I mean one doesn't need more than:
>>> exit
Use exit() or Ctrl-D (i.e. EOF) to exit
to see. They have that special handling, but they still don't want to let you out, because... IMO, they just want to be annoying.When a solution could be something like:
Note: exit() is needed in scripts. In this prompt Ctrl-D (i.e. EOF) can also be used.
Exiting.I generally agree with your overall sentiment, but I think it's important to note that the behavior is _not_ special handling; it's just the normal `repr` behavior at the REPL, where `exit` is an object like any other, and `repr(exit)` is that message.
Python has a lot excuses about the need to remain "consistent" but the reality isn't so:
>>> x = open( "/tmp/whatever123", "w" )
>>> close( x )
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
NameError: name 'close' is not defined. Did you mean: 'cosh'?
...
>>> x.close()
>>>
>>> x = "tttt"
>>> x.len()
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
...
>>> len( x )
4To compare with Javascript: if I remember, as it appeared, V8 JIT was orders of magnitude faster, compared to the interpreted code.
At the same time, I also agree with your sentiment.
[1] https://github.com/faster-cpython/ideas/wiki/Workflow-for-3....
NumPy, SciPy, TensorFlow, PyTorch, JAX, Pandas, Pillow, lxml, cjson, PyCapnP, Tornado, fast-avro, etc. all get it right. They are wrappers around C (or in some cases: Fortran/assembly/CUDA) code, where the overflow of Python method dispatch is dwarfed by the hundreds of thousands of iterations of an inner loop that's in optimized, vectorized assembly. Django, Protobufs, and Avro get it wrong (often for portability or developer velocity sake), where they wrote the whole library in Python at the expense of performance.
I was briefly tempted to write an API-compatible reimplementation of Django with the core in C++ when I left Google, but by then Django (and server-side web programming) was already falling out of favor, and if you're just shipping JSON to a SPA you can use cjson with any number of fast wsgi or asgi gateways.
I feel like as an industry we should step back and take a serious look at front end frameworks from first principles. One aspect that's clear to me, is that we should make modifications to HTML to support HTMX like transactions.
My other two main options were Tcl and Perl. Tcl was excellent at gluing, but worse at scaling, with no namespaces (then) and OO only as third-party add-ons.
Perl extensions were not so easy (better with Perl 5), and much as I enjoyed the language, handling complex data structures, was not for the faint-hearted.
Gluing pieces of C together is why we have NumPy, PyQt, pywin32, wxPython, and tens of thousands of other packages that work with C/C++/Fortran libraries.
https://news.ycombinator.com/item?id=9777816
Root of the subthread: https://news.ycombinator.com/item?id=9775799
Where is this single use case intent articulated, by whom, and what year was it? What is the use case?
Today, it seems that Python is pitched for almost everything, short of ethernet drivers.
'Python is a programming language that lets you work more quickly and integrate your systems more effectively.'
I think the "Python is pitched for almost everything" in that sentence shows a misinterpretation of gp's phrasing of "use case".
The "use case" isn't about different subject matter domains as if it was a claim about using Python as a universal language for writing database kernels or AAA games.
Instead, the "use case" is about the 2-level 2-language architecture of (1) a high-level scripting language and (2) extension modules that can be written in low-level C and imported into the interpreter. That's the "glue language" + "C Language" to combine the strengths of each language approach. (In contrast, Julia took approach of designing a language that was "fast enough" to avoid the "2 languages issue".)
>Where is this single use case intent articulated, by whom, and what year was it? What is the use case?
The Python "ergonomics use case" (not "domains use case") was originated by Python's inventor Guido van Rossum from the beginning in 1991. A clone of the 1991 Python source code has Guido's commentary for importing C modules:
https://github.com/smontanaro/python-0.9.1
Modern frameworks and libraries like TensorFlow and Pytorch continue the same use case of "high-level script glue code calling low-level C code" that was there in 1991. You can't write a tight cpu loop in pure Python code to paint a 60fps video game. That's not Python's intended use case. That philosophy is why a library like Tensorflow only has Python code for users to "glue" together the neural network graph which then calls out to the C++ code for the expensive cpu loops of backpropagation, gradient descent, etc.
I like Django. I need to process data on the server side and like to write that in python because it is more convenient than C. I also built my GUI in Django without knowing JavaScript. "just shipping JSON" seems like a different use case.
I have a piece of hardware (a laboratory hardware switch) that exposes a REST API for CRUD. I wanted to build a GUI that formats and summarizes information and offers convenient control. The data is small enough so that python can process it without becoming the bottleneck. I used Django ORM to model the data and django forms with htmx for the GUI. Authentication was easily added to Django.
The ORM part was a bit painful as Django forms expect a queryset and a queryset is not be the result of a raw sql query. There is a way to feed list of tuples into a choices argument but I decided against that , and instead dumbed down my query so I was able to write it as an django ORM language query.
https://stackoverflow.com/questions/17330158/django-how-to-u...
This is most effective for reducing startup time of short lived scripts, where the runtime is dominated by many thousands of trivial mallocs right at startup. But in general if you can establish a bound on memory, it will be faster to allocate it in one shot.
The repo states that even this dummy implementation:
> has a 60% faster startup as compared to base CPython, and in some test cases has marginally better runtime performance as well.
If I know anything about programmers, it's that everyone would just use the "go faster" flag by default.
Not only is there branches to a ton of special things but also macros that hides even more lines (IncRef/DecRef probably has a lot of magic behind there).
Nanobind is by the same author that started pybind11, used by Tensorflow and PyTorch. The web site [1] contains a bit more of the rationale.
For scientific applications you'd typically want floating points and Python's floats are just regular ieee-754 doubles (or whatever "double" meant to the compiler used to compile that python interpreter).
For example: I do a lot of work with financial software and also some basic applied cryptography. And the essential rule is to never ever use floats. Where 'decimals' are needed you want them to be simulated using integers. Python has a module called decimal which I think helps mitigate some of these issues.
I've written some code to work with accurate, large precision numbers in Python and C before. Mostly the annoying part with this is having data types for the numbers that correspond well to database fields (like uint64) or are portable (in C its easiest if you can get a u128 but this type is very compiler-specific so some hacking may be needed.)
It's fun to work on code like this but definitely needs to be precise and have good test coverage. Writing your own math libraries that are going to be used for such important operations is hair-raising stuff.
https://stratoflow.com/efficient-and-environment-friendly-pr...
And why Mojo could be an answer for high (well, higher) performance Python: https://stratoflow.com/introduction-to-mojo-programming-lang...
The entire Lua core is 15kLoC. Is that more than or less than what's needed for python's "a + b", assuming a and b are defined.
I'm genuinely curious.
To answer your question, 15kLoC is more than enough to implement dynamic dispatch and the PyObject base struct, along with the special method logic for __add__ on any python object type.. but still a lot less than what's needed for all the special method types and a lot of boilerplate for the C-compatible interface around those methods.
I understand what your clarification is trying to provide but it isn't relevant to this particular thread's article. The article is not about "performance benchmarks" where you need cpu instructions as a definitive unit-of-measure for comparisons.
Instead of measuring performance, the author's theme in this case is more akin to "decompiling" or "reverse-engineering". He takes a tiny piece of Python code and then maps it back to the actual CPython source *.c and *.h files that implements the Python vm. He added several deep links to the relevant sections of CPython source code on Github to help illustrate the mappings between Python's BINARY_OP to the .c and .h files. The article is sharing the type of knowledge you'd gain by loading up CPython in a debugger and single-stepping through the source code line-by-line.
In other words, the article's title could also have been: "Which Lines of CPython does it Take to Execute a + b in Python?"
For the scope of this particular article, the "lines of C" _are_ the defining factor because the subject of dissection is CPython's .c/.h files.
Replacing 2 lines of python code (with tens of glue code in Numba) with hundreds lines of C++ with glue code.
The actual posts are here: https://snarky.ca/tag/syntactic-sugar/ (multiple pages!)
Or, the generic, useless but correct, answer: it depends (as the linked article said, too)
If you only limit yourself to numbers (the title doesn't specify that) it should be bounded, but the article goes into some depth here, so I'll leave it at that.
typeobj->tp_as_number->nb_add()
when `tp_as_number` is a pointer to a `struct float_as_number`, and `nb_add` is a pointer to `float_add`?Do struct definitions count as "lines of code called"?
A more interesting example IMO would be something like "how many lines of C it takes to execute person.name='Bob' in Python, where person.name is undefined".
That would better demonstrate why we use Python in the first place (hint: it's not "to add integers"), while also indicating why it is slow.
The whole point of high level languages is that you put in effort upfront to make everyone elses job easier.
By you logic, using machine code directly onto toggle switches is easiest. No assembler to write, no test editor to write.....
(Jokes aside, it would really be more complicated in C if a and b were actually strings or lists.)
There's a surprising amount of depth to adding two numbers.
You don't event need strings or lists for that. Just imagine bigger numbers for a and b. Arbitrarily long integer addition is not a native language feature in C.
edit: formatting
fac = lambda n: 1 if n<=1 else n*fac(n-1)
a, b = fac(42), fac(69)
a + b
is 3 lines of python; how many lines of C code would it take to execute?Joking aside: these are not catch-all comparisons. Nor is the article. But Python is mucher slower and much safer than C. It's easier to start in Python than in C.
return 0;
in the hopes that compiles down to something like: xor rax, rax
ret#! /Usr/bin/python
Print(a+b)
V
#include <stdio.h>
Int main (){ Printf("%i", a+b); Return 0; }
And printf is basically a DSL, so it still isn't 'simple' And this is assuming a+b fits into an integer
$ cat tmp.c
main() {
printf("%d\n", 3+4);
}
$ gcc -w -o tmp.out tmp.c && ./tmp.out
7
You can even do away with types entirely, as long as you're working with ints: $ cat tmp.c
foo(x) {
return x+5;
}
bar() {
return 4;
}
main() {
printf("%d\n", foo(bar()));
}
$ gcc -w -o tmp.out tmp.c && ./tmp.out
9
Rarely a good idea, though!Anyway, my main point was that c has more boilerplate. It's never 'a+b'.
Second, printf is complicated, that's aimed more at the op though.
How much more, though? The conventional wisdom here seems to be that it's worth taking the unavoidable performance hit of dynamically typed scripting languages because the productivity boost to programmers balances it out... but I don't believe I've seen that productivity boost measured. Once you know what you're doing in C, you can do that same things you can do in Python. There's some (fascinating) syntactic sugar in there, but Python can easily be just as incomprehensible as C.