Convert macros to functions in the Python C API
python.org
python.org
In my case it's .NET and P/Invoke, where I'm facing challenges similar to those faced by Python.NET, but the same things apply to other FFIs, (although they are ameliorated if you assume a common, system-wide C compiler as is usually the case with linux, which is probably why this case seems rarely to be considered).
Macros are zero use in these situations. There should be exported functions corresponding to these. This PEP sounds like a move in the right direction, but I don't see it mentioned that these macro-replacing inline functions will also be exported from the library so they're accessible with GetProcAddress (windows) or dlsym (linux). If they will be, it will make me very happy.
There's an asssumption in cpython "C API" development that clients are only written in C/C++. Even the Stable ABI and its improvements (which have been a great boon) are only discussed in terms of the ABI not changing, never in terms of actually defining an ABI. It would be great to have an interface you could interact with via an FFI without having to make assumptions about how it was compiled, even if it involves stuff like calling a bunch of functions to discover how big a ssize_t is, etc.
> When a C compiler decides to not inline, there is likely a good reason. For example, inlining would reuse a register which require to save/restore the register value on the stack and so increase the stack memory usage or be less efficient.
> On the other side, the Py_NO_INLINE macro can be used to disable inlining. It is useful to reduce the stack memory usage.
This seems completely backwards, what are they talking about?
> When a C compiler decides to not inline, there is likely a good reason. For example, inlining would reuse a register which require to save/restore the register value on the stack and so increase the stack memory usage or be less efficient.
And ... that removes all doubt. They are wrong. If a calculation requires an extra register, doing a function call won't conjure it out of thin air. It has to spill the register too, and it will push the PC along with it.
It's still possible a doing function call rather than inlining will speed up code on modern CPU's. The repeated code the inlining generates extra demands on the caches.
flatten
Generally, inlining into a function is limited.
For a function marked with this attribute, every call
inside this function is inlined, if possible.
Functions declared with attribute noinline and similar are not inlined.
Whether the function itself is considered for inlining depends on
its size and the current inlining parameters.Long running function calling short lived but high memory use function is a bad candidate for inlining because of that.
There are three cases: neither `b` or `c` is inlined; both are inlined; or only one is inlined. If neither is inlined, the memory is always reused. If both are inlined, the memory is usually reused but not always. If only one is inlined, the memory is never reused.
Therefore, marking functions no-inline can indeed reduce stack usage, but it depends on the situation.
Details:
Case 1: Neither `b` or `c` is inlined. Then `b` will push its stack frame and pop it when it's done, then `c` will do the same with its stack frame, reusing the same memory.
Case 2: Both `b` and `c` are inlined. Then both of their variables will be part of `a`'s stack frame. A naive compiler would put each variable in a separate location in the stack frame, so the size of the stack frame would be at least the sum of `b` and `c`'s variables, wasting stack space. However, most compilers can determine that the variables' lifetimes don't overlap and reuse the same part of the stack frame for both. (LLVM calls this "stack coloring", for reference.)
Most compilers, but not all. In a simple test on gcc.godbolt.org, GCC, MSVC, and Clang all normally perform this optimization at all optimization levels [1]. But ICC (Intel C compiler) fails to perform it, allocating space for both variables even at -O3. And there are many more obscure C compilers (not that commonly used these days, but they exist), some of which presumably have the same problem.
Case 3: One of `b` and `c` is inlined but the other isn't. Suppose `b` is the one inlined. `b`'s variable will be incorporated into `a`'s stack frame, but when `a` then calls `c`, `c` will push its stack frame on top of `a`'s. In theory, `a` could dynamically reduce the size of its stack frame before calling `c`, but as far as I know, no major compilers do this, regardless of the optimization level. Therefore, the memory cannot be reused.
[1] Test case: https://gcc.godbolt.org/z/nY4ddz7q1
Note: Clang actually doesn't perform the optimization at -O0, but at -O0 functions are never inlined unless forced to be with always_inline, which should be used sparingly. So it's not a concern in typical situations. MSVC, for its part, doesn't inline functions at /O0 even if they are marked __forceinline.
Actually just the name is undefined behavior per C and C++. The standards reserve all identifiers starting with underscore and followed by a capital letter to the implementation.