An exploration of why Python doesn't require a 'main' function
utcc.utoronto.ca
utcc.utoronto.ca
I believe this comes from the fact that local-variable lookup is faster than global-variable lookup - because the former is lookup-by-index and the latter is lookup-by-name.
The same is true in Julia: “Any code that is performance critical or being benchmarked should be inside a function.” [0]
[0] https://docs.julialang.org/en/v1/manual/performance-tips/ind...
>> Would it be OK if we add this to 3.7beta2?
> It's up to the release manager, but personally it feels like you're pushing too hard. Since you've done that frequently in the past I think we should help you by declining your request. 3.8 is right around the corner (on a Python release timescale) and this will give people a nice incentive to upgrade. Subsequent betas exist to stabilize the release, not to put in more new stuff.
>PS. I like you a lot but there's no way I'm going to "bare" with you. :-)
if __name__ = "__main__":
mainMethod()Of course I could just use a real debugger...
Edit: Just saw that ktpsns and I made pretty much the same comment. I'll add that when I learned C I found the idea of a main function so confusing at first. I wanted to know why it had to be called main and not whatever else I wanted it to be. On the plus side, it led to me exploring how compilers and linkers work, which made me understand hardware a lot better.
https://devblogs.microsoft.com/dotnet/welcome-to-c-9-0/#top-...
if __name__ == "__main__":
...
especially in contrast with c and other languages that shoehorn you into defining a main function. Namely, you can put a "main" into a module, and import that into another module without causing trouble.I use this a lot during a development cycle (not so much in prod). Say I'm trying to replace a function foo with a new implementation foo_new in a large package. I write some rudimentary tests and benchmarks in "main" which compare the two methods (including a late-import of the top-level package), and directly execute that file until I'm satisfied. At that point, I'll switch (foo, foo_new) -> (foo_old, foo) and start running unit tests. Once those unit tests pass, I remove "main" (usually, promoting some fragments of it to new unit tests) and foo_old.
i don't think that's a problem in languages with namespaces/modules -- a module might define a main(), it just wouldn't be used when it's imported. (unless you're in C, then you'd get a name clash)
Nor did BASIC.
I think the better question is, why does C (and C derived languages) require a main function?
I suppose one of the advantages is that it makes separate compilation easier. Take several .c files, compile each one separately into a .o file, and what happens when you try to link them together? Where should you start executing?
Defining a function with a well-known name is a simple, easy solution to that.
Python doesn't need to solve this problem because you give it a file to start interpreting, and that file is the entry point. You either pass the file name as an argument or you use #! and that determines it.
I think more important than having a main() function is to clearly delineate the entry point. And in Python, the entry point is the first line of the program.
Now I've noticed, at least within the Arduino environment, that it uses setup() and loop() as entry points, but code is executed before these functions are called, if for instance a global variable definition invokes a class constructor. I don't know if that's how "regular" C works, but it creates ambiguity as to what runs first.
There are other languages where it seems hard to guess where the actual entry point is located unless you know where to look.
However, many common C compilers provide methods to declare functions to run at initialization time, e.g. gcc's `constructor` attribute, opening this rabbit hole also for C programmers.
The OS determines the exact place in the code where execution starts. In DOS I think it was address 0x100. A more modern OS probably uses an executable file format that allows you to specify the starting address.
> ... why does C (and C derived languages) require a main function?
It doesn't actually. You just need an entrypoint defined for the loader. In practice most people don't want to go around libc, or fiddle with internals. This is much less true in embedded development, where main functions are sometimes missing.
Why doesn't C have global code? Probably because it would require abstracting further away from the generated assembly. If there is code that is not inside a function, how can it be called? Remember that the entry point to all programs is a function.
The following is a completely valid .f90 file that will compile and run:
program p
do i=1,10
print *, "Hello World ", i
end do
end programI compiled the above program with gfortran, and then objdump -d on it, and found this line:
0000000000001283 <main>:
Under which the cli flags are setup, and then the program is called with: callq 1199 <MAIN__>
Which is where you'll find the "true" main function.To my mind, the question is "does the programmer have to write a main function", not "does the toolchain create a main function somewhere".
There is a "true" main function for the interpreter itself. From the perspective of the operating system and CPU the python code isn't executing, the interpreter is.
But conceptually, the python code probably doesn't have main function. Python code is so far abstracted from how the CPU works that it likely doesn't need one.
How would you place Bash scripts then? Or a JITed language like Julia?
1) How does a computer start executing your code? 2) What kind of abstractions does a programming language provide?
Having anything other than a function as a top level most likely implies an abstraction. The second question is pretty much irrelevant when you are talking about interpreted languages. These aren't programs at all from the perspective of the CPU.
JIT is a different story, and one that I know very little about on a technical level. What happens when pypy generates some native code? Does it generate a function and call it? Does it have some other convention? I have no idea.
That's not the true question. The true question would be, "why does C (and C derived languages) require that all executable statements be within a function?" Requiring a main function is just a natural consequence of that.
So... why does it?
Arguably it also keeps things like libraries simpler. For example, what would it mean for me to have top-level expressions in a dynamic shared library? Does it execute as soon as I dlopen() it? Does it create an implicit "init" function that I need to call?
I used it to inject my own code in foreign executables.
C has no concept of a module A depending on a module B.
If C had top-level executable statements, there would be no defined order to them. Proof: C++ has a way of executing top-level code via file-scope definitions of class objects that have constructors and destructors. The order of these is not defined by the language; it depends on how the object files are linked and such.
Niklaus Wirth's Modula-2 language has syntax for this. Every module has an anonymous block of code that is executed at program initialization time. Modules declare dependencies on each other, which cannot be cyclic. If A uses B, then B's initialization code runs before that of A.
There is a single root module on which nothing else depends, and that module's initialization code runs last. In that initialization code, you can do the application startup; it serves as main.
I think some other languages in this broad family are that way also, like Ada.
Embedded systems written in free-standing dialects of C still choose to start with a function, though not necessarily main. For instance, Linux starts with a function called start_linux. That's not the first code in vmlinux that executes; there is assembly code above that, typically in a module called entry.S which will branch to start_linux.
This is very easy for embedded developers to deal with. There is no special run-time support for special global initialization blocks: it's just a function call named by a symbol that you can use as a branch target in the assembly code.
Even in Lisps, though there is no concept of a main function in dialects like Common Lisp, you have to specify a startup function when you save an executable image. This is because an image is not simply a collection of top-level forms that are to be evaluated, unlike a Lisp source file. It is not obvious what has to be executed when the image is restarted. Without the ability to specify a startup function, a possible choice might be this: that when the image is re-started, the save-image function (whatever it is called in the implementation) will appear to return a second time, perhaps with a value indicating "I'm now returning again due to the image being re-animated".
A compiled C program, at a very basic level, is analogous to a Lisp image.
A Lisp system loads some .fasl files, and then saves an image, specifying a startup function.
Similarly, in a C toolchain, the linker loads .o files and saves an image, recording in that image an entry point where to begin executing.
As a rule of thumb, if it's image-based, then it probably needs a startup function, unless there is an equivalent language feature for handling startup.
That applies to the booting of operating system images also. Consider:
MS-DOS has "autoexec".
Unixes run /sbin/init.
U-Boot has the bootcmd variable.
All of these are like a main startup function.
Specifying a startup function when saving an image is usually optional in Lisp.
Also the capability to save an image is optional: Lisp systems without saving images: ECL, mocl, ABCL, ...
With this in mind, it is weird to even ask for a main function. :-)
Modern PHP usage will generally concede the point now and pipe all requests through internal routing executed off of a single script exposed to the webserver/whatever.
C's entry point is in the runtime too, it's not your main function (but yeah, the optimizer will probably inline your main). At the end of the day, there's no real difference between compiled and interpreted languages, all the difference is on the toolset.
You may enjoy reading this: http://www.muppetlabs.com/~breadbox/software/tiny/teensy.htm... It explains a lot of how linux loads programs. It's a bit dated though: everything is 32 bit.
Python, Perl, Ruby, Bash and every other scripted interpreted language IMPLICITLY wraps "main" around your program based on the context, otherwise it wouldn't know where to start.
I find this article obfuscating a simple point.
EDIT: When I say .ORG, I'm referring to the standard assembly language designation for "execute this instruction first", and not ICANN. ;)
This article is bordering on nitpicking.
1. Or whatever other webserver - and actually you can directly invoke PHP via CLI, but I'm ignoring those for the most common way people tend to use it.
http://www.99-bottles-of-beer.net/language-fortran-77-264.ht...
void _start(void) {
extern int main(...);
init_clib();
exit(main(...));
}
The exact procedure heavily depends on a build target. Also, C compiler cli tool (e.g. gcc) is usually a frontend for a compiler and a linker, among other tools, which may add to the confusion.Object files are collections of symbols, either external, or defined in that file. Every non-extern symbol is a region (addr, len) with either data or code. There is no unnamed/main symbols by design. In theory, C could abstract that and allow code at column zero, but there is no rationale for that.
This whole topic is like comparing apples to oranges. The python interpreter(which IS compiled and has a main function) defines the rules of whats required in python. And it can reasonable assume that the entry point is the beginning of the .py file its running.
I can think of a real problem with your approach however: in languages with a main function you usually can't have top level statements, only declarations. So in C if you write at the top level `while (1) { printf("Hi\n"); }` you get a compilation error. Now as compilers get smarter you can get constexprs in some languages but it's still very specific and entirely static. I guess you could also mention global constructors, but that generally comes with severe caveats and should be used very carefully.
Python on the other hand treats top level code like anything else. That's why the `if __name__ == "__main__"` pattern even works in the first place. Having a real-looking main function would send the wrong message IMO, because it would make people that it's the real entry point of the program which it's not, in Python the entry point is the first line of the file being executed, end of story.
I do think that Python would benefit having a clearer and more explicitly named helper function or variable for this use case however, something like "if this_file_is_executed()" or something like that. At least this way you'd know what the code is doing instead of this awkward anti-pattern being used these days. But of course by now it's probably not worth changing.
It's a good technique prior to introducing formal unit testing (and even then, supplemental ad hoc has advantages).
2) This is not a hurdle to hello world.
print("Hello World")
is a complete program that does exactly what you want. Needing to wrap it in a function would be an additional hurdle.For complex programs; that is a hurdle well worth jumping. But for simple scripts (python is a scripting language), or teaching purposes, it is not always needed.
the current idiom uses dunder name, which is implicitly defined and looks special.
Making main special would be about the same level of a solution, yet more conventional
If “main” has to be a special function, it must keep the dunder.
I wouldn't really say that it's a hurdle to greeting the world since you can just throw `print("Hello world")` at the top level.
Here, name is the name of the module. When the module is run by itself, this is the same as __main__, in which case I want the things specified by the if to run right away. This could be calling a main function, but a lot of times it's just is some tests or examples.
When the module is imported into another file, it's name is not the same as __main__. In this case I usually only want things in it to do stuff when they are called by the other file.
This doesn't have anything to do with a function literally called "main" though.
if __name__ == "__main__":
main()
Is a common idiom, but there's no connection between the two uses of the word main.C# literally got the ability to System.out.writeLine("hello world") (probably wrong, I'm not a C# programmer) in 2020.
I suspect the answer is "we care about existing programmers more than people who are learning to program".
Of course syntactically-speaking, compiled languages could allow users to interleave these static declarations amid instructions and make it the compiler's job to correctly sort out the static from the dynamic (and indeed many languages do this); however, this makes it trickier to write good parsers and it also makes it harder for humans to distinguish at-a-glance the static from the dynamic (and this distinction is important in programming).
> I suspect the answer is "we care about existing programmers more than people who are learning to program".
Arguably if you care about people who are learning to program, you don't just punt on teaching them the distinction between static and dynamic. I think it would be better stated, "we care more about the long-term interests of novice programmers than we do about their short term frustrations".
'hinting' seems vague. Programmers who know what they're doing know that calling functions happens at run time. Obviously you can call a function without needing a wrapper function or a wrapper class + method.
It’s fine if you don’t agree with the value judgments—for example, if you would rather save a handful of characters in “hello world” at the expense of some clarity or technical simplicity. I’m just giving the rationale.
You are arguing against `print('hello world')` working.