Web CGI programs aren't particularly slow these days
utcc.utoronto.ca
utcc.utoronto.ca
If you don't have these issues than by all means use CGI. There is nothing wrong with it. But at some point having an in-memory cache is just going to be faster, even if your executable and various other dependence are cached by the OS page cache. Your in-memory "cache" will have all of the common data it needs pre-parsed and ready to use and may have backend connections already warmed up.
AWS lambda at least will reuse the same execution environment for multiple requests. So simple things like caching in a hashmap or using a database connection pool will work and avoid these startup costs.
You still get CGI-like behaviour on cold starts but once warm they perform much more like a persistent process. This ends up being a pretty good trade off for many use cases because you use little resources when idle and get pretty good latency and resource usage when busy.
Startup cost is potentially an issue, but the trick there is to make sure the CGI bits are tiny and can delegate heavy duty work to longer running processes. You don't have to execute the entire workflow within the CGI script.
Edit: missed a bit of context, so here's a more relevant answer:
Long lived process with extra steps, yes, but the key part is that the CGI binary (and the webserver) are sacrificial. They handle input validation and all that nasty stuff, while the real application is safely tucked away. If the CGI binary dies, or the web server crashes, things become unreachable but otherwise operational. It's occasionally a useful characteristic.
But it's been years since the last time I felt that win was sufficiently worth it to design a system that way. In theory, I still love the idea of a guaranteed clean slate, but I'm not willing to give up the slow starting language runtimes I prefer today (e.g. Ruby) for the sake of it.
That is: Compiled, statically linked code, that also took advantage of the fact the process would die (which meant e.g. not freeing memory, or doing any other unnecessary cleanup), and get external APIs out of band where possible unless their latency were low.
I built a webmail service like that back then, and we measured vs. FastCGI and other options and it performed well.
Wouldn't necessarily do it again, because it was a product of competing against slow interpreted languages on a whole server park slower than my laptop, where you could recover from the process creation overhead just by being faster.
Things like the static linking gave us a double digit percentage reduction in latency over the "naive" CGI approach, and cut the advantage of persistent processes right down. Being able to drop almost all memory frees helped a lot too (we'd only free particularly large objects, otherwise just let the OS free everything at once when the process died)
Overall we were mostly IO bound despite using CGI even with the significantly worse process creation overhead of Linux back then (it was never "bad" compared to e.g. Windows as far as I remember, but there were significant efforts to improve it)
Convert this one to fcgi, and it will get as fast as any other solution. In fcgi mode it runs as a service listening to a socket and reads and sends response over it while never terminating.
Even php running on fcgi mode is far faster than the old styled mod_php.
If you don't have enough ram for it to be in the disk cache, chances are your fcgi daemon is going to get swapped out too.
Now if your program insists on taking forever to start up, that's something that a long running daemon can help.
I haven't benchmarked, but I really don't see why php in fcgi mode would be much faster than mod_php. I've run both modes and they're both very fast. But then someone uses one of the popular frameworks and hello world takes 20-40ms, so it barely matters how you run php.
I have no figures on PHP, though if its execution model hasn’t changed from a decade ago most of this probably won’t affect it much because it effectively didn’t even persist modules. But I don’t know.
On my system in hosting (FreeBSD 13.2-RELEASE, 2x Xeon L5640), time perl -e '' with perl5-5.34.3_3 reports 3ms. On the same system time php -r '' takes 16 ms with php82-8.2.11 and a handful of libraries installed. By my own criteria, php as a cgi sucks; but I did say if your program takes too long to startup, run it as a daemon, which I think includes linking it in like mod_php or mod_perl.
On the topic of mod_php vs php_fcgi; with apache mod_php, an empty php file takes about 1.2 ms to serve, as measured by curl from when curl sends the GET request to when the content ends. I don't have a php fcgi server to try, but my recollection was something similar for minimum php output from lighttpd with php_fcgi on newer hardware.
Finally, of course, the runtime needs to start up again, compile/interpret the code, and, finally, execute it.
A binary will have the fastest startup. I'm sure there may be things that could be done. I don't know if perl mmap'd the .pl file as a readonly buffer, if each invocation would share that file (probably), and thus prevent that initial duplication. Don't know if any of the interpreters do this. Odds are this gain is lost in the whole startup of the interpreter anyways.
Also, using a binary, by a similar mechanic, lowers overall memory consumption. Have 100 connections all running the binary, they all share the code pages, and thus only "pay" for their individual data use. Share 100 connections with an interpreter, and they each pay their own cost for the entire program file plus any runtime data. Again, not so much a problem today with our ample memory resources, and mmap the source may remedy this, but this actually was an issue back in the day.
Node.js is firmly designed for daemon operation, where having a high cost of instantiating a module, something that you only do at startup, is fine, if it means you later get better throughput because you’ve done JIT compilation or whatever. (It may also just be that Node.js is unnecessarily slow at these things because no one’s cared enough. But module loading being slow is reasonably well-known in the ecosystem at large so I doubt it’s particularly that.)
PHP might have lower engine startup costs because of its history of CGI usage, and lower code loading costs because of its start-from-scratch-for-each-request model—but both are probably at a later runtime cost.
Perl might have lower engine startup costs because it’s tended to be in a similar basket to PHP, plus more general scripting use, where you just want to get through things quickly: faster for one-offs, slower for repeated things. And my understanding of Perl is that it’s somewhat closer to just being an interpreter (which is excellent for latency but terrible for throughput) than are the other languages we’re discussing.
$ hyperfine --shell=none "node -e ''" "python -c pass" "perl -e ''" "php -r ''" "bash -c ''" "echo"
Benchmark 1: node -e ''
Time (mean ± σ): 87.3 ms ± 8.0 ms [User: 69.8 ms, System: 18.4 ms]
Range (min … max): 77.9 ms … 111.1 ms 31 runs
Benchmark 2: python -c pass
Time (mean ± σ): 24.4 ms ± 5.6 ms [User: 17.5 ms, System: 6.7 ms]
Range (min … max): 15.8 ms … 31.3 ms 97 runs
Benchmark 3: perl -e ''
Time (mean ± σ): 2.8 ms ± 0.6 ms [User: 0.9 ms, System: 1.8 ms]
Range (min … max): 1.7 ms … 3.9 ms 863 runs
Benchmark 4: php -r ''
Time (mean ± σ): 19.3 ms ± 4.2 ms [User: 8.9 ms, System: 10.0 ms]
Range (min … max): 12.1 ms … 24.8 ms 241 runs
Benchmark 5: bash -c ''
Time (mean ± σ): 4.9 ms ± 0.8 ms [User: 2.8 ms, System: 1.8 ms]
Range (min … max): 2.6 ms … 6.4 ms 552 runs
Warning: Statistical outliers were detected. Consider re-running this benchmark on a quiet system without any interferences from other programs. It might help to use the '--warmup' or '--prepare' options.
Benchmark 6: echo
Time (mean ± σ): 1.2 ms ± 0.3 ms [User: 0.7 ms, System: 0.3 ms]
Range (min … max): 0.8 ms … 2.1 ms 2237 runs
Summary
echo ran
2.41 ± 0.89 times faster than perl -e ''
4.17 ± 1.39 times faster than bash -c ''
16.37 ± 5.91 times faster than php -r ''
20.74 ± 7.64 times faster than python -c pass
74.21 ± 22.49 times faster than node -e ''
(tested with node 21.5.0, python 3.11.6, perl 5.38.1, php 8.2.14, and bash 5.2.21 on Ryzen 2700X)If you trade low/no build time for high startup latency, by all means stay clear of CGI at all cost.
I've built fast CGI based systems, but not recently, because I prefer the convenience of an interpreted language, so I am not arguing you should prefer CGI.
Just that comparing a design pattern that is a consequence of not using CGI is a poor measure of CGI. If you were to seriously want to use CGIs, you wouldn't use runtimes that slow to start, or spread everything over hundreds of files, at least without a build step.
A real weakness cgi has is that if your application involves a lot of in-memory configuration that has a terrible impact on startup time even if it has an excellent effect on execution time.
Typo?
fastcgi c/c++ server development was essentially stopped many years ago, would love to see a c version that I can use for embedded boards, though cgi does the job 90% of the time just fine.
https://svn.apache.org/viewvc/httpd/httpd/trunk/
It’s a fork of what you linked (and was more popular afaik back when fastcgi was state of the art, and apache was the undisputed champion of web servers).
These days, nginx has more market share than apache, and its fastcgi module is one of the more recently updated ones in its source tree (5 months vs multiple years):
https://github.com/nginx/nginx/tree/master/src/http/modules
If I was going to build an embedded web server, I’d start with nostd rust, probably with though axum + tokio, since thats already memory safe-ish.
If I needed fastcgi for some reason (dynamically loadable endpoints, or os-level isolation), there are at least four implementations of fastcgi for it. No idea if any are decent though.
Maybe not by much though.
https://www.netcraft.com/blog/december-2023-web-server-surve...
What's the point of using this over serving HTTP directly? The whole point of CGI was providing an easy to implement interface to proper HTTP servers, but FastCGI is such as step up in complexity compared to CGI that you're going to need a library anyway.
Maybe FastCGI made sense back when companies tried to convince us we needed their heavy-duty software to handle any HTTP traffic at all, but now HTTP libraries and reverse proxies are dime a dozen.
* I agree that you often will want a server such as Caddy/nginx/Apache httpd to be the Internet-facing http server. They give you a lot of nice things like DoS resistance, logging, robust TLS implementations, and dispatch to all the various applications you might be running.
* But you need some protocol between that proxy and your application. Like debugnik, I'd prefer to just use HTTP as that protocol rather than something else like FastCGI. I already have to be familiar with HTTP/1.1 in the sense that I both actually know the wire format by heart and know of a huge variety of tools that can directly speak it. So I'd rather just have my application speak HTTP over a Unix or IP socket that my frontend webserver proxies to. "Each HTTP-serving lib is another liability" is at least as true for "each FastCGI-serving lib" given that the FastCGI libs will be obscure (less well-tested and maintained) and (due to poorer familiarity/tooling) harder to debug.
True-ish, but complicated a bit because (as far as I can tell) the FastCGI spec is tiny compared to HTTP... even just HTTP 1.1, let alone 2 and (yikes) 3.
That's true-ish also. For this back leg, it's fine to only support HTTP/1.1; the frontend server can handle translation from HTTP/{2,3} as necessary. And the HTTP/1.1 spec defines a lot of stuff like the syntax/semantics of the headers and semantics of different request methods that the FastCGI spec doesn't address. It's still probably true that the HTTP/1.1 wire format a bit more complex than FastCGI, but it's much closer than it might first appear. On balance I'd still take HTTP any day.
[EDIT] And, as noted in my response to the other reply here, the size of the FastCGI spec, and the size of implementations, are far smaller than HTTP, which is nice.
I have an internal app that runs as a CGI but is written in a compiled language. It averages 100ms response times, mostly due to database queries.
I wouldn’t write a major app using cgi but for small sites and internal services it’s fine
If it's then also prepared properly, e.g statically linked position independent code, you won't even have relocation overhead and the only thing you "pay" is the process creation overhead, which hasn't been a major issue on at least Linux for a couple of decades at least.
Whether the difference will matter really depends on your specific runtime. With an interpreter, startup can easily bite you hard, sure.
PHP is a bad example for that reason - the moment you need a bunch of extra syscalls on startup, the context switch latency alone will kill performance vs. what you can achieve.
That's where the author is wrong. Complex Python, Ruby or Perl scripts can easily take several hundred milliseconds just to import required modules.
Edit: see the responses for why “not necessarily terrible” is kind of annoying, now I have a comment that says there exists a use case where they care about it, and there exists one where they don’t.
The last time I used CGI was for the "sign up to my newsletter" on my website; the endpoint for that would lock a file, append the email from the form, and close it. Even if that takes a second (which it didn't, more like 50ms) then ... that would actually be fine for this particular use case of a small newsletter.
Of course for many other things it would be horrible.
I think "not necessarily terrible" is a reasonable way to phrase things.
It's slightly ironic to be reading this phrase via the front page of HN.
(I am the author of the linked-to article and also the author of the software it's running. Said software also has (on-disk) caching, but that's not why the last-modified is back in December of last year.)
If a page is using CGI, it's often (but not always) because it's interactive in some way. If users are interacting and generate 10 requests total, on average, you're down to only 100ms per request (which is about how long it takes my system to spin up a new python3 process and `import numpy`).
Anyone remember when you offloaded something to the server to make it faster? Now it's like "well this is blazing fast on my machine, but we'll see how it does when it's deployed to prod..."
All that monolithic kernel performance wasted in endless layers of virtualization.
[EDIT] Incidentally, I don't hate all this stuff—I like docker-as-a-cross-distro-package-and-process-manager better than most of the rest Linux has to offer on that front, I find k8s kinda-nifty, but wish it 1) didn't always seem to involve so much goddamn yaml, which is so shockingly horrible that I can't believe it's still in so many things, and 2) were better-integrated with the underlying operating system(s) rather than reinventing so very many wheels and being another layer, rather than part of the system. Like I can vaguely imagine, through a hazy cloud of dream, some kind of hypothetical k8s-alike integrated into FreeBSD that I'd absolutely love.
We were serving hundreds of hits per second on 200MHz web servers in the late 90s. Everything is _so much faster_ now that surely the performance must be acceptable nowadays?
In fact, I drastically prefer having no front-end requirement. Even if node is "slow" compared to some C FastCGI benchmark, I like the operational simplicity of deploying code onto some stripped down Linux with minimal explicit dependencies. Go offering a single binary is even more attractive. I was always frustrated running Apache/Nginx, then maybe some process manager like PPM, fighting with the configuration across a bunch of different formats.
While I admit that modern container based orchestration can be overly-complex, it doesn't leave me nostalgic for the days of CGI.
For my customers, no CGIs. They do run on Python and Ruby though.
I think at the extremes process-per-request and persistent processes look very similar.
A fresh process will read data from the disk cache, but it may need to parse or transform that data before it is usable. A persistent process will have (some) pre-parsed data in memory. If you have swap enabled it will also be swapped out to disk if the process is idle. This may be faster if parsing that data into an in-memory structure is slow or slower if this results in more IO than the source data (which may be stored more efficently).
Even on "slow" hardware, they could often run in about a second. Which was still fast enough for a small business with a form that gets submitted a few times per day.
For a lot of a small sites, needing to serve more than one request per second would be a great problem to have.
So that's a reason not to not use it; but are there any reasons to use it?
Even in TFA, he demonstrates using a simple Golang CGI binary; but writing a simple Golang server that you pass through as a reverse proxy is just as simple.
Apache had an option to run the scripts as your user, rather than default of "nobody". That allows segmenting the processes of different users on the same server.
CGI's biggest weakness is the IO & CPU needed to "boot" the interpreter. If you can keep the interpreter resident, you can reach acceptable performance.
Like an array of Haskell CGI processes or something...
You could redesign CGI to change this, of course. But at that point, you've lost compatibility with existing CGI scripts, and you might as well use FastCGI and/or HTTP over a socket instead.
No, it absolutely would not be worthwhile to go through such hoops, but fun thought experiment.
(only slightly tongue-in-cheek)
https://www.theserverside.com/blog/Coffee-Talk-Java-News-Sto...
Although CGI is a defined protocol. Serverless is a more broad term, not limited to any particular protocol.
you can serve most moderate-traffic services with any tech you like
its time for smart people to move on to other problems