Reminiscing CGI Scripts
rednafi.com
rednafi.com
Stuff still works just fine, a nice simple way to expose things to users without them having to ssh into a server.
Not everyone is trying to build the next adladen POS unicorn financed from abusing millions of users for a couple of cents, some are just trying to make peoples lives a little easier and make their jobs a little more efficient, and throwing some Perl, Python, PHP or whatever together often accomplishes that just fine.
FWIW/FYI, Fossil's developers do not typically run it that way - we use its builtin server primarily for the "ui" command and ad-hoc syncing across systems where setting up a web server is unnecessary or undesirable. The public-facing Fossil instances, for the core project and all of its sibling projects[^1], have Fossil running as a CGI.
That's not to say that you cannot or should not run the Fossil standalone server, just that those closest to the project typically do not run it that way (though i believe that Warren does so, via Docker, behind an nginx reverse proxy?).
[1]: that is, all projects headed up by Richard Hipp.
Having cPanel-style user instantiation and management functions then became more and more critical as IP addresses ran out and virtual hosting took hold - the amount of manual config tweaking otherwise required (vhost, database, filesystem, (s)ftp access, etc.) exceeded a rational human workload for medium sized service providers, especially those with large numbers of users spread across an array of non-uniform infrastructure.
By the time the 'cloud' arrived, this stuff at typical ISPs and web hosts had become a mish-mash of band-aids so advanced that in hindsight we might well wonder why it took so damn long for the paradigm to shift.
mod_php, on the other hand, runs inside the HTTP server process and with the same uid and capabilities, with no way to change that.
> Modern application servers like Uvicorn, Gunicorn, Puma, Unicorn or even Go’s standard server have addressed these inefficiencies by maintaining persistent server processes.
PHP and Perl were doing this 20 years ago, slightly better / more flexibly than these modern app servers. Apache2 and mod_php or mod_perl would let you choose between a threaded or forked model and give you the same choices you have today. New languages are just more trendy so they reinvent the wheel in them, but it's the same old thing.
Another advantage of using Apache was you had a ton of modules, could use multiple languages to write different web apps in one "app server", and serve both static and dynamic content out of the one server, or choose to move static out to a separate array of servers to lighten load. Today you'd use Nginx for your static content and Unicorn for your dynamic, so there's more and different moving pieces to do the same thing, and you lack features in both.
And people today while a lot about platforms and want to run their own VPS just to serve a CRUD app. But you know what we used to do? Pay $2 a month to someone who managed a web server for us, giving us a /cgi-bin/ we could upload our code to, along with .htaccess files to control Apache. Totally managed server, with a management web interface, and all you do is upload a script and you have a website. No server to maintain, no operations, just pure web dev. Now, they were laughably insecure, and things like Cron jobs and persistent dynamic applications kind of didn't exist. But as they say for VPSes today "you probably don't need all that, you just need to run a basic web app" and a PHP or Perl script was good enough for 99% of the web.
The switch was 100% made due to ease of use and lack of desire to learn/manage/debug server architecture; the slight performance boost modules offered was a small cherry on top.
Funnily enough: i once maintained a site for a client whose hoster implemented cron jobs as client-defined URLs, which could point to your CGI or PHP scripts. The provider would hit the URL (which was presumably private) on a cron-style schedule, providing a castrated semblance of cron jobs.
Weren’t most CGI scripts written in Perl? At least that’s what I remember.
PHP was basically the server-side JavaScript of its time. But as programming trends changed, and more and more went "client-side", more JS frameworks popped up, until Node appeared, and then there was no need to learn two languages to make a web app. Python and Ruby had their time, but neither were as dominant as PHP was or JS became.
In general, CGI is more secure.
They actually did, they are called serveless nowadays, and the more the merrier, taken up by a generation that didn't understand why we moved away from them.
Hence why we moved on to application servers built on top of Apache, IIS, Tomcat, and so on.
Basically what people end up doing to try to improve their serveless workloads, reduce their execution costs, is reinventing application servers, poorly.
Stuff like lets put severless in containers, managed by Kubernetes, compiled into WASM, for example.
https://www.thran.uk/cgi-win/fleg.pl
Source: https://github.com/lordfeck/cgi-win
Needs improvement but it does the job, and I had fun.
Serverless and WASM runtimes are similar and Firecracker are in this space.
It would have to be multithreaded, it could communicate with nginx via domain socket as php-fpm does. I don't know if there's a way to multiplex data over a domain socket, maybe you could use multiple domain sockets and load balance between them? Or custom framing.
Does head of line blocking come into this too?
> It would have to be multithreaded, it could communicate with nginx via domain socket as php-fpm does.
If we define the communication with nginx over domain socket to be HTTP/1.1, then you have basically reinvented Rails / Flask / Django / Hyper / ... - you can define each script as a Python/Ruby/Rust function and everything runs in shared processes so you can cache database connections and other things if you fancy.
This. It's an area where PHP really pulls ahead. php-fpm has preload, persistent db connections, worker pre-warm, etc. all built in- all of these optimisations combined makes CGI very viable.
> CGI scripts mostly went out of fashion because of their limitations around performance.
Nah, performance was fine. It's mostly always been the case that the code running on the server isn't the bottleneck.
CGI scripts went out of fashion for several reasons that fall into the categories "too complex to maintain" or "there's this new thing called PHP...".
CGI scripts were often written in Perl, but also sometimes good ol' C, and occasionally some other oddball thing. But, Perl ruled the CGI script space. Over time, CGI scripts became more complex, as software often does. When people wanted something more than a simple mail-form, or a page visit counter, or a guestbook (remember those?), they'd often be faced with downloading, installing, and then maintaining a complex Perl application, which was a nightmare. Wanna see what a big Perl application looks like? Check out https://github.com/movabletype/movabletype.
Perl just didn't lend itself to well-structured code. I say this as someone who really liked Perl, TMTOWTDI and all that, and resisted moving away from it for years. But really, once your CGI script started spanning multiple .pl files, you were gonna have a bad time.
The other thing we didn't have was good templating. Handlebars and the like now just didn't really exist then. There were all kinds of ways to sort of gin up a templating system, but there were no standards and everything was homebrewed. This was really icky if you wanted to do something like make a site with a web editor -- forget all the wysiwyg stuff, just rendering the html in a clean and safe way and spitting it back out was a bit of a faff.
Along came PHP. It had a few advantages right out of the gate: (a) you could inline it in html, and I really can't overstate just how amazing that was at the time -- suddenly you didn't need templates anymore, you just used the html you already had; (b) it came with a good enough standard library of calls that were mostly comprehensible; and (c) contrary to Perl, which Perl hackers readily referred to as "line noise", PHP's syntax was pretty clean. (Younger developers may scoff at the idea of PHP being easier to read than any other language, but it was true at the time, and older developers may grumble that it was possible to write clean Perl, and that's true too, but it's also true that Perl culture encouraged and delighted at horrid and inscrutable gibberish.)
The one other thing PHP had going for it was mod-php, which was easy for sysadmins to install and worked right alongside mod-perl, which pretty much all of them knew how to install already. So, every little web host added support for PHP practically overnight.
> When a CGI script is executed, it initiates a new process for each request. While this approach is straightforward, it becomes increasingly inefficient as web traffic volume grows.
This is true, but largely irrelevant to why CGI scripts fell out of favor. Lots of people on-prem'd or colo'd their own stuff back then (I had a beige box on an ISDN once upon a time!), and it was really hard to get enough traffic to make a machine fall over because it was spawning too many processes. Usually your bandwidth would get saturated before that happened.
There was, maybe, a brief period where this was sort of a thing, where 56k modems were everywhere that DSL wasn't and people started to pay companies to run stuff for them, but even then -- as now -- the bottleneck was usually not in the number of running processes.
> Modern application servers like Uvicorn, Gunicorn, Puma, Unicorn or even Go’s standard server have addressed these inefficiencies by maintaining persistent server processes. This, along with the advantage of not having to bear the VM startup cost, has led people to opt for these alternatives.
Heh, heh. Some people may be using those because they read on somebody's blog that everybody else is using those so they should probably use those too, but I could become a wealthy man betting $100 to every dollar that any web dev who thinks switching from LAMP, php-fpm, nginx, or whathaveyou to Spangly Animal Server is gonna be their big performance win hasn't actually done a comprehensive performance profile of their application. You are burning waaaaay more milliseconds on your JS dependencies than you are on spawning a new thread.
You would split your app across .pm files (modules) not .pl (scripts). At the control layer this meant classes based in CGI::Prototype, CGI::Application or the like. At the model layer it would be DBIx::Class or Class::DBI. View would be Template. Then in each layer you’d have a hierarchy of classes like every other language (as .pm files).
Perl has its issues but simply organizing code in a tidy way was not a problem. Problem had more to do with the readability of code, lack of basic OO facilities, awkward split between “references” (pointers) and direct values.
> Along came PHP. It had a few advantages right out of the gate: (a) you could inline it in html, and I really can't overstate just how amazing that was at the time -- suddenly you didn't need templates anymore, you just used the html you already had
While this was absolutely a win in terms of convenience when taking a site dynamic for the first time, and made PHP wildly popular especially for replacing old cgi scripts (va apps) it seems to have made PHP worse from a maintainability standpoint than well written Perl when it comes to sizable web apps.
In terms of libraries Perl had a very nice standard lib plus CPAN where the code and docs were generally to a very high standard. I don’t think PHP won at a code architecture level, it won at the very things that tended to produce the messiest, most amateurish Perl code that gave the language a poor reputation - hacking up a quick and dirty solution. It was superb at that and I admire it for that. But let’s not overstate the thoughtfulness of early PHP. It was an abomination of a language in many ways.
I was already using it a year later, for our distributed computing labs, where doing a CGI in C and Perl was part of one exercise.
This is how I remember it also. Doing CGI "well" was significantly harder. And new ways of making dynamic web content was rapidly coming out.
There were no HTTP server libraries in any languages. All web servers were large codebases. Most of the web sites ran on Apache in multiuser environments. Configuring (and compiling) Apache correctly and securely took a bit of voodoo in those environments.
The security issues with CGI wasn't just injection attacks. How should admins setup safeguards when every user could write a binary/script that remote people could execute? By the time best practices evolved PHP was taking off. And PHP offered a little bit more of a sandbox.
And then... ColdFusion. Java Servlets. JSP. ASP. And many more ways of creating dynamic websites that were often easier, had better libraries tailored to webdev, and included better sandboxing.
Web servers started to become proxies to long running processes and stateful web apps became a thing. For better or worse. (for worse IMHO. :D)
CGI with Go or Rust could be pretty interesting and significantly easier than my first C-based CGI binaries. Mostly because of their extensive web-dev library options and dependency management tools.
https://metacpan.org/pod/CGI::Application
Using that module gave a good foundation for structuring code in a maintainable way, and also providing test-cases alongside it.
One of the things I like about coding in golang is the strong emphasis on testing, but I'd say that this was also a big deal for people writing in Perl. Sure there is a reputation for line-noise, but CPAN is/was full of well-tested modules and extensions and I always appreciated the built in test::tap/prove support.
I like your comment about templating, it's something I'd never considered before. At the time I used HTML::Template, or some other module, and found it was "good enough".
At the time I remember bumping into a lot of PHP, but due to settings available to mod_php5 you'd find code that worked on one host didn't necessarily work on another host. That was something that was never a problem with perl. Though I suspect it wasn't so often that a site got so complex/slow that I had to resort to mod_perl.
I didn't know about the code readability issues, since the few perl I've seen was more readable than the PHP I saw (wp sociable plugin still haunts me).
In the end I thought it was mostly a mod perl issue causing memory bloat and security issues (and thus costs)
An older version of the article[1] also handled the headers in the web server instead of leaving it up to the script as you normally would with CGI scripts. It's really more of an example of executing a subprocess than implementing CGI.
Also, the "significant vulnerability to injection attacks" has nothing to do with CGI. It comes from taking plain text as input and then treating it as HTML without actually converting it to HTML. The solution is not to "sanitize" it but to encode it into the format you want to output.
[0] https://datatracker.ietf.org/doc/html/rfc3875.html
[1] https://web.archive.org/web/20231226033541/https://rednafi.c...
But that's kind of the point. Ye Olde "request -> stdin -> program -> stdout -> response" makes the computing piece very tech agnostic.
My second foray was using SIOD (Scheme in One Defun). It had marginal HTTP/CGI support. More than enough for my purposes: running, caching, and rendering SQL reports from text files to PDF so they can be downloaded.
Those were internal projects, nothing on the public internet. But it just helps to demonstrate the flexibility of CGI in a new world of interconnectivity when the tech and tools were still forming and cooling out of the heated plasma of the new age.
In my mind I was looking to do a bit of a compare and contrast exercise with a very modern solution like Envoy.
We would write the code in VB, compile the exe file (files? I forget), and FTP them to somewhere in c:\program files\
Pretty sure that system was a result of "when all you have is a hammer" from the previous IT manager.
Eventually the three of us managed to convince the owner that maybe RoR porting wasn't such a great idea and joined forces to make the PHP thing work.
It was the time that PHP had, basically, only Drupal and Joomla to organize a typical backend for this kind of business. Zend framework was in the early stages, but definitely didn't catch on yet. Kohana, Laravel haven't been around yet.
Now, I'm not really even a Web developer, and my familiarity with PHP is... not very deep. Before then, I'd only worked on Web projects that were in either Java, Python or RoR, and usually the Web part wasn't where I was assigned to. I'd usually work on internal company's infrastructure for the project and less on the project itself.
Anyways, the first thing that was drastically different in the existing CGI from anything I'd work on before then was the "routing". I.e. all other Web frameworks tried to keep the definition of what are the things accessible on your Web site in one place. In some hierarchical fashion. This, I believe, why the ideas s.a. REST and Swagger caught on: conceptually they resembled the way developers thought about their Web applications. If there was a $site/user URL, then it was natural and expected that $site/user/shopping-cart would be where you'd find the information about what that user wanted to buy. CGI scripts, at least those I've inherited had nothing of the sort. Since this was ASP Classic, it was just a directory with files named like shopping_cart.vb or checkout.js.
Why this was awful: in a hierarchical routing definition a developer had no problem answering questions s.a. "is this endpoint supposed to be reachable?". Also, since these routing schemes often mapped to classes / some kind of other internal program hierarchy, it was quite obvious to the developer in what state (what session variables or cookies will be present) a certain endpoint might be visited. With CGI it was neigh impossible to tell if it's legal for the particular page to be displayed when a particular session variable isn't set or not.
Especially, given the age of the project and the lack of discipline and general understanding of version control on the part of the previous authors the project often contained multiple but slightly different copies of what conceptually would've been the same page. There'd be stuff like "shopping_cart Copy (1).vb" which you'd think was definitely a garbage leftover from someone copying files accidentally by dragging them with the mouse, and you'd be wrong... There'd also be "user_login_from_company_X_using_phone.js" which you'd later discover that company X went out of business at least five years ago.
A lot of these things are automatically prevented by using frameworks, where the authors would have to expend extra effort to cause the mayhem I've witnessed in this CGI application. While frameworks "stifle the creativity" of the Web programmers by making them all follow the same beaten path, knowing the average quality of Web programmer output, when not confined to the rules of such a framework, things go south very fast.
So, in practice, I'll say, it's a good thing people don't use CGI as much anymore. Average programmer is a bad programmer. If you have commercial goals to hit, the loss of freedom and creativity are more than compensated by the better guarantees on the lower margin on the product quality.
So this was quite an interesting read and long overdue!
There's so many ways to do dynamic HTTP, the real win though is identifying how much can be done as clientside javascript so that you don't have to dynamically generate a page and just serve those as static.